Get Full Government Meeting Transcripts, Videos, & Alerts Forever!
Get email alerts on the Dataset Bias topic
No spam. Unsubscribe anytime.
How training data and model design create bias, experts say
Summary
Professor Dave Gadeau explained that biases in training datasets and sampling choices drive many common AI failures, citing examples from maps, baby photos, hiring tools and medical imaging, and noting fixes require data and design changes.
Get email alerts on the Dataset Bias topic
No spam. Unsubscribe anytime.
Dave Gadeau, a computing science professor at Finger Lakes Community College, told attendees that the core cause of many AI failures is the makeup of training data and the statistical sampling approach models use. "AI has examined all the information and what we call the training..." he said, describing how models ingest text, images and audio and learn probabilistic associations that later determine outputs.
He illustrated consequences with several concrete examples: the prevalence of Mercator maps in training data can skew geographic representation, baby‑photo color distributions produce gendered image outputs, and resume corpora reflecting past hiring practices led an Amazon screening prototype to favor male patterns. He said correcting these problems requires curated datasets, careful label design, and targeted prompting or model constraints rather than only relying on downstream filtering.
