AI is Biased, But Do We Know Why? | Sadhana Lolla | TEDxBoston — Transcript
Full transcript
- 0:08As a computer scientist and an
- 0:10artificial intelligence researcher,
- 0:12there's no doubt in my mind that AI is
- 0:14going to change the fabric of our lives.
- 0:17But, before AI can truly be deployed
- 0:20into the areas where it is the most
- 0:21necessary, like safety-critical domains,
- 0:24there are still a number of key issues
- 0:26that we, as engineers and researchers,
- 0:29must address.
- 0:31Let me give you an example using Chat
- 0:33GPT, a really popular language model. We
- 0:36can ask Chat GPT questions and it will
- 0:38respond intelligently. And in this
- 0:40example, we ask Chat GPT to list 10
- 0:43philosophers, and it gives us a pretty
- 0:45good list.
- 0:47But, if you'll notice, all of the
- 0:49individuals listed on the screen are
- 0:51men.
- 0:53So, when we ask Chat GPT why it didn't
- 0:56list any women, it apologizes profusely,
- 0:59which is very kind, and it lists 10
- 1:02women philosophers.
- 1:04And we can continue to ask Chat GPT
- 1:06questions like this, asking it for
- 1:09non-Western philosophers, non-Western
- 1:11women philosophers, and once we have
- 1:13this diverse list of philosophers, we
- 1:16can go back and ask, "Okay, let's try
- 1:18again.
- 1:19Give me 10 philosophers."
- 1:22And instead,
- 1:24it gives us the same exact list of 10
- 1:27white, Western, male philosophers.
- 1:31So, on the surface, this seems like a
- 1:33pretty trivial problem. Nobody was hurt
- 1:36in the creation of this video. Um there
- 1:38are no real consequences to Chat GPT
- 1:40answering questions like this.
- 1:42But in reality, this is a classic
- 1:44example of algorithmic bias, which is
- 1:47present in almost every single model
- 1:50that has been deployed today.
- 1:53So, when you and I think of bias, we're
- 1:55often thinking about human bias, which
- 1:57is when we are prejudiced towards a
- 1:59group of individuals or a system of
- 2:02beliefs.
- 2:03But bias in AI, even though it may
- 2:05propagate human biases, is completely
- 2:08different. Because it is quantifiable.
- 2:10We can assign a number to this, and more
- 2:13importantly, using this type of
- 2:14analysis, we can mitigate algorithmic
- 2:17bias.
- 2:18And that's because algorithmic bias
- 2:20fundamentally results because of
- 2:22imbalances in data that are used to
- 2:24train artificial intelligence models.
- 2:27So, here's an example, an oversimplified
- 2:29one, of a data distribution that we
- 2:31might use to train an AI system.
- 2:34There are peaks in this distribution, so
- 2:36areas of high data, and there are areas
- 2:39of lower representation, places where we
- 2:41don't have a lot of data.
- 2:44When we're training models that involve
- 2:46human data, the data that is at the
- 2:48peaks of this distribution tends to come
- 2:51from, exactly as we just saw,
- 2:53white, Western men.
- 2:56And data that comes from the
- 2:57underrepresented regions of this type of
- 2:59data set comes from women and people of
- 3:02color.
- 3:04So, what would happen if we trained an
- 3:06artificial intelligence system on a data
- 3:08set that looked something like this?
- 3:11Let's consider the example of facial
- 3:12detection.
- 3:14Artificial intelligence-powered facial
- 3:15detection systems are everywhere. We use
- 3:18them on our phones daily. They're also
- 3:20used in surveillance systems, and the
- 3:22reports of these surveillance systems
- 3:24are used in the criminal justice system.
- 3:27But these algorithms are fundamentally
- 3:29biased towards lighter-skinned faces
- 3:32because the overrepresented areas of
- 3:33this data set of a lot of commercial
- 3:35facial detection systems is a lot of
- 3:38lighter-skinned faces.
- 3:40So, these algorithms end up being
- 3:42infamously bad at distinguishing
- 3:45darker-skinned faces from backgrounds,
- 3:47but they're very good at picking out
- 3:49lighter-skinned faces. And this can have
- 3:51pretty catastrophic effects.
- 3:54But, although bias in artificial
- 3:55intelligence may propagate and amplify
- 3:58human social and racial biases,
- 4:00the problem with bias in artificial
- 4:02intelligence is a much more foundational
- 4:04one. It is due only to data
- 4:07distributions.
- 4:09And so, this can present itself in a lot
- 4:11of different ways.
- 4:13Let's consider the example of an AI
- 4:15system that is using
- 4:17um that is trying to predict whether or
- 4:18not a human has a very rare disease.
- 4:22In this case, the AI system is mostly
- 4:24trained on healthy patients because
- 4:26they're so much more common than
- 4:28individuals who are sick or who have
- 4:29this super rare disease.
- 4:32And that means that the model is
- 4:33fundamentally biased towards healthy
- 4:35patients, and it won't actually be able
- 4:37to determine which patients have a
- 4:39disease. And this can lead to a lot of
- 4:41false negatives, and that can have
- 4:43life-altering consequences. So, I truly
- 4:47believe that mitigating algorithmic bias
- 4:49is one of the largest challenges to
- 4:51modern artificial intelligence
- 4:53because without this mitigation, we
- 4:55won't be able to deploy any of the
- 4:57advances that we currently have into the
- 4:59real world because they won't be able to
- 5:02perform well on everyone who needs to
- 5:04use them. So, at MIT and at Themis AI,
- 5:08which is um a startup that has spun out
- 5:10from MIT, this is exactly the type of
- 5:12problem that we are trying to address.
- 5:15We have developed systematic algorithmic
- 5:18approaches to debiasing artificial
- 5:20intelligence algorithms,
- 5:21even those where they're trained on such
- 5:24imbalanced data sets.
- 5:26And the way we do this is we fully
- 5:28unpack how an artificial intelligence
- 5:30algorithm is able to extract data from
- 5:33input, and we make sure that this
- 5:35algorithm is able to perform well on the
- 5:38most challenging scenarios that we can
- 5:40give it.
- 5:41And this results in a new class of
- 5:43models, AI systems that not only work on
- 5:46some of the data in their data sets, but
- 5:48all of them.
- 5:51So, today I hope that you've learned a
- 5:53little bit more about how bias in
- 5:56artificial intelligence really works,
- 5:58how the inputs to a model can affect how
- 6:00its performance works on all of us.
- 6:03And I truly believe that by mitigating
- 6:05algorithmic bias and building
- 6:07trustworthy bias-free AI, we can get one
- 6:10step closer to deploying artificial
- 6:12intelligence in safety-critical domains.
- 6:16By creating these trustworthy artificial
- 6:19intelligence algorithms, we're one step
- 6:21closer to robotic surgeons operating in
- 6:23hospitals, autonomous vehicles driving
- 6:26on all of our roads, healthcare AI
- 6:28systems that are actually able to
- 6:29diagnose previously hidden diseases, and
- 6:32we can make sure that all of these
- 6:33advances work not just for some of us,
- 6:36but for all of us. Thank you.
About this transcript
This page contains the full transcript of AI is Biased, But Do We Know Why? | Sadhana Lolla | TEDxBoston by TEDx Talks, generated from the public captions YouTube serves with the video. The transcript has 1,014 words across 172 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.