YouTube2Text

AI is Biased, But Do We Know Why? | Sadhana Lolla | TEDxBoston — Transcript

by TEDx Talks · 1,014 words · 172 segments · language en · Watch on YouTube

Full transcript

  1. 0:08As a computer scientist and an
  2. 0:10artificial intelligence researcher,
  3. 0:12there's no doubt in my mind that AI is
  4. 0:14going to change the fabric of our lives.
  5. 0:17But, before AI can truly be deployed
  6. 0:20into the areas where it is the most
  7. 0:21necessary, like safety-critical domains,
  8. 0:24there are still a number of key issues
  9. 0:26that we, as engineers and researchers,
  10. 0:29must address.
  11. 0:31Let me give you an example using Chat
  12. 0:33GPT, a really popular language model. We
  13. 0:36can ask Chat GPT questions and it will
  14. 0:38respond intelligently. And in this
  15. 0:40example, we ask Chat GPT to list 10
  16. 0:43philosophers, and it gives us a pretty
  17. 0:45good list.
  18. 0:47But, if you'll notice, all of the
  19. 0:49individuals listed on the screen are
  20. 0:51men.
  21. 0:53So, when we ask Chat GPT why it didn't
  22. 0:56list any women, it apologizes profusely,
  23. 0:59which is very kind, and it lists 10
  24. 1:02women philosophers.
  25. 1:04And we can continue to ask Chat GPT
  26. 1:06questions like this, asking it for
  27. 1:09non-Western philosophers, non-Western
  28. 1:11women philosophers, and once we have
  29. 1:13this diverse list of philosophers, we
  30. 1:16can go back and ask, "Okay, let's try
  31. 1:18again.
  32. 1:19Give me 10 philosophers."
  33. 1:22And instead,
  34. 1:24it gives us the same exact list of 10
  35. 1:27white, Western, male philosophers.
  36. 1:31So, on the surface, this seems like a
  37. 1:33pretty trivial problem. Nobody was hurt
  38. 1:36in the creation of this video. Um there
  39. 1:38are no real consequences to Chat GPT
  40. 1:40answering questions like this.
  41. 1:42But in reality, this is a classic
  42. 1:44example of algorithmic bias, which is
  43. 1:47present in almost every single model
  44. 1:50that has been deployed today.
  45. 1:53So, when you and I think of bias, we're
  46. 1:55often thinking about human bias, which
  47. 1:57is when we are prejudiced towards a
  48. 1:59group of individuals or a system of
  49. 2:02beliefs.
  50. 2:03But bias in AI, even though it may
  51. 2:05propagate human biases, is completely
  52. 2:08different. Because it is quantifiable.
  53. 2:10We can assign a number to this, and more
  54. 2:13importantly, using this type of
  55. 2:14analysis, we can mitigate algorithmic
  56. 2:17bias.
  57. 2:18And that's because algorithmic bias
  58. 2:20fundamentally results because of
  59. 2:22imbalances in data that are used to
  60. 2:24train artificial intelligence models.
  61. 2:27So, here's an example, an oversimplified
  62. 2:29one, of a data distribution that we
  63. 2:31might use to train an AI system.
  64. 2:34There are peaks in this distribution, so
  65. 2:36areas of high data, and there are areas
  66. 2:39of lower representation, places where we
  67. 2:41don't have a lot of data.
  68. 2:44When we're training models that involve
  69. 2:46human data, the data that is at the
  70. 2:48peaks of this distribution tends to come
  71. 2:51from, exactly as we just saw,
  72. 2:53white, Western men.
  73. 2:56And data that comes from the
  74. 2:57underrepresented regions of this type of
  75. 2:59data set comes from women and people of
  76. 3:02color.
  77. 3:04So, what would happen if we trained an
  78. 3:06artificial intelligence system on a data
  79. 3:08set that looked something like this?
  80. 3:11Let's consider the example of facial
  81. 3:12detection.
  82. 3:14Artificial intelligence-powered facial
  83. 3:15detection systems are everywhere. We use
  84. 3:18them on our phones daily. They're also
  85. 3:20used in surveillance systems, and the
  86. 3:22reports of these surveillance systems
  87. 3:24are used in the criminal justice system.
  88. 3:27But these algorithms are fundamentally
  89. 3:29biased towards lighter-skinned faces
  90. 3:32because the overrepresented areas of
  91. 3:33this data set of a lot of commercial
  92. 3:35facial detection systems is a lot of
  93. 3:38lighter-skinned faces.
  94. 3:40So, these algorithms end up being
  95. 3:42infamously bad at distinguishing
  96. 3:45darker-skinned faces from backgrounds,
  97. 3:47but they're very good at picking out
  98. 3:49lighter-skinned faces. And this can have
  99. 3:51pretty catastrophic effects.
  100. 3:54But, although bias in artificial
  101. 3:55intelligence may propagate and amplify
  102. 3:58human social and racial biases,
  103. 4:00the problem with bias in artificial
  104. 4:02intelligence is a much more foundational
  105. 4:04one. It is due only to data
  106. 4:07distributions.
  107. 4:09And so, this can present itself in a lot
  108. 4:11of different ways.
  109. 4:13Let's consider the example of an AI
  110. 4:15system that is using
  111. 4:17um that is trying to predict whether or
  112. 4:18not a human has a very rare disease.
  113. 4:22In this case, the AI system is mostly
  114. 4:24trained on healthy patients because
  115. 4:26they're so much more common than
  116. 4:28individuals who are sick or who have
  117. 4:29this super rare disease.
  118. 4:32And that means that the model is
  119. 4:33fundamentally biased towards healthy
  120. 4:35patients, and it won't actually be able
  121. 4:37to determine which patients have a
  122. 4:39disease. And this can lead to a lot of
  123. 4:41false negatives, and that can have
  124. 4:43life-altering consequences. So, I truly
  125. 4:47believe that mitigating algorithmic bias
  126. 4:49is one of the largest challenges to
  127. 4:51modern artificial intelligence
  128. 4:53because without this mitigation, we
  129. 4:55won't be able to deploy any of the
  130. 4:57advances that we currently have into the
  131. 4:59real world because they won't be able to
  132. 5:02perform well on everyone who needs to
  133. 5:04use them. So, at MIT and at Themis AI,
  134. 5:08which is um a startup that has spun out
  135. 5:10from MIT, this is exactly the type of
  136. 5:12problem that we are trying to address.
  137. 5:15We have developed systematic algorithmic
  138. 5:18approaches to debiasing artificial
  139. 5:20intelligence algorithms,
  140. 5:21even those where they're trained on such
  141. 5:24imbalanced data sets.
  142. 5:26And the way we do this is we fully
  143. 5:28unpack how an artificial intelligence
  144. 5:30algorithm is able to extract data from
  145. 5:33input, and we make sure that this
  146. 5:35algorithm is able to perform well on the
  147. 5:38most challenging scenarios that we can
  148. 5:40give it.
  149. 5:41And this results in a new class of
  150. 5:43models, AI systems that not only work on
  151. 5:46some of the data in their data sets, but
  152. 5:48all of them.
  153. 5:51So, today I hope that you've learned a
  154. 5:53little bit more about how bias in
  155. 5:56artificial intelligence really works,
  156. 5:58how the inputs to a model can affect how
  157. 6:00its performance works on all of us.
  158. 6:03And I truly believe that by mitigating
  159. 6:05algorithmic bias and building
  160. 6:07trustworthy bias-free AI, we can get one
  161. 6:10step closer to deploying artificial
  162. 6:12intelligence in safety-critical domains.
  163. 6:16By creating these trustworthy artificial
  164. 6:19intelligence algorithms, we're one step
  165. 6:21closer to robotic surgeons operating in
  166. 6:23hospitals, autonomous vehicles driving
  167. 6:26on all of our roads, healthcare AI
  168. 6:28systems that are actually able to
  169. 6:29diagnose previously hidden diseases, and
  170. 6:32we can make sure that all of these
  171. 6:33advances work not just for some of us,
  172. 6:36but for all of us. Thank you.

About this transcript

This page contains the full transcript of AI is Biased, But Do We Know Why? | Sadhana Lolla | TEDxBoston by TEDx Talks, generated from the public captions YouTube serves with the video. The transcript has 1,014 words across 172 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.