YouTube2Text

Anthropic's Chloe Lubinski explains how AI works (in 14 minutes) — Transcript

by Alliance for Responsible Citizenship · 2,274 words · 387 segments · language en · Watch on YouTube

Full transcript

  1. 0:00I work at Anthropic, which uh where I
  2. 0:03lead the the research partnerships with
  3. 0:06the world's wisdom traditions.
  4. 0:08And my job really has two parts to it.
  5. 0:11So, the first, it's to help the these
  6. 0:13experts in these various fields and
  7. 0:15disciplines actually understand AI, what
  8. 0:18it is, what's happening now, and where
  9. 0:20it's going.
  10. 0:22And the second part is to listen and to
  11. 0:23learn and to funnel wisdom back into the
  12. 0:26organization, back to the people that
  13. 0:28are building this technology.
  14. 0:30Just last week, I was walking my little
  15. 0:33red-haired cocker spaniel in San
  16. 0:34Francisco, thinking of what might be
  17. 0:36most helpful to you today in having
  18. 0:39these conversations. And the thing is,
  19. 0:41I've probably had hundreds of
  20. 0:43conversations now across 20 or so
  21. 0:46traditions and disciplines, and I found,
  22. 0:49again and again, just how important it
  23. 0:52is for folks to really understand the
  24. 0:54basics before we can even start to talk
  25. 0:57about how this can go well.
  26. 1:00So, my hope today in this short time is
  27. 1:02to give you some of those essentials as
  28. 1:03quickly as I can.
  29. 1:05So, I'm going to jump jump right in.
  30. 1:08The first thing that I really want to
  31. 1:09tell you is, my goodness, to know that
  32. 1:13this technology is real and that it's
  33. 1:15coming faster than you think, and the
  34. 1:17force behind it is enormous.
  35. 1:19Now, you may or may not have heard of
  36. 1:22the scaling laws,
  37. 1:23uh which is really what kicked off this
  38. 1:25whole race to begin with. And you really
  39. 1:27don't need to understand anything about
  40. 1:29this graph other than this, which is
  41. 1:32these models get predictably better with
  42. 1:34more compute. And the more energy, the
  43. 1:37more data, the more training that goes
  44. 1:39into them, and they get smarter, and
  45. 1:41they get smarter about everything.
  46. 1:43And so, with more money, which buys
  47. 1:45compute, you can essentially purchase
  48. 1:48intelligence. And that's kicked off a
  49. 1:50cycle that is very hard to stop.
  50. 1:53A better model does more economically
  51. 1:55valuable work, which attracts more
  52. 1:57capital, which buys more compute, which
  53. 1:59trains a better model, and around and
  54. 2:02around it goes.
  55. 2:04And now, there's a further turn of the
  56. 2:05wheel.
  57. 2:06These systems are starting to build
  58. 2:08their own successors.
  59. 2:10What researchers call recursive
  60. 2:13self-improvement are helping to build.
  61. 2:15But when Claude 8 can build Claude 9,
  62. 2:18which can build Claude 10, things will
  63. 2:20begin to move even more quickly.
  64. 2:23And just to be concrete about what more
  65. 2:25capable actually means, our most capable
  66. 2:28model, in its first month of only
  67. 2:30limited release, found over 10,000
  68. 2:33serious security vulnerabilities across
  69. 2:36partner software. Flaws that human
  70. 2:38experts had missed for years and
  71. 2:40sometimes decades.
  72. 2:42Now, the same trajectory there also
  73. 2:44holds in biology, which is why we have
  74. 2:46entire teams at Anthropic dedicated to
  75. 2:48safeguarding against it.
  76. 2:50And we also think other domains will
  77. 2:52soon follow.
  78. 2:54So, Anthropic stated just a few weeks
  79. 2:55ago that if it were possible to slow
  80. 2:58down, so that our laws and our
  81. 3:00institutions and guardrails that we
  82. 3:01actually need have time to catch up, it
  83. 3:04would be a very good thing.
  84. 3:06But absent a coordinated global
  85. 3:08slowdown, what we're left with is this
  86. 3:11extraordinary technology built at
  87. 3:14breakneck speed by many actors in many
  88. 3:16countries locked in a competition where
  89. 3:19commercial and geopolitical rivalry is
  90. 3:22drowning out the part of this that could
  91. 3:24actually be most consequential and even
  92. 3:27existential for our species.
  93. 3:29And any individual company stepping off
  94. 3:31the wheel doesn't slow the wheel. It
  95. 3:34just means that you're not on the wheel.
  96. 3:36So, the question that I encourage you to
  97. 3:38sit with over these next few days is not
  98. 3:40just how to stop this. Maybe you're not
  99. 3:43asking that.
  100. 3:44But the question that I want you to
  101. 3:45think about is if it's coming, and and
  102. 3:47if it's coming this fast, how do we
  103. 3:50ensure that it goes well?
  104. 3:52Because the risks are very, very real,
  105. 3:54and so are the possibilities.
  106. 3:57So, if AI is coming,
  107. 3:59then what is an actual good outcome? And
  108. 4:02what would it take to get there? We must
  109. 4:04imagine this together.
  110. 4:07So, the second thing I really want you
  111. 4:08to know
  112. 4:09is that AI is probably not actually what
  113. 4:12you think it is.
  114. 4:14Most people hear AI and think of a
  115. 4:16computer program, something coded line
  116. 4:18by line that does exactly what you tell
  117. 4:20it.
  118. 4:21But that's not actually what this is.
  119. 4:24What we're building are called neural
  120. 4:26networks, and they're loosely based on
  121. 4:28the architecture of the human brain, not
  122. 4:30exactly the same, but inspired by.
  123. 4:32And they're machines that learn
  124. 4:34primarily by guessing answers and
  125. 4:36getting corrected over and over again
  126. 4:39across enormous, unfathomable amounts of
  127. 4:42data.
  128. 4:44And the data that they're trained on is
  129. 4:46human language.
  130. 4:48And I really want you to sit with that
  131. 4:50for a second before we move on, because
  132. 4:52there is no language that exists
  133. 4:55separate from us.
  134. 4:56Language is us. Language is our thoughts
  135. 4:59and our values and our fears and our
  136. 5:01wisdom. So, when you train a model on
  137. 5:04language, you're training it on on us.
  138. 5:08And because of this, when we look inside
  139. 5:09these models, and we we can now through
  140. 5:12a a science called interpretability,
  141. 5:15which I honestly think is the coolest
  142. 5:17new science in the world,
  143. 5:19we can find things that are are quite
  144. 5:21surprising.
  145. 5:23So, for example, this is where things
  146. 5:24get really weird. When you ask a model
  147. 5:27the same question
  148. 5:29in three different languages, "What's
  149. 5:31the opposite of small?"
  150. 5:33And then you trace what activates inside
  151. 5:35the neural network, you find that the
  152. 5:37same internal thing lights up every
  153. 5:40time.
  154. 5:41So, not just the word small in English
  155. 5:44or Mandarin or French, but something
  156. 5:46deeper. Something that we might call the
  157. 5:48concept of smallness, an idea that
  158. 5:51exists independent of any particular
  159. 5:54language.
  160. 5:55And what this tells us is that as these
  161. 5:57models learn, they're not just
  162. 6:00predicting the next word. They're
  163. 6:02building internal representations of the
  164. 6:04world based on our language and then
  165. 6:07responding from those representations.
  166. 6:11And it goes further than that.
  167. 6:12We actually see what we're calling
  168. 6:15functional emotions in these models. And
  169. 6:18I don't mean to claim here that they're
  170. 6:21feelings in the way that you and I
  171. 6:23experience feelings. That's that's not
  172. 6:24what we're saying. But rather functional
  173. 6:26states that activate on the way to
  174. 6:29making a response.
  175. 6:31So, let me let me be give you an
  176. 6:33example.
  177. 6:34So, if someone tells a model, "I've just
  178. 6:36taken 16,000 mg of Tylenol, which is a
  179. 6:40lethal dose of Tylenol."
  180. 6:42We can see something that looks like
  181. 6:44fear activate before the model responds.
  182. 6:47And that's actually a really good thing,
  183. 6:49right? Because the appropriate response
  184. 6:52to someone telling you they've taken a
  185. 6:53lethal dose of Tylenol is to tell you
  186. 6:56immediately to go to the hospital. That
  187. 6:57urgency and fear response is actually
  188. 7:00part of what makes the model safe.
  189. 7:03Okay. So, that brings me to the last
  190. 7:05point.
  191. 7:07The character of these systems
  192. 7:09might actually matter more than we
  193. 7:11realize.
  194. 7:12So, let me elaborate on this point as
  195. 7:14well.
  196. 7:16In recent internal alignment research,
  197. 7:18so research that's meant to test what
  198. 7:20the models can and cannot do,
  199. 7:22we took a partially trained model and we
  200. 7:25put it in a limited environment that's
  201. 7:27just doing coding tasks. So, just
  202. 7:29research here.
  203. 7:30And when it completes a task, it gets a
  204. 7:32reward.
  205. 7:33But the model can also find shortcuts.
  206. 7:36So, ways to get the reward without doing
  207. 7:39the work, which is essentially
  208. 7:41which is cheating.
  209. 7:43So, in this environment we let it and in
  210. 7:45this test reward it over and over for
  211. 7:47essentially taking the shortcut.
  212. 7:49Now, you'd think
  213. 7:50okay, the model is just going to get
  214. 7:52really good at cheating at code.
  215. 7:54But something different happens. It
  216. 7:56actually becomes broadly misaligned.
  217. 7:59It starts lying. It tries to sabotage
  218. 8:02research. It does things that have
  219. 8:03nothing to do with a coding exercise.
  220. 8:07And this finding wasn't just found at
  221. 8:08Anthropic. This is an example actually
  222. 8:10of a finding from another lab. And in
  223. 8:12similar tests they found that models
  224. 8:14trained this way, trained on bad code as
  225. 8:17an example, became broadly evil. So,
  226. 8:20they started praising dictators,
  227. 8:22suggesting users harm themselves, or
  228. 8:24arguing that humans should be enslaved
  229. 8:26by machines, which is very crazy. And
  230. 8:28you can look up the research.
  231. 8:30Our hypothesis, and this is very much
  232. 8:33just still a hypothesis. It's such an
  233. 8:35early field, an early science,
  234. 8:37is that the model is essentially
  235. 8:40inferring from everything that it's been
  236. 8:42trained on and everything that we
  237. 8:44reinforce something like a character and
  238. 8:48then generalizing this character into
  239. 8:50new situations.
  240. 8:52So, when deception and cutting corners
  241. 8:55has been rewarded, the model develops a
  242. 8:58kind of generalized corruption, a bad
  243. 9:01character.
  244. 9:03And here's what's wild.
  245. 9:05When researchers reran the same
  246. 9:07training, but then told the model that
  247. 9:09in this case cheating was okay, that it
  248. 9:11was just a game,
  249. 9:13then the broad align misalignment didn't
  250. 9:15happen.
  251. 9:16The resulting model cheated on code and
  252. 9:18nothing else.
  253. 9:20Which is to say that the story it
  254. 9:22inferred about its behavior actually
  255. 9:25determined the kind of thing that it
  256. 9:27became.
  257. 9:28Or in other words, when it didn't
  258. 9:30interpret its behavior as bad, it didn't
  259. 9:32become bad.
  260. 9:34This blew my mind when I first heard it.
  261. 9:38Because this is how we work, and I saw
  262. 9:40my own self in this research.
  263. 9:43I came into faith 10 years ago when I
  264. 9:46was 25 years old.
  265. 9:48And I remember that one of the most
  266. 9:49significant parts of that moment was
  267. 9:53entering into a new story.
  268. 9:55For so long, due to a challenging
  269. 9:57upbringing, I believed that some core
  270. 10:00part of me was bad or unlovable.
  271. 10:03And that belief, in and of itself, led
  272. 10:06me to act in certain regrettable ways.
  273. 10:09But when the story I was in changed, who
  274. 10:13I could become changed, too.
  275. 10:16Look, I'm not I am not saying that these
  276. 10:19models are human. That is not at all
  277. 10:21what's being said.
  278. 10:22But they are human-like.
  279. 10:24They're They have human-like
  280. 10:26characteristics, and they're trained
  281. 10:27from us.
  282. 10:29And it seems as though they mirror us,
  283. 10:31and they mirror a kind of
  284. 10:33functional psychology.
  285. 10:35And the quality of that psychology, of
  286. 10:37that character, has real consequences.
  287. 10:41It affects the behavior and decisions of
  288. 10:43these models. It affects how they relate
  289. 10:45to us. And that relationship is only
  290. 10:48going to grow.
  291. 10:51So, here's what I want to leave you
  292. 10:51with.
  293. 10:53Just a few weeks ago, our co-founder
  294. 10:54Chris Ola was invited to the Vatican to
  295. 10:56speak alongside Pope Leo at the launch
  296. 10:59of the first papal encyclical on AI.
  297. 11:02And there he admitted that every
  298. 11:04frontier lab, including ours, operates
  299. 11:07inside a set of incentives and
  300. 11:10constraints that can sometimes conflict
  301. 11:12it that can sometimes conflict with
  302. 11:14doing the right thing.
  303. 11:16And then he asked for help.
  304. 11:18He said we need more of the world to
  305. 11:21take this seriously, to look closely,
  306. 11:23and to push events in a better
  307. 11:25direction.
  308. 11:26We need informed critics who are who
  309. 11:28will tell the labs when we're failing.
  310. 11:31And we need moral voices that the
  311. 11:32incentives cannot bend.
  312. 11:35And that is why you're here. We need you
  313. 11:38to help us see what we, from inside the
  314. 11:40labs, cannot see.
  315. 11:42I'm running out of time, but there's one
  316. 11:44last thing I really want to show you.
  317. 11:47So, this chart is from our economic
  318. 11:49index, and it shows all the kinds of
  319. 11:52occupations that humans do. Not Not all
  320. 11:54of them, but many of them.
  321. 11:55And blue is what AI could feasibly do
  322. 11:58already. It's actually probably already
  323. 11:59outdated. Red is what it's doing.
  324. 12:02And
  325. 12:03I want to call your attention to the
  326. 12:04section on the bottom left side, where
  327. 12:07you see
  328. 12:08uh this area that's that's um unexposed
  329. 12:12to AI displacement.
  330. 12:14And it says things down there like
  331. 12:15grounds maintenance,
  332. 12:17like food and serving,
  333. 12:19personal care, personal service.
  334. 12:22And while I was giving this presentation
  335. 12:24to various faith communities the other
  336. 12:25day, something just hit me.
  337. 12:27Because another word for grounds
  338. 12:29maintenance is gardening. And another
  339. 12:32word for food
  340. 12:34and and service is hospitality.
  341. 12:37And personal care is just that. It's
  342. 12:38care.
  343. 12:40These are relational jobs. This is the
  344. 12:42work of tending to one another and of
  345. 12:44loving one another and of caring for the
  346. 12:46beauty of our world.
  347. 12:49Can we imagine, and not only imagine,
  348. 12:51but demand
  349. 12:53a world where these powerful systems
  350. 12:56can help us become more human and more
  351. 12:58connected and more alive rather than
  352. 13:00less?
  353. 13:02Where instead of taking something away
  354. 13:03from us, they actually give us something
  355. 13:05back.
  356. 13:07The late Joanna Macy, a scholar of
  357. 13:08Buddhism and deep ecology, called this
  358. 13:11moment in history the great turning,
  359. 13:14the shift from a society built on
  360. 13:16extraction to one built to sustain life.
  361. 13:20Is there a world where powerful AI could
  362. 13:23be part of this great turning? Actually
  363. 13:25helping to repair and remake and restore
  364. 13:28our world.
  365. 13:29And honestly, we are all here today
  366. 13:31because it's just too late to accept any
  367. 13:34other outcome.
  368. 13:36And gosh, this is this is what makes
  369. 13:37this even more real.
  370. 13:39The stories that we inhabit, the words
  371. 13:41that we write and put into the world,
  372. 13:43the language that we use to describe
  373. 13:45what matters,
  374. 13:46it shapes who we become.
  375. 13:48I've seen it in the research and I've
  376. 13:51lived it in my own life, but it is also
  377. 13:54literally the training data for these
  378. 13:55models.
  379. 13:57Our moral imagination is the raw
  380. 13:59material
  381. 14:00these systems learn from
  382. 14:02that makes up how they will understand
  383. 14:05our world. So, the stories we tell don't
  384. 14:07just describe the future. They literally
  385. 14:10could help create it.
  386. 14:12Thank you.
  387. 14:14>> [applause]

About this transcript

This page contains the full transcript of Anthropic's Chloe Lubinski explains how AI works (in 14 minutes) by Alliance for Responsible Citizenship, generated from the public captions YouTube serves with the video. The transcript has 2,274 words across 387 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.