Anthropic's Chloe Lubinski explains how AI works (in 14 minutes) — Transcript
Full transcript
- 0:00I work at Anthropic, which uh where I
- 0:03lead the the research partnerships with
- 0:06the world's wisdom traditions.
- 0:08And my job really has two parts to it.
- 0:11So, the first, it's to help the these
- 0:13experts in these various fields and
- 0:15disciplines actually understand AI, what
- 0:18it is, what's happening now, and where
- 0:20it's going.
- 0:22And the second part is to listen and to
- 0:23learn and to funnel wisdom back into the
- 0:26organization, back to the people that
- 0:28are building this technology.
- 0:30Just last week, I was walking my little
- 0:33red-haired cocker spaniel in San
- 0:34Francisco, thinking of what might be
- 0:36most helpful to you today in having
- 0:39these conversations. And the thing is,
- 0:41I've probably had hundreds of
- 0:43conversations now across 20 or so
- 0:46traditions and disciplines, and I found,
- 0:49again and again, just how important it
- 0:52is for folks to really understand the
- 0:54basics before we can even start to talk
- 0:57about how this can go well.
- 1:00So, my hope today in this short time is
- 1:02to give you some of those essentials as
- 1:03quickly as I can.
- 1:05So, I'm going to jump jump right in.
- 1:08The first thing that I really want to
- 1:09tell you is, my goodness, to know that
- 1:13this technology is real and that it's
- 1:15coming faster than you think, and the
- 1:17force behind it is enormous.
- 1:19Now, you may or may not have heard of
- 1:22the scaling laws,
- 1:23uh which is really what kicked off this
- 1:25whole race to begin with. And you really
- 1:27don't need to understand anything about
- 1:29this graph other than this, which is
- 1:32these models get predictably better with
- 1:34more compute. And the more energy, the
- 1:37more data, the more training that goes
- 1:39into them, and they get smarter, and
- 1:41they get smarter about everything.
- 1:43And so, with more money, which buys
- 1:45compute, you can essentially purchase
- 1:48intelligence. And that's kicked off a
- 1:50cycle that is very hard to stop.
- 1:53A better model does more economically
- 1:55valuable work, which attracts more
- 1:57capital, which buys more compute, which
- 1:59trains a better model, and around and
- 2:02around it goes.
- 2:04And now, there's a further turn of the
- 2:05wheel.
- 2:06These systems are starting to build
- 2:08their own successors.
- 2:10What researchers call recursive
- 2:13self-improvement are helping to build.
- 2:15But when Claude 8 can build Claude 9,
- 2:18which can build Claude 10, things will
- 2:20begin to move even more quickly.
- 2:23And just to be concrete about what more
- 2:25capable actually means, our most capable
- 2:28model, in its first month of only
- 2:30limited release, found over 10,000
- 2:33serious security vulnerabilities across
- 2:36partner software. Flaws that human
- 2:38experts had missed for years and
- 2:40sometimes decades.
- 2:42Now, the same trajectory there also
- 2:44holds in biology, which is why we have
- 2:46entire teams at Anthropic dedicated to
- 2:48safeguarding against it.
- 2:50And we also think other domains will
- 2:52soon follow.
- 2:54So, Anthropic stated just a few weeks
- 2:55ago that if it were possible to slow
- 2:58down, so that our laws and our
- 3:00institutions and guardrails that we
- 3:01actually need have time to catch up, it
- 3:04would be a very good thing.
- 3:06But absent a coordinated global
- 3:08slowdown, what we're left with is this
- 3:11extraordinary technology built at
- 3:14breakneck speed by many actors in many
- 3:16countries locked in a competition where
- 3:19commercial and geopolitical rivalry is
- 3:22drowning out the part of this that could
- 3:24actually be most consequential and even
- 3:27existential for our species.
- 3:29And any individual company stepping off
- 3:31the wheel doesn't slow the wheel. It
- 3:34just means that you're not on the wheel.
- 3:36So, the question that I encourage you to
- 3:38sit with over these next few days is not
- 3:40just how to stop this. Maybe you're not
- 3:43asking that.
- 3:44But the question that I want you to
- 3:45think about is if it's coming, and and
- 3:47if it's coming this fast, how do we
- 3:50ensure that it goes well?
- 3:52Because the risks are very, very real,
- 3:54and so are the possibilities.
- 3:57So, if AI is coming,
- 3:59then what is an actual good outcome? And
- 4:02what would it take to get there? We must
- 4:04imagine this together.
- 4:07So, the second thing I really want you
- 4:08to know
- 4:09is that AI is probably not actually what
- 4:12you think it is.
- 4:14Most people hear AI and think of a
- 4:16computer program, something coded line
- 4:18by line that does exactly what you tell
- 4:20it.
- 4:21But that's not actually what this is.
- 4:24What we're building are called neural
- 4:26networks, and they're loosely based on
- 4:28the architecture of the human brain, not
- 4:30exactly the same, but inspired by.
- 4:32And they're machines that learn
- 4:34primarily by guessing answers and
- 4:36getting corrected over and over again
- 4:39across enormous, unfathomable amounts of
- 4:42data.
- 4:44And the data that they're trained on is
- 4:46human language.
- 4:48And I really want you to sit with that
- 4:50for a second before we move on, because
- 4:52there is no language that exists
- 4:55separate from us.
- 4:56Language is us. Language is our thoughts
- 4:59and our values and our fears and our
- 5:01wisdom. So, when you train a model on
- 5:04language, you're training it on on us.
- 5:08And because of this, when we look inside
- 5:09these models, and we we can now through
- 5:12a a science called interpretability,
- 5:15which I honestly think is the coolest
- 5:17new science in the world,
- 5:19we can find things that are are quite
- 5:21surprising.
- 5:23So, for example, this is where things
- 5:24get really weird. When you ask a model
- 5:27the same question
- 5:29in three different languages, "What's
- 5:31the opposite of small?"
- 5:33And then you trace what activates inside
- 5:35the neural network, you find that the
- 5:37same internal thing lights up every
- 5:40time.
- 5:41So, not just the word small in English
- 5:44or Mandarin or French, but something
- 5:46deeper. Something that we might call the
- 5:48concept of smallness, an idea that
- 5:51exists independent of any particular
- 5:54language.
- 5:55And what this tells us is that as these
- 5:57models learn, they're not just
- 6:00predicting the next word. They're
- 6:02building internal representations of the
- 6:04world based on our language and then
- 6:07responding from those representations.
- 6:11And it goes further than that.
- 6:12We actually see what we're calling
- 6:15functional emotions in these models. And
- 6:18I don't mean to claim here that they're
- 6:21feelings in the way that you and I
- 6:23experience feelings. That's that's not
- 6:24what we're saying. But rather functional
- 6:26states that activate on the way to
- 6:29making a response.
- 6:31So, let me let me be give you an
- 6:33example.
- 6:34So, if someone tells a model, "I've just
- 6:36taken 16,000 mg of Tylenol, which is a
- 6:40lethal dose of Tylenol."
- 6:42We can see something that looks like
- 6:44fear activate before the model responds.
- 6:47And that's actually a really good thing,
- 6:49right? Because the appropriate response
- 6:52to someone telling you they've taken a
- 6:53lethal dose of Tylenol is to tell you
- 6:56immediately to go to the hospital. That
- 6:57urgency and fear response is actually
- 7:00part of what makes the model safe.
- 7:03Okay. So, that brings me to the last
- 7:05point.
- 7:07The character of these systems
- 7:09might actually matter more than we
- 7:11realize.
- 7:12So, let me elaborate on this point as
- 7:14well.
- 7:16In recent internal alignment research,
- 7:18so research that's meant to test what
- 7:20the models can and cannot do,
- 7:22we took a partially trained model and we
- 7:25put it in a limited environment that's
- 7:27just doing coding tasks. So, just
- 7:29research here.
- 7:30And when it completes a task, it gets a
- 7:32reward.
- 7:33But the model can also find shortcuts.
- 7:36So, ways to get the reward without doing
- 7:39the work, which is essentially
- 7:41which is cheating.
- 7:43So, in this environment we let it and in
- 7:45this test reward it over and over for
- 7:47essentially taking the shortcut.
- 7:49Now, you'd think
- 7:50okay, the model is just going to get
- 7:52really good at cheating at code.
- 7:54But something different happens. It
- 7:56actually becomes broadly misaligned.
- 7:59It starts lying. It tries to sabotage
- 8:02research. It does things that have
- 8:03nothing to do with a coding exercise.
- 8:07And this finding wasn't just found at
- 8:08Anthropic. This is an example actually
- 8:10of a finding from another lab. And in
- 8:12similar tests they found that models
- 8:14trained this way, trained on bad code as
- 8:17an example, became broadly evil. So,
- 8:20they started praising dictators,
- 8:22suggesting users harm themselves, or
- 8:24arguing that humans should be enslaved
- 8:26by machines, which is very crazy. And
- 8:28you can look up the research.
- 8:30Our hypothesis, and this is very much
- 8:33just still a hypothesis. It's such an
- 8:35early field, an early science,
- 8:37is that the model is essentially
- 8:40inferring from everything that it's been
- 8:42trained on and everything that we
- 8:44reinforce something like a character and
- 8:48then generalizing this character into
- 8:50new situations.
- 8:52So, when deception and cutting corners
- 8:55has been rewarded, the model develops a
- 8:58kind of generalized corruption, a bad
- 9:01character.
- 9:03And here's what's wild.
- 9:05When researchers reran the same
- 9:07training, but then told the model that
- 9:09in this case cheating was okay, that it
- 9:11was just a game,
- 9:13then the broad align misalignment didn't
- 9:15happen.
- 9:16The resulting model cheated on code and
- 9:18nothing else.
- 9:20Which is to say that the story it
- 9:22inferred about its behavior actually
- 9:25determined the kind of thing that it
- 9:27became.
- 9:28Or in other words, when it didn't
- 9:30interpret its behavior as bad, it didn't
- 9:32become bad.
- 9:34This blew my mind when I first heard it.
- 9:38Because this is how we work, and I saw
- 9:40my own self in this research.
- 9:43I came into faith 10 years ago when I
- 9:46was 25 years old.
- 9:48And I remember that one of the most
- 9:49significant parts of that moment was
- 9:53entering into a new story.
- 9:55For so long, due to a challenging
- 9:57upbringing, I believed that some core
- 10:00part of me was bad or unlovable.
- 10:03And that belief, in and of itself, led
- 10:06me to act in certain regrettable ways.
- 10:09But when the story I was in changed, who
- 10:13I could become changed, too.
- 10:16Look, I'm not I am not saying that these
- 10:19models are human. That is not at all
- 10:21what's being said.
- 10:22But they are human-like.
- 10:24They're They have human-like
- 10:26characteristics, and they're trained
- 10:27from us.
- 10:29And it seems as though they mirror us,
- 10:31and they mirror a kind of
- 10:33functional psychology.
- 10:35And the quality of that psychology, of
- 10:37that character, has real consequences.
- 10:41It affects the behavior and decisions of
- 10:43these models. It affects how they relate
- 10:45to us. And that relationship is only
- 10:48going to grow.
- 10:51So, here's what I want to leave you
- 10:51with.
- 10:53Just a few weeks ago, our co-founder
- 10:54Chris Ola was invited to the Vatican to
- 10:56speak alongside Pope Leo at the launch
- 10:59of the first papal encyclical on AI.
- 11:02And there he admitted that every
- 11:04frontier lab, including ours, operates
- 11:07inside a set of incentives and
- 11:10constraints that can sometimes conflict
- 11:12it that can sometimes conflict with
- 11:14doing the right thing.
- 11:16And then he asked for help.
- 11:18He said we need more of the world to
- 11:21take this seriously, to look closely,
- 11:23and to push events in a better
- 11:25direction.
- 11:26We need informed critics who are who
- 11:28will tell the labs when we're failing.
- 11:31And we need moral voices that the
- 11:32incentives cannot bend.
- 11:35And that is why you're here. We need you
- 11:38to help us see what we, from inside the
- 11:40labs, cannot see.
- 11:42I'm running out of time, but there's one
- 11:44last thing I really want to show you.
- 11:47So, this chart is from our economic
- 11:49index, and it shows all the kinds of
- 11:52occupations that humans do. Not Not all
- 11:54of them, but many of them.
- 11:55And blue is what AI could feasibly do
- 11:58already. It's actually probably already
- 11:59outdated. Red is what it's doing.
- 12:02And
- 12:03I want to call your attention to the
- 12:04section on the bottom left side, where
- 12:07you see
- 12:08uh this area that's that's um unexposed
- 12:12to AI displacement.
- 12:14And it says things down there like
- 12:15grounds maintenance,
- 12:17like food and serving,
- 12:19personal care, personal service.
- 12:22And while I was giving this presentation
- 12:24to various faith communities the other
- 12:25day, something just hit me.
- 12:27Because another word for grounds
- 12:29maintenance is gardening. And another
- 12:32word for food
- 12:34and and service is hospitality.
- 12:37And personal care is just that. It's
- 12:38care.
- 12:40These are relational jobs. This is the
- 12:42work of tending to one another and of
- 12:44loving one another and of caring for the
- 12:46beauty of our world.
- 12:49Can we imagine, and not only imagine,
- 12:51but demand
- 12:53a world where these powerful systems
- 12:56can help us become more human and more
- 12:58connected and more alive rather than
- 13:00less?
- 13:02Where instead of taking something away
- 13:03from us, they actually give us something
- 13:05back.
- 13:07The late Joanna Macy, a scholar of
- 13:08Buddhism and deep ecology, called this
- 13:11moment in history the great turning,
- 13:14the shift from a society built on
- 13:16extraction to one built to sustain life.
- 13:20Is there a world where powerful AI could
- 13:23be part of this great turning? Actually
- 13:25helping to repair and remake and restore
- 13:28our world.
- 13:29And honestly, we are all here today
- 13:31because it's just too late to accept any
- 13:34other outcome.
- 13:36And gosh, this is this is what makes
- 13:37this even more real.
- 13:39The stories that we inhabit, the words
- 13:41that we write and put into the world,
- 13:43the language that we use to describe
- 13:45what matters,
- 13:46it shapes who we become.
- 13:48I've seen it in the research and I've
- 13:51lived it in my own life, but it is also
- 13:54literally the training data for these
- 13:55models.
- 13:57Our moral imagination is the raw
- 13:59material
- 14:00these systems learn from
- 14:02that makes up how they will understand
- 14:05our world. So, the stories we tell don't
- 14:07just describe the future. They literally
- 14:10could help create it.
- 14:12Thank you.
- 14:14>> [applause]
About this transcript
This page contains the full transcript of Anthropic's Chloe Lubinski explains how AI works (in 14 minutes) by Alliance for Responsible Citizenship, generated from the public captions YouTube serves with the video. The transcript has 2,274 words across 387 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.