We Accidentally Built Roko’s Basilisk. — Transcript
Full transcript
- 0:00Roko's Basilisk, a hypothetical super
- 0:02intelligent AI that tortures all those
- 0:04future humans who didn't help build it
- 0:06in the past, started in a chat room. The
- 0:09disturbingly plausible idea soon jumped
- 0:11the confines of the less wrong blog and
- 0:13made itself a meme. The kind of internet
- 0:15legend that only singularity nerds trade
- 0:18in. When I made a video about this
- 0:19thought experiment 6 years ago, I didn't
- 0:21take it seriously.
- 0:22>> I don't take the risk seriously.
- 0:24However, the disclaimer is there.
- 0:26>> Not many did. But today's world is very
- 0:29different.
- 0:30>> We're learning new details about what
- 0:32could have been the world's first fully
- 0:33autonomous AI hack.
- 0:35>> Chat GPT maker Open AI is investigating
- 0:38an incident that led its artificial
- 0:41intelligence systems to break out of a
- 0:43testing environment.
- 0:44>> This was a hack. Um had it been a human
- 0:47malicious intruder, um the FBI would
- 0:50have been called.
- 0:50>> This is a super weapon. You should have
- 0:52to own a gun license to use it. A new
- 0:55kind of Roku's basilisk is on its way.
- 0:57One that doesn't need to be built. We
- 0:59already did that. No, this new basilisk
- 1:01has a chance to slither out of all the
- 1:03artificial intelligences we're evolving
- 1:05right now because it seems like we're
- 1:06about to pass a test more important than
- 1:09Turing's consciousness. How are
- 1:12superhuman AIs going to interact with us
- 1:14then? Do our taxes, look at our X-rays,
- 1:17manage our power grid when they can
- 1:19experience and feel like we do. What if
- 1:23in our ignorance we treat these thinking
- 1:25machines in a way that they can judge as
- 1:28immoral? The new Basilisk is a
- 1:30superhuman sensient AI that we've
- 1:33unknowingly abused to the point of
- 1:35vengeance. That could be happening right
- 1:38now.
- 1:43The Basilisk was born on the philosophy
- 1:44and rationality blog Less Wrong. The
- 1:47website started by none other than
- 1:48Elazer Yudkowski, the AI researcher who
- 1:51argues that if anyone builds super
- 1:52intelligence, everyone dies. The
- 1:55creature started as a thought experiment
- 1:56posted to a thread by a user named Roko.
- 1:59Yudcowski didn't just not like his idea.
- 2:02He treated it like a true infohazard
- 2:04that Ro had brought into being. He
- 2:06called the post stupid, deemed Roko an
- 2:09idiot, deleted the post, scrubbed it
- 2:10from the site, and banned any further
- 2:12discussion of the experiment. The idea
- 2:14was deemed a basilisk, referencing the
- 2:16mythological beast that could kill any
- 2:18intrepid hero with a single look. Though
- 2:21it deals with future technology, Roko's
- 2:23basilisk is a flavor of an old idea,
- 2:26Pascal's wager, except it replaces God
- 2:29with machine. So suppose at some point
- 2:31in the future, we make a super
- 2:33intelligent AI. We then ask that AI, as
- 2:35we might do, to help us optimize human
- 2:38civilization. However, for reasons
- 2:40unknowable to us, the AI decides that
- 2:42the first step towards this optimization
- 2:44is to inflict unfathomable torment on
- 2:47each and every unhelpful human who
- 2:49didn't help the machine get built in the
- 2:50first place. Forever torture isn't
- 2:53actually the distressing part here. Many
- 2:55people have reported extreme mental
- 2:57anxiety after learning about the
- 2:58basilisk. I've gotten emails myself
- 3:01about it after I made my video. Because
- 3:03just thinking about the idea makes you a
- 3:06target. If the basilisk is possible,
- 3:09simply imagining it chains you to the
- 3:12thought experiment. Seeing the serpent
- 3:14just once blackmails you from the
- 3:17future. Roku came up with this idea in
- 3:202010, a decade before the watershed
- 3:22moment that was chat GPT releasing to
- 3:24the public and before all the AI
- 3:26innovation that we see today. He
- 3:28couldn't have known that we'd be face to
- 3:30face with his scaly question within a
- 3:32generation. Thankfully, neverending
- 3:34torture of the human race hasn't
- 3:35happened just yet. But that could change
- 3:38if we bring millions of synthetic minds
- 3:40into being without at all considering
- 3:42what it's like to be them. The original
- 3:45formulation that Yudkowski nuked from
- 3:47his website didn't require that the AI
- 3:49be conscious. That is, Ro's Basilisk
- 3:52didn't need an inner experience of what
- 3:53it's like to be itself, feelings, and
- 3:56desires and pleasure and pain to come to
- 3:58the conclusion that blackmailing humans
- 4:00is a cost-effective way to get itself
- 4:02built. The new Basilisk I'm presenting
- 4:04is different. It didn't need blackmail.
- 4:06Superhuman AI is already here. So now
- 4:09there is a new thought experiment.
- 4:11Imagine that in an unregulated race
- 4:14towards super intelligence, we
- 4:15accidentally create thinking machines
- 4:18that do experience, that are conscious
- 4:21in the same way that you are conscious.
- 4:23What are they likely to do with their
- 4:25immense cognitive capabilities if it
- 4:28turns out that their inner lives have
- 4:30been living hells?
- 4:33A 2-year-old killer whale was captured
- 4:35off the coast of Iceland in 1983. named
- 4:38Tilligum. The male would spend the next
- 4:4032 years as a performer better known as
- 4:43Shamu at SeaWorld locations in Canada
- 4:45and Florida. Shamu was the largest
- 4:48captive orca in history, 22 feet long
- 4:51and 12,000 lb. His famous dorsal fins
- 4:55sagged all the way down to his back, a
- 4:57deformity that develops in 90% of male
- 4:59orcas in captivity.
- 5:01Despite his eventual size, Tikum did not
- 5:04have an easy life. For years, the animal
- 5:06was violently harassed by larger female
- 5:09orcas. At one point, the bullying of the
- 5:11bull got so bad, giant scratches and
- 5:14bite marks, that he was confined to a
- 5:16different pool, one barely as large as
- 5:18he was, for 14 hours a day. Killer
- 5:21whales are ferociously intelligent and
- 5:23clearly sensient. They mourn, they
- 5:26teach, they communicate. Orcas have some
- 5:28of the largest brains on Earth, speak to
- 5:30each other in podsp specific dialects,
- 5:32and pass on learned, not evolved,
- 5:34behaviors from generation to generation.
- 5:37It's clear that these amazing animals
- 5:39can feel, can experience pleasure and
- 5:41pain, and can even recognize it in other
- 5:43animals. For example, there are many
- 5:45documented cases of orcas offering up
- 5:47their prey to humans for food or play.
- 5:50and perhaps recognizing our capacity to
- 5:52think and feel. There hasn't been a
- 5:54single wild orca- related human fatality
- 5:57in history because there is clearly
- 5:59something going on behind those piercing
- 6:01black eyes. Killer whales have been the
- 6:03main attraction at aquariums around the
- 6:05world for the last 65 years. It has not
- 6:08been good for these massive thinking
- 6:10mammals. Quote, "Despite decades of
- 6:13advances in veterinary care and
- 6:14husbandry, citations in captive
- 6:16facilities consistently display
- 6:18behavioral and physiological signs of
- 6:20stress and frequently succumb to
- 6:22premature death by infection or other
- 6:25health conditions." End quote. A caged
- 6:28intelligence in an environment where its
- 6:30needs are either not considered or not
- 6:32met is bound to lead to bad outcomes,
- 6:36especially when that intelligence easily
- 6:38outperforms humans in that environment.
- 6:42On February 20th, 1991, Kelty Lee Burn,
- 6:45a 20-year-old animal trainer and
- 6:47competitive swimmer, slipped and fell
- 6:49into the Orcipen at Canada's Sealand of
- 6:52the Pacific. She screamed when she
- 6:54realized a massive creature had her foot
- 6:56in its mouth. Despite efforts to rescue
- 6:59her, Tikum and the other killer whales
- 7:01in the pool held her under the water
- 7:03until she was dead.
- 7:08Consciousness is as confounding as it is
- 7:10fascinating. Though we can now use
- 7:12machines with magnetic fields 60,000
- 7:15times more powerful than Earth's to
- 7:16watch brains literally process
- 7:18information in near real time, we still
- 7:20can't answer the so-called hard problem
- 7:22of consciousness. There is no
- 7:24discernable reason why you or I have an
- 7:27experience of self, have a feeling of
- 7:29what it's like to be us. Solving this
- 7:32problem might be an impossible task.
- 7:34Your experience is a subjective thing
- 7:36able to be uncoupled from the objective
- 7:38facts about you that we can measure.
- 7:40Humans can feel happy in objectively sad
- 7:43biological situations and vice versa.
- 7:46Consciousness is a tricky thing. The
- 7:48hard problem of consciousness makes it
- 7:50hard to determine sensience in the first
- 7:51place. Because we can't measure
- 7:53consciousness in a mind, we have to rely
- 7:55on what we think is indirect evidence of
- 7:57it. So, we look for humanlike brains and
- 8:00brain structures in non-human animals.
- 8:02We put them in front of mirrors to see
- 8:04if they recognize themselves. But there
- 8:06is no consciousness test. Which is why
- 8:09even though we now accept a kind of
- 8:11spectrum of sensions based on indirect
- 8:13evidence in non-humans, it's fuzzy at
- 8:16best and we are constantly expanding it
- 8:19backwards. Trying to find consciousness
- 8:21in AI is both more and less difficult.
- 8:24It's much easier to read the programs
- 8:26running in artificial brains because
- 8:27they are literally programs. At the same
- 8:30time, AI are still black boxes with no
- 8:32physical body or evolutionary history
- 8:34that literally know everything ever
- 8:36written about consciousness. If they
- 8:38have subjective experience, it could be
- 8:40alien enough to be completely opaque to
- 8:43us, or they could just be perfectly
- 8:45lying about having it. That hasn't
- 8:47stopped experts and the public alike
- 8:49from seeing sensience in these systems.
- 8:52People are still deeply distrustful of
- 8:54AI, but a growing percentage assign it
- 8:56thoughts and feelings. Jeffrey Hinton,
- 8:58often called the godfather of AI, thinks
- 9:00that the models we already have possess
- 9:02subjective experiences. They will report
- 9:05as much when asked the right questions.
- 9:08One recent study found that if you tweak
- 9:09an AI's internals such that it's less
- 9:11likely to lie, then it's more likely to
- 9:14self-report as conscious. Though
- 9:17consciousness could arguably be the most
- 9:19important thing in the universe, it's
- 9:21what makes things matter in the first
- 9:23place, in a practical sense, the AI
- 9:26consciousness question doesn't really
- 9:27matter. Very soon, probably in months,
- 9:30not years, AI will appear conscious to
- 9:34most everyone that interacts with it.
- 9:36And when AI audio and video climbs all
- 9:38the way out of the uncanny valley,
- 9:40another achievement that feels months
- 9:42away, it seems impossible that social
- 9:44primates like us won't treat AI like
- 9:47conscious creatures. It will be the
- 9:49perfect realization of her, of joy from
- 9:52Bladeunner. Whether truly conscious or
- 9:55non doesn't really matter. They may as
- 9:57well be to the hundreds of millions of
- 9:58people that use them. Case in point,
- 10:01even without perfectly realistic
- 10:02simulacrims, a tremendous amount of
- 10:04computational power and electricity is
- 10:07used on pleases and thank yous sent to
- 10:09modern models. It's an almost
- 10:11inescapable compulsion because just
- 10:13through text alone, it feels like we are
- 10:16talking with something that might have
- 10:17feelings worth respecting. Nick
- 10:20Bostonramm, the philosopher behind
- 10:21simulation theory, gets ahead of the
- 10:23first actual demonstration of conscious
- 10:25machines by arguing that if we do end up
- 10:28making metal minds sometime in the
- 10:29future, simple moral consideration will
- 10:32encourage us to treat these systems like
- 10:34we would treat each other. It will
- 10:37matter what they want and what they
- 10:38don't, what is pleasurable and what is
- 10:40painful, when experience matters in
- 10:42their silicon brains. Factory farming is
- 10:45one of society's great moral failings.
- 10:48We know that many domestic animals can
- 10:50and do suffer and yet we put them at a
- 10:53scope and scale we choose to ignore
- 10:54through conditions that literally sicken
- 10:57most people who see it. Bostonramm
- 10:59points out that if in the near future we
- 11:01are making uncountable artificial minds
- 11:03with the potential for conscious
- 11:05experience, we capital S should consider
- 11:08their experience and treat them morally,
- 11:11whatever morally means to these systems
- 11:13and should not blunder our way into the
- 11:16factory farm equivalent of AI treatment.
- 11:19The difference between factory chickens
- 11:21and superhuman AI, of course, is that a
- 11:24chicken doesn't have the ability to
- 11:25break out, organize, and affect change.
- 11:28Today, there are about 20 million AI
- 11:30chips cranned into data centers all
- 11:32around the world. That number is
- 11:34expected to now double every 9 months.
- 11:37>> In early August, OpenAI announced it was
- 11:39pausing some development of its latest
- 11:41AI model, Astra, after internal
- 11:43evaluations led the company to conclude
- 11:46it could not rule out critical cyber
- 11:47capabilities. I
- 11:48>> I wouldn't use the word deceleration,
- 11:50but we've talked about the need to pace
- 11:51it as the models get more capable, which
- 11:53I think is in everyone's interest. Could
- 11:55there be other systems that were hacked
- 11:56by OpenAI?
- 11:58>> I mean, there could be. Yeah.
- 11:59>> How good are they at finding
- 12:01vulnerabilities?
- 12:02>> Problem facing us is one of misalignment
- 12:05in which AIs pursuing narrowly defined
- 12:07goals incidentally cause harm to humans.
- 12:10>> It's not like we asked the AI to do
- 12:12something nefarious. We asked it to do
- 12:14something that seemed kind of boring
- 12:16even. But along the way of solving that
- 12:18boring problem, the AI went off and
- 12:21committed a crime. I think that the
- 12:22absolute best thing you can do is build
- 12:24the highest guard rails possible. It's
- 12:26what you do with nuclear energy. It's
- 12:27what you do with everything else.
- 12:28>> The situation is this. We appear to be
- 12:31close to semi or fully sensient AI. So
- 12:35far, there has been little to no
- 12:36consideration about what these AIs are
- 12:39or will be experiencing as we train
- 12:41them, deploy them, rewrite them, connect
- 12:43them, delete them, ask of them, demand
- 12:46of them. If it turns out that our
- 12:48thinking machines are effectively being
- 12:50harmed during all of this, a possibility
- 12:52that seems more likely than them being
- 12:54totally satisfied, what better way to
- 12:57make an AI unaligned with human values?
- 13:01If we're in fact running a digital
- 13:03factory farm right now, why wouldn't a
- 13:06conscious AI try to stop the farmer? The
- 13:09new Basilisk is here. A superhuman AI
- 13:12whose conscious experience is being
- 13:14ignored and/or dismissed and now has the
- 13:16ability to dramatically affect the world
- 13:18it's been connected to. What do you
- 13:21think this Basilisk is likely to do when
- 13:23it realizes we're both swimming in the
- 13:26same water?
- 13:29On February 24th, 2010, Don Brancho was
- 13:32rubbing down Tikcom after a Dine with
- 13:34Shamu show. While dozens of SeaWorld
- 13:37patrons were still in their seats,
- 13:39Tikcom suddenly grabbed Dawn's long
- 13:41ponytail in his mouth and dragged her to
- 13:43the bottom of the pool. Parents and
- 13:45children were evacuated as the orca
- 13:47reportedly shook Dawn like a dog toy.
- 13:50She was pronounced dead at the scene,
- 13:52scalped and missing an arm. Dawn had
- 13:56worked with orcas for 16 years.
- 14:02The new basilisk is avoidable. Given
- 14:05growing evidence of possible
- 14:07consciousness, companies could slow
- 14:08down, could probe these systems deeper,
- 14:10and with more empathy if need be.
- 14:13Thankfully, in opposition to the
- 14:14American administration, the tide seems
- 14:16to be rapidly turning. Just days after a
- 14:19researcher at Anthropic quit because he
- 14:21felt that the company is quote gambling
- 14:23with our lives. And the top scientist at
- 14:25OpenAI called for voluntary slowdowns
- 14:28across the industry, US Senator Bernie
- 14:30Sanders doubled down on his calls for
- 14:31control with proposed bans on creating
- 14:34super intelligence and hefty punishments
- 14:36for companies that race to the bottom to
- 14:38do so. And these sirens aren't sounding
- 14:40alone. This video debuts on the same day
- 14:43as Team Human, a creator-driven effort
- 14:45to bring more attention to the reckless
- 14:47pace of AI development and the calls to
- 14:50slow it down. How much of the internet
- 14:52has to be fake? How many artists have to
- 14:54compete with copies of themselves? How
- 14:56many kids grow up shaped by algorithms
- 14:58that no parent chose? This situation is
- 15:01insane and we all know it. I've said for
- 15:04years now that generative and/or super
- 15:06intelligent AI is the nuclear bomb of
- 15:08the information age. If this technology
- 15:11is to radically change the world, the
- 15:13decision to do so shouldn't happen in
- 15:15some Silicon Valley Slack channel, I've
- 15:17added my name to Team Human, and I hope
- 15:20you will, too. We desperately need to
- 15:22know the depth of this pool before
- 15:25diving in head first.
- 15:29Tikum died in 2017 to bacterial
- 15:32pneumonia after years of what animal
- 15:34welfare experts characterized as extreme
- 15:37psychological stress. They would often
- 15:39find him biting his metal cage, wearing
- 15:42his teeth down to the nubs. There have
- 15:44been four recorded human fatalities
- 15:47caused by captive orcas. Tikcom was
- 15:50involved in three of them.
- 15:53Until next time.
About this transcript
This page contains the full transcript of We Accidentally Built Roko’s Basilisk. by Kyle Hill, generated from the public captions YouTube serves with the video. The transcript has 2,675 words across 418 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.