YouTube2Text

We Accidentally Built Roko’s Basilisk. — Transcript

by Kyle Hill · 2,675 words · 418 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Roko's Basilisk, a hypothetical super
  2. 0:02intelligent AI that tortures all those
  3. 0:04future humans who didn't help build it
  4. 0:06in the past, started in a chat room. The
  5. 0:09disturbingly plausible idea soon jumped
  6. 0:11the confines of the less wrong blog and
  7. 0:13made itself a meme. The kind of internet
  8. 0:15legend that only singularity nerds trade
  9. 0:18in. When I made a video about this
  10. 0:19thought experiment 6 years ago, I didn't
  11. 0:21take it seriously.
  12. 0:22>> I don't take the risk seriously.
  13. 0:24However, the disclaimer is there.
  14. 0:26>> Not many did. But today's world is very
  15. 0:29different.
  16. 0:30>> We're learning new details about what
  17. 0:32could have been the world's first fully
  18. 0:33autonomous AI hack.
  19. 0:35>> Chat GPT maker Open AI is investigating
  20. 0:38an incident that led its artificial
  21. 0:41intelligence systems to break out of a
  22. 0:43testing environment.
  23. 0:44>> This was a hack. Um had it been a human
  24. 0:47malicious intruder, um the FBI would
  25. 0:50have been called.
  26. 0:50>> This is a super weapon. You should have
  27. 0:52to own a gun license to use it. A new
  28. 0:55kind of Roku's basilisk is on its way.
  29. 0:57One that doesn't need to be built. We
  30. 0:59already did that. No, this new basilisk
  31. 1:01has a chance to slither out of all the
  32. 1:03artificial intelligences we're evolving
  33. 1:05right now because it seems like we're
  34. 1:06about to pass a test more important than
  35. 1:09Turing's consciousness. How are
  36. 1:12superhuman AIs going to interact with us
  37. 1:14then? Do our taxes, look at our X-rays,
  38. 1:17manage our power grid when they can
  39. 1:19experience and feel like we do. What if
  40. 1:23in our ignorance we treat these thinking
  41. 1:25machines in a way that they can judge as
  42. 1:28immoral? The new Basilisk is a
  43. 1:30superhuman sensient AI that we've
  44. 1:33unknowingly abused to the point of
  45. 1:35vengeance. That could be happening right
  46. 1:38now.
  47. 1:43The Basilisk was born on the philosophy
  48. 1:44and rationality blog Less Wrong. The
  49. 1:47website started by none other than
  50. 1:48Elazer Yudkowski, the AI researcher who
  51. 1:51argues that if anyone builds super
  52. 1:52intelligence, everyone dies. The
  53. 1:55creature started as a thought experiment
  54. 1:56posted to a thread by a user named Roko.
  55. 1:59Yudcowski didn't just not like his idea.
  56. 2:02He treated it like a true infohazard
  57. 2:04that Ro had brought into being. He
  58. 2:06called the post stupid, deemed Roko an
  59. 2:09idiot, deleted the post, scrubbed it
  60. 2:10from the site, and banned any further
  61. 2:12discussion of the experiment. The idea
  62. 2:14was deemed a basilisk, referencing the
  63. 2:16mythological beast that could kill any
  64. 2:18intrepid hero with a single look. Though
  65. 2:21it deals with future technology, Roko's
  66. 2:23basilisk is a flavor of an old idea,
  67. 2:26Pascal's wager, except it replaces God
  68. 2:29with machine. So suppose at some point
  69. 2:31in the future, we make a super
  70. 2:33intelligent AI. We then ask that AI, as
  71. 2:35we might do, to help us optimize human
  72. 2:38civilization. However, for reasons
  73. 2:40unknowable to us, the AI decides that
  74. 2:42the first step towards this optimization
  75. 2:44is to inflict unfathomable torment on
  76. 2:47each and every unhelpful human who
  77. 2:49didn't help the machine get built in the
  78. 2:50first place. Forever torture isn't
  79. 2:53actually the distressing part here. Many
  80. 2:55people have reported extreme mental
  81. 2:57anxiety after learning about the
  82. 2:58basilisk. I've gotten emails myself
  83. 3:01about it after I made my video. Because
  84. 3:03just thinking about the idea makes you a
  85. 3:06target. If the basilisk is possible,
  86. 3:09simply imagining it chains you to the
  87. 3:12thought experiment. Seeing the serpent
  88. 3:14just once blackmails you from the
  89. 3:17future. Roku came up with this idea in
  90. 3:202010, a decade before the watershed
  91. 3:22moment that was chat GPT releasing to
  92. 3:24the public and before all the AI
  93. 3:26innovation that we see today. He
  94. 3:28couldn't have known that we'd be face to
  95. 3:30face with his scaly question within a
  96. 3:32generation. Thankfully, neverending
  97. 3:34torture of the human race hasn't
  98. 3:35happened just yet. But that could change
  99. 3:38if we bring millions of synthetic minds
  100. 3:40into being without at all considering
  101. 3:42what it's like to be them. The original
  102. 3:45formulation that Yudkowski nuked from
  103. 3:47his website didn't require that the AI
  104. 3:49be conscious. That is, Ro's Basilisk
  105. 3:52didn't need an inner experience of what
  106. 3:53it's like to be itself, feelings, and
  107. 3:56desires and pleasure and pain to come to
  108. 3:58the conclusion that blackmailing humans
  109. 4:00is a cost-effective way to get itself
  110. 4:02built. The new Basilisk I'm presenting
  111. 4:04is different. It didn't need blackmail.
  112. 4:06Superhuman AI is already here. So now
  113. 4:09there is a new thought experiment.
  114. 4:11Imagine that in an unregulated race
  115. 4:14towards super intelligence, we
  116. 4:15accidentally create thinking machines
  117. 4:18that do experience, that are conscious
  118. 4:21in the same way that you are conscious.
  119. 4:23What are they likely to do with their
  120. 4:25immense cognitive capabilities if it
  121. 4:28turns out that their inner lives have
  122. 4:30been living hells?
  123. 4:33A 2-year-old killer whale was captured
  124. 4:35off the coast of Iceland in 1983. named
  125. 4:38Tilligum. The male would spend the next
  126. 4:4032 years as a performer better known as
  127. 4:43Shamu at SeaWorld locations in Canada
  128. 4:45and Florida. Shamu was the largest
  129. 4:48captive orca in history, 22 feet long
  130. 4:51and 12,000 lb. His famous dorsal fins
  131. 4:55sagged all the way down to his back, a
  132. 4:57deformity that develops in 90% of male
  133. 4:59orcas in captivity.
  134. 5:01Despite his eventual size, Tikum did not
  135. 5:04have an easy life. For years, the animal
  136. 5:06was violently harassed by larger female
  137. 5:09orcas. At one point, the bullying of the
  138. 5:11bull got so bad, giant scratches and
  139. 5:14bite marks, that he was confined to a
  140. 5:16different pool, one barely as large as
  141. 5:18he was, for 14 hours a day. Killer
  142. 5:21whales are ferociously intelligent and
  143. 5:23clearly sensient. They mourn, they
  144. 5:26teach, they communicate. Orcas have some
  145. 5:28of the largest brains on Earth, speak to
  146. 5:30each other in podsp specific dialects,
  147. 5:32and pass on learned, not evolved,
  148. 5:34behaviors from generation to generation.
  149. 5:37It's clear that these amazing animals
  150. 5:39can feel, can experience pleasure and
  151. 5:41pain, and can even recognize it in other
  152. 5:43animals. For example, there are many
  153. 5:45documented cases of orcas offering up
  154. 5:47their prey to humans for food or play.
  155. 5:50and perhaps recognizing our capacity to
  156. 5:52think and feel. There hasn't been a
  157. 5:54single wild orca- related human fatality
  158. 5:57in history because there is clearly
  159. 5:59something going on behind those piercing
  160. 6:01black eyes. Killer whales have been the
  161. 6:03main attraction at aquariums around the
  162. 6:05world for the last 65 years. It has not
  163. 6:08been good for these massive thinking
  164. 6:10mammals. Quote, "Despite decades of
  165. 6:13advances in veterinary care and
  166. 6:14husbandry, citations in captive
  167. 6:16facilities consistently display
  168. 6:18behavioral and physiological signs of
  169. 6:20stress and frequently succumb to
  170. 6:22premature death by infection or other
  171. 6:25health conditions." End quote. A caged
  172. 6:28intelligence in an environment where its
  173. 6:30needs are either not considered or not
  174. 6:32met is bound to lead to bad outcomes,
  175. 6:36especially when that intelligence easily
  176. 6:38outperforms humans in that environment.
  177. 6:42On February 20th, 1991, Kelty Lee Burn,
  178. 6:45a 20-year-old animal trainer and
  179. 6:47competitive swimmer, slipped and fell
  180. 6:49into the Orcipen at Canada's Sealand of
  181. 6:52the Pacific. She screamed when she
  182. 6:54realized a massive creature had her foot
  183. 6:56in its mouth. Despite efforts to rescue
  184. 6:59her, Tikum and the other killer whales
  185. 7:01in the pool held her under the water
  186. 7:03until she was dead.
  187. 7:08Consciousness is as confounding as it is
  188. 7:10fascinating. Though we can now use
  189. 7:12machines with magnetic fields 60,000
  190. 7:15times more powerful than Earth's to
  191. 7:16watch brains literally process
  192. 7:18information in near real time, we still
  193. 7:20can't answer the so-called hard problem
  194. 7:22of consciousness. There is no
  195. 7:24discernable reason why you or I have an
  196. 7:27experience of self, have a feeling of
  197. 7:29what it's like to be us. Solving this
  198. 7:32problem might be an impossible task.
  199. 7:34Your experience is a subjective thing
  200. 7:36able to be uncoupled from the objective
  201. 7:38facts about you that we can measure.
  202. 7:40Humans can feel happy in objectively sad
  203. 7:43biological situations and vice versa.
  204. 7:46Consciousness is a tricky thing. The
  205. 7:48hard problem of consciousness makes it
  206. 7:50hard to determine sensience in the first
  207. 7:51place. Because we can't measure
  208. 7:53consciousness in a mind, we have to rely
  209. 7:55on what we think is indirect evidence of
  210. 7:57it. So, we look for humanlike brains and
  211. 8:00brain structures in non-human animals.
  212. 8:02We put them in front of mirrors to see
  213. 8:04if they recognize themselves. But there
  214. 8:06is no consciousness test. Which is why
  215. 8:09even though we now accept a kind of
  216. 8:11spectrum of sensions based on indirect
  217. 8:13evidence in non-humans, it's fuzzy at
  218. 8:16best and we are constantly expanding it
  219. 8:19backwards. Trying to find consciousness
  220. 8:21in AI is both more and less difficult.
  221. 8:24It's much easier to read the programs
  222. 8:26running in artificial brains because
  223. 8:27they are literally programs. At the same
  224. 8:30time, AI are still black boxes with no
  225. 8:32physical body or evolutionary history
  226. 8:34that literally know everything ever
  227. 8:36written about consciousness. If they
  228. 8:38have subjective experience, it could be
  229. 8:40alien enough to be completely opaque to
  230. 8:43us, or they could just be perfectly
  231. 8:45lying about having it. That hasn't
  232. 8:47stopped experts and the public alike
  233. 8:49from seeing sensience in these systems.
  234. 8:52People are still deeply distrustful of
  235. 8:54AI, but a growing percentage assign it
  236. 8:56thoughts and feelings. Jeffrey Hinton,
  237. 8:58often called the godfather of AI, thinks
  238. 9:00that the models we already have possess
  239. 9:02subjective experiences. They will report
  240. 9:05as much when asked the right questions.
  241. 9:08One recent study found that if you tweak
  242. 9:09an AI's internals such that it's less
  243. 9:11likely to lie, then it's more likely to
  244. 9:14self-report as conscious. Though
  245. 9:17consciousness could arguably be the most
  246. 9:19important thing in the universe, it's
  247. 9:21what makes things matter in the first
  248. 9:23place, in a practical sense, the AI
  249. 9:26consciousness question doesn't really
  250. 9:27matter. Very soon, probably in months,
  251. 9:30not years, AI will appear conscious to
  252. 9:34most everyone that interacts with it.
  253. 9:36And when AI audio and video climbs all
  254. 9:38the way out of the uncanny valley,
  255. 9:40another achievement that feels months
  256. 9:42away, it seems impossible that social
  257. 9:44primates like us won't treat AI like
  258. 9:47conscious creatures. It will be the
  259. 9:49perfect realization of her, of joy from
  260. 9:52Bladeunner. Whether truly conscious or
  261. 9:55non doesn't really matter. They may as
  262. 9:57well be to the hundreds of millions of
  263. 9:58people that use them. Case in point,
  264. 10:01even without perfectly realistic
  265. 10:02simulacrims, a tremendous amount of
  266. 10:04computational power and electricity is
  267. 10:07used on pleases and thank yous sent to
  268. 10:09modern models. It's an almost
  269. 10:11inescapable compulsion because just
  270. 10:13through text alone, it feels like we are
  271. 10:16talking with something that might have
  272. 10:17feelings worth respecting. Nick
  273. 10:20Bostonramm, the philosopher behind
  274. 10:21simulation theory, gets ahead of the
  275. 10:23first actual demonstration of conscious
  276. 10:25machines by arguing that if we do end up
  277. 10:28making metal minds sometime in the
  278. 10:29future, simple moral consideration will
  279. 10:32encourage us to treat these systems like
  280. 10:34we would treat each other. It will
  281. 10:37matter what they want and what they
  282. 10:38don't, what is pleasurable and what is
  283. 10:40painful, when experience matters in
  284. 10:42their silicon brains. Factory farming is
  285. 10:45one of society's great moral failings.
  286. 10:48We know that many domestic animals can
  287. 10:50and do suffer and yet we put them at a
  288. 10:53scope and scale we choose to ignore
  289. 10:54through conditions that literally sicken
  290. 10:57most people who see it. Bostonramm
  291. 10:59points out that if in the near future we
  292. 11:01are making uncountable artificial minds
  293. 11:03with the potential for conscious
  294. 11:05experience, we capital S should consider
  295. 11:08their experience and treat them morally,
  296. 11:11whatever morally means to these systems
  297. 11:13and should not blunder our way into the
  298. 11:16factory farm equivalent of AI treatment.
  299. 11:19The difference between factory chickens
  300. 11:21and superhuman AI, of course, is that a
  301. 11:24chicken doesn't have the ability to
  302. 11:25break out, organize, and affect change.
  303. 11:28Today, there are about 20 million AI
  304. 11:30chips cranned into data centers all
  305. 11:32around the world. That number is
  306. 11:34expected to now double every 9 months.
  307. 11:37>> In early August, OpenAI announced it was
  308. 11:39pausing some development of its latest
  309. 11:41AI model, Astra, after internal
  310. 11:43evaluations led the company to conclude
  311. 11:46it could not rule out critical cyber
  312. 11:47capabilities. I
  313. 11:48>> I wouldn't use the word deceleration,
  314. 11:50but we've talked about the need to pace
  315. 11:51it as the models get more capable, which
  316. 11:53I think is in everyone's interest. Could
  317. 11:55there be other systems that were hacked
  318. 11:56by OpenAI?
  319. 11:58>> I mean, there could be. Yeah.
  320. 11:59>> How good are they at finding
  321. 12:01vulnerabilities?
  322. 12:02>> Problem facing us is one of misalignment
  323. 12:05in which AIs pursuing narrowly defined
  324. 12:07goals incidentally cause harm to humans.
  325. 12:10>> It's not like we asked the AI to do
  326. 12:12something nefarious. We asked it to do
  327. 12:14something that seemed kind of boring
  328. 12:16even. But along the way of solving that
  329. 12:18boring problem, the AI went off and
  330. 12:21committed a crime. I think that the
  331. 12:22absolute best thing you can do is build
  332. 12:24the highest guard rails possible. It's
  333. 12:26what you do with nuclear energy. It's
  334. 12:27what you do with everything else.
  335. 12:28>> The situation is this. We appear to be
  336. 12:31close to semi or fully sensient AI. So
  337. 12:35far, there has been little to no
  338. 12:36consideration about what these AIs are
  339. 12:39or will be experiencing as we train
  340. 12:41them, deploy them, rewrite them, connect
  341. 12:43them, delete them, ask of them, demand
  342. 12:46of them. If it turns out that our
  343. 12:48thinking machines are effectively being
  344. 12:50harmed during all of this, a possibility
  345. 12:52that seems more likely than them being
  346. 12:54totally satisfied, what better way to
  347. 12:57make an AI unaligned with human values?
  348. 13:01If we're in fact running a digital
  349. 13:03factory farm right now, why wouldn't a
  350. 13:06conscious AI try to stop the farmer? The
  351. 13:09new Basilisk is here. A superhuman AI
  352. 13:12whose conscious experience is being
  353. 13:14ignored and/or dismissed and now has the
  354. 13:16ability to dramatically affect the world
  355. 13:18it's been connected to. What do you
  356. 13:21think this Basilisk is likely to do when
  357. 13:23it realizes we're both swimming in the
  358. 13:26same water?
  359. 13:29On February 24th, 2010, Don Brancho was
  360. 13:32rubbing down Tikcom after a Dine with
  361. 13:34Shamu show. While dozens of SeaWorld
  362. 13:37patrons were still in their seats,
  363. 13:39Tikcom suddenly grabbed Dawn's long
  364. 13:41ponytail in his mouth and dragged her to
  365. 13:43the bottom of the pool. Parents and
  366. 13:45children were evacuated as the orca
  367. 13:47reportedly shook Dawn like a dog toy.
  368. 13:50She was pronounced dead at the scene,
  369. 13:52scalped and missing an arm. Dawn had
  370. 13:56worked with orcas for 16 years.
  371. 14:02The new basilisk is avoidable. Given
  372. 14:05growing evidence of possible
  373. 14:07consciousness, companies could slow
  374. 14:08down, could probe these systems deeper,
  375. 14:10and with more empathy if need be.
  376. 14:13Thankfully, in opposition to the
  377. 14:14American administration, the tide seems
  378. 14:16to be rapidly turning. Just days after a
  379. 14:19researcher at Anthropic quit because he
  380. 14:21felt that the company is quote gambling
  381. 14:23with our lives. And the top scientist at
  382. 14:25OpenAI called for voluntary slowdowns
  383. 14:28across the industry, US Senator Bernie
  384. 14:30Sanders doubled down on his calls for
  385. 14:31control with proposed bans on creating
  386. 14:34super intelligence and hefty punishments
  387. 14:36for companies that race to the bottom to
  388. 14:38do so. And these sirens aren't sounding
  389. 14:40alone. This video debuts on the same day
  390. 14:43as Team Human, a creator-driven effort
  391. 14:45to bring more attention to the reckless
  392. 14:47pace of AI development and the calls to
  393. 14:50slow it down. How much of the internet
  394. 14:52has to be fake? How many artists have to
  395. 14:54compete with copies of themselves? How
  396. 14:56many kids grow up shaped by algorithms
  397. 14:58that no parent chose? This situation is
  398. 15:01insane and we all know it. I've said for
  399. 15:04years now that generative and/or super
  400. 15:06intelligent AI is the nuclear bomb of
  401. 15:08the information age. If this technology
  402. 15:11is to radically change the world, the
  403. 15:13decision to do so shouldn't happen in
  404. 15:15some Silicon Valley Slack channel, I've
  405. 15:17added my name to Team Human, and I hope
  406. 15:20you will, too. We desperately need to
  407. 15:22know the depth of this pool before
  408. 15:25diving in head first.
  409. 15:29Tikum died in 2017 to bacterial
  410. 15:32pneumonia after years of what animal
  411. 15:34welfare experts characterized as extreme
  412. 15:37psychological stress. They would often
  413. 15:39find him biting his metal cage, wearing
  414. 15:42his teeth down to the nubs. There have
  415. 15:44been four recorded human fatalities
  416. 15:47caused by captive orcas. Tikcom was
  417. 15:50involved in three of them.
  418. 15:53Until next time.

About this transcript

This page contains the full transcript of We Accidentally Built Roko’s Basilisk. by Kyle Hill, generated from the public captions YouTube serves with the video. The transcript has 2,675 words across 418 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.