YouTube2Text

Claude for Long-Horizon Tasks — Lance Martin, Anthropic — Transcript

by AI Engineer · 4,469 words · 743 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[music]
  2. 0:12>> Good to go.
  3. 0:14All right. Well,
  4. 0:16take a quick sip and then let's start.
  5. 0:19It is great to be here. Um this is like
  6. 0:21like my third year coming to this
  7. 0:22conference and I always really enjoy it.
  8. 0:25And thank you for coming to this
  9. 0:26workshop. I know there's many
  10. 0:27interesting talks.
  11. 0:29Let me talk a little bit about um
  12. 0:31our view of async agents at Anthropic
  13. 0:33and some things we've been up to lately.
  14. 0:36So, this is kind of a way I think about
  15. 0:38models in product. So, you can think
  16. 0:40about Claude as a light source and you
  17. 0:42can think about products as windows that
  18. 0:44allow the light to pass through.
  19. 0:46And what's kind of interesting is over
  20. 0:48time, the window that you need to
  21. 0:50actually kind of see the light of the
  22. 0:51model kind of shifts.
  23. 0:54And we've seen this over the past few
  24. 0:56years. So, I'm plotting here different
  25. 0:58Claude models and their task horizon.
  26. 1:00So, how much autonomous work can they do
  27. 1:03over time?
  28. 1:04And you might recall back in like the
  29. 1:06Opus 3 days, this was kind of like 2024,
  30. 1:09models could only do, you know, maybe 10
  31. 1:11to 20 minutes of autonomous work. This
  32. 1:13is measured by meter.
  33. 1:15And in that regime, only certain product
  34. 1:17surfaces made sense. Things like
  35. 1:18autocomplete, things like chat, where
  36. 1:20your human is very in the loop cuz the
  37. 1:22model's really only doing a very short
  38. 1:23amount of work before you're steering
  39. 1:25it.
  40. 1:26Now, the past year we saw the rise of
  41. 1:28synchronous coding agents like Claude
  42. 1:29code. And this is, you know, a kind of a
  43. 1:31shift because then models could do maybe
  44. 1:33an hour of work.
  45. 1:35So, it made sense to have them run, but
  46. 1:37typically locally, where you could still
  47. 1:39steer them easily.
  48. 1:41And it's kind of interesting because
  49. 1:42during this regime, I remember efforts
  50. 1:44and I was involved in some efforts to
  51. 1:45build kind of async agents.
  52. 1:48But when models can only do like an hour
  53. 1:50of work, async as an experience is kind
  54. 1:52of bad. Um the model goes off and it
  55. 1:55like hits an error and it comes back to
  56. 1:57you over a short period of time.
  57. 1:59In order to really unlock async, we
  58. 2:00needed longer task horizons.
  59. 2:03And so we're starting to see that now.
  60. 2:06And kind of with this shift in
  61. 2:08capability and time horizon came a shift
  62. 2:10in the API surfaces. So if you look at
  63. 2:12the lower left,
  64. 2:15Messages API came out like 2 years ago.
  65. 2:17It's basically prompt response. It's
  66. 2:19great for building harnesses,
  67. 2:22but it's a very simple API. Again,
  68. 2:24you're just passing in a message and you
  69. 2:25get a response out.
  70. 2:27There's no sense of deployment with
  71. 2:29that. So you basically take Messages API
  72. 2:30and you can roll your own harness, you
  73. 2:32can deploy that harness, and you have an
  74. 2:33agent.
  75. 2:34Now over the past year, we saw the rise
  76. 2:37of, you know, coding agents in
  77. 2:38particular. So we released Agent SDK. So
  78. 2:40that's basically a way to
  79. 2:41programmatically call Claude code.
  80. 2:43And that's like basically us giving you
  81. 2:45a harness.
  82. 2:47But over the past few months, since
  83. 2:49April, as we've seen longer longer task
  84. 2:51horizons, we released a new API called
  85. 2:53Managed Agents,
  86. 2:55which basically packages both the
  87. 2:57harness as well as all the managed
  88. 2:58deployment infrastructure for you.
  89. 3:01And I'll talk about some of the themes
  90. 3:02that underpin this new uh surface,
  91. 3:04Claude Managed Agents, and and some of
  92. 3:06the themes that kind of extend beyond
  93. 3:08just Managed Agents broadly to think
  94. 3:10about this kind of new type of
  95. 3:11asynchronous agents, uh which can apply
  96. 3:13of course to Claude and other types of
  97. 3:15kind of longer running long horizon
  98. 3:17agents.
  99. 3:19So theme one is decoupling the brain
  100. 3:21from the hands.
  101. 3:23Um
  102. 3:24so when we first set out to build
  103. 3:26Managed Agents, we started with the
  104. 3:27container.
  105. 3:28We put the harness in the sandbox in the
  106. 3:30same container.
  107. 3:32Now the problem here is, what happens if
  108. 3:36the harness dies or the container dies?
  109. 3:39What we saw is you actually lose the
  110. 3:40session.
  111. 3:41So basically, this architecture is kind
  112. 3:44of tricky for long horizon agents
  113. 3:46because what can happen is your agent's
  114. 3:48running, and And container dies, you
  115. 3:51lose everything with it.
  116. 3:54Also, as models get more capable,
  117. 3:56putting the credentials in the same
  118. 3:58container with the agent itself can be
  119. 4:01problematic.
  120. 4:02So, for example, giving Claude access to
  121. 4:03a bunch of your secrets and letting it
  122. 4:05run for 10 hours and not watching it can
  123. 4:07be a little bit spooky and have some
  124. 4:08security concerns, especially as models
  125. 4:10get extremely capable.
  126. 4:13So, for this reason, we kind of decouple
  127. 4:15what we call the brain, that's the
  128. 4:16harness, from the hands, the execution
  129. 4:18environments, and manage agents to set
  130. 4:19up like this.
  131. 4:21So, the story here is that the harness
  132. 4:25becomes a stateless process
  133. 4:27that talks to a session. The session is
  134. 4:30an append-only event log
  135. 4:33and that can reach out to hands, which
  136. 4:34are just containers.
  137. 4:36So, that's just Those are sandboxes
  138. 4:38where work is done.
  139. 4:39And one thing that's interesting is
  140. 4:41Claude is increasingly capable of
  141. 4:42managing many hands.
  142. 4:44So, that is you can give one harness
  143. 4:46access to many different containers to
  144. 4:48perform to execution.
  145. 4:50And Claude can manage this very easily
  146. 4:52and and and and effectively.
  147. 4:55If the session, uh sorry, if the harness
  148. 4:58dies or sandbox dies, it's completely
  149. 5:01fine because the session is always
  150. 5:03backed up in this append-only log and
  151. 5:05credentials are never actually added to
  152. 5:07the sandbox.
  153. 5:08They're stored in a separate vault. So,
  154. 5:10this decoupling actually makes it quite
  155. 5:12reliable and safe, particularly for
  156. 5:14long-horizon tasks. And this is kind of
  157. 5:15one of the core ideas that underpins
  158. 5:17manage agents architecturally.
  159. 5:20And I think an interesting thing that
  160. 5:21falls out of this is related to you guys
  161. 5:24may have kind of seen or come across the
  162. 5:26recursive language models work.
  163. 5:28The session becomes an external context
  164. 5:30object that the model interrogates.
  165. 5:33And this has all sorts of benefits for
  166. 5:35context management. So, you think about
  167. 5:37it, when you're doing something like
  168. 5:38compaction, you're choosing some logic
  169. 5:40to retain some amount of context, and
  170. 5:43naively in a typical in a kind of a
  171. 5:45typical step, you're discarding all the
  172. 5:46context that you didn't compact.
  173. 5:49In this architecture, and also more
  174. 5:50broadly with recursive language models,
  175. 5:52the idea is that the context object is
  176. 5:54persistent and is unadulterated. So,
  177. 5:57it's append-only.
  178. 5:58And the model can always go back and
  179. 6:01fetch old context. So, basically creates
  180. 6:02a very very nice architecture for
  181. 6:04context engineering because the core
  182. 6:06context object is immutable in the sense
  183. 6:08that it's it's it's non-destructive
  184. 6:11and you can only and you only append to
  185. 6:13it over time.
  186. 6:14So, we've seen this be to be quite nice
  187. 6:16in terms of long horizon context
  188. 6:17engineering as well.
  189. 6:20So, the second theme is use verifiers.
  190. 6:25And one of the problems that we've seen
  191. 6:28with Claude and other models in general
  192. 6:31is that when you ask them to do a bunch
  193. 6:32of work and then say, "Okay, grade your
  194. 6:35work."
  195. 6:37If that same context is being used to
  196. 6:39both do the work and grade, you can get
  197. 6:41lots of odd artifacts and confabulation
  198. 6:44and and basically odd behavior.
  199. 6:47For example, this is just an image
  200. 6:48showing you can think about that that
  201. 6:50context window is filled with lots of
  202. 6:51different information and the model is
  203. 6:53grading itself, often it's not properly
  204. 6:56tuned to do kind of critical
  205. 6:58verification.
  206. 7:00And so, what we found is it's quite
  207. 7:02effective to separate verification into
  208. 7:04a separate context window.
  209. 7:06This is a very general trend.
  210. 7:08We talked about it in a number of
  211. 7:09different engineering blogs.
  212. 7:11Um and the reason is the verifier
  213. 7:13context can be tuned very specifically
  214. 7:15for the critical verification task.
  215. 7:18>> [snorts]
  216. 7:18>> And so, the way this works in practice
  217. 7:20is
  218. 7:21when you build loops,
  219. 7:23you can have a loop of a build context
  220. 7:24and a verifier context. And this can be
  221. 7:25a build agent verifier agent. And what
  222. 7:27happens is
  223. 7:28the verifier has some goal
  224. 7:31or rubric
  225. 7:32and it's verifying the result or work of
  226. 7:35the build agent. And this continues in a
  227. 7:37loop until verification is complete.
  228. 7:40And this is really the big idea behind
  229. 7:41this whole loops trend that you might
  230. 7:42have heard about and we found it to be a
  231. 7:44kind of a very powerful paradigm
  232. 7:45especially for working with some of the
  233. 7:47higher capacity models.
  234. 7:50And here are some of the primitives. So
  235. 7:51in Claude code you have goal and manage
  236. 7:54agents you have outcomes and the
  237. 7:55principles are really the same.
  238. 7:58You're setting up a measurable end state
  239. 7:59in both cases.
  240. 8:01You're using an independent context
  241. 8:02model.
  242. 8:03You're using independent context to
  243. 8:05grade
  244. 8:07over the course of this loop.
  245. 8:09The loop can run and you only exit the
  246. 8:11loop until this independent verifier has
  247. 8:13verified that it has the outcomes or
  248. 8:14outputs that you want. That's the key
  249. 8:16idea.
  250. 8:18Now let me tell you a story about how
  251. 8:20I've used this. So this is a kind of a
  252. 8:22fun and interesting challenge called
  253. 8:23parameter golf. It's a benchmark that
  254. 8:25set up that was put up by Open AI.
  255. 8:28And it tests models ability to
  256. 8:30effectively do kind of ML research.
  257. 8:33So it asks the model to basically take a
  258. 8:36small um kind of model and train it in
  259. 8:40with eight with eight 8100 GPUs in less
  260. 8:44than 10 minutes. And what you see on the
  261. 8:46Y is basically you can think about it as
  262. 8:48loss. So lower is better, okay?
  263. 8:51And what I did was I set up a kind of a
  264. 8:53verifier loop using manage agents and
  265. 8:55outcomes
  266. 8:57to test the ability for Opus 4.7 and one
  267. 9:00of our frontier models like mythos class
  268. 9:02models on this task. And what you see is
  269. 9:06basically I allow the model to continue
  270. 9:08to iterate until the outcome that I
  271. 9:10specify is is satisfied which is it
  272. 9:13finished exactly 20 iterations and kind
  273. 9:17of it kind of met all the experimental
  274. 9:18criteria as defined by the benchmark.
  275. 9:21What you see is
  276. 9:23the frontier capability models
  277. 9:25are extremely good with this pattern of
  278. 9:27kind of loops in software and and kind
  279. 9:29of verification because what happens is
  280. 9:33instead of
  281. 9:35encoding steering me and into like me as
  282. 9:38the human, you're encoding the signal
  283. 9:40into the environment. So, the model can
  284. 9:42self-correct when it receives feedback
  285. 9:44from, for example, the verifier.
  286. 9:47And using this kind of paradigm with
  287. 9:48very high-capacity models, you can get
  288. 9:50very strong results.
  289. 9:52So, the main point I'm trying to make
  290. 9:54here is that this paradigm of loops,
  291. 9:56which a lot of people been talking about
  292. 9:57today,
  293. 9:59paired with very capacity models is a
  294. 10:01very good general primitive for
  295. 10:03long-running asynchronous work.
  296. 10:05That's really the key point here.
  297. 10:07>> [snorts]
  298. 10:09>> Now, let me talk another about another
  299. 10:11theme of self-learning.
  300. 10:14So, the human brain has two kind of
  301. 10:18interesting systems for memory.
  302. 10:21So, one is as you go about your day, the
  303. 10:22hippocampus is kind of writing traces of
  304. 10:25kind of short-term, very fast kind of
  305. 10:26experiential memory. Like, you might
  306. 10:28remember what you had for lunch today.
  307. 10:29You had lunch an hour ago, you kind of
  308. 10:31remember, that's kind of written to
  309. 10:32short-term memory.
  310. 10:33When you go to bed at night, though, an
  311. 10:35offline process or out-of-band process,
  312. 10:37dreams. And dreaming stores certain
  313. 10:41important details to long-term memory in
  314. 10:42the cortex. So, for example, tomorrow,
  315. 10:44you might not remember what you ate for
  316. 10:45lunch today. That's kind of a local
  317. 10:46trace. But, like, if you had a very
  318. 10:48important experience today, maybe this
  319. 10:50talk, maybe not this talk, but if you
  320. 10:53had an interesting experience today,
  321. 10:54that might be written to long-term
  322. 10:55memory. That's kind of the point. These
  323. 10:56two subsystems in human work like this.
  324. 10:59And we've actually found memory systems
  325. 11:00with Claude actually can employ these
  326. 11:02same two principles.
  327. 11:04So, this is showing Claude's capacity as
  328. 11:08an in-band memory writer. So, you
  329. 11:10basically give Claude memory tools. And
  330. 11:12when I say memory tools, I mean it's
  331. 11:13basically the ability to write to a file
  332. 11:15system, that is basically a memory
  333. 11:16directory. That's really it.
  334. 11:19Now, this is showing some work on Claude
  335. 11:21plays Pokémon with Claude's on it 3.5.
  336. 11:23And here's the key point. When Sonnet
  337. 11:263.5 is given this access to a memory
  338. 11:28directory, and it can write memory,
  339. 11:30{quote} in-band, as it progresses
  340. 11:32through this game, it's not very good.
  341. 11:35So, the memories it writes are pretty
  342. 11:36crappy. It's kind of um
  343. 11:38it's kind of uh
  344. 11:40tactical notes. It's it's not very
  345. 11:42strategic. And the game progress is
  346. 11:44quite limited.
  347. 11:46But with more recent models like this is
  348. 11:48looking at 4 6,
  349. 11:50the notes are much more strategic
  350. 11:52and game progress is much further. So,
  351. 11:53the key point I'm making here is that
  352. 11:56Claude has gotten much better at this
  353. 11:57in-band memory writing across model
  354. 11:59generations.
  355. 12:02And this is another way to show that
  356. 12:04same result.
  357. 12:05So, this is a benchmark that I ran
  358. 12:07called Continual Learning Bench. It's an
  359. 12:08open-source benchmark. I took one of the
  360. 12:10tasks. This is a task that basically
  361. 12:11asked the model to perform a sequential
  362. 12:14question answering with a SQL database
  363. 12:16and it can write memory in between each
  364. 12:17step.
  365. 12:19And what you see is basically the
  366. 12:21performance improves across models.
  367. 12:25Um
  368. 12:26and so what this is kind of showing is
  369. 12:29that models get natively better at this
  370. 12:32in-band memory writing with respect to
  371. 12:34model capability.
  372. 12:35>> [snorts]
  373. 12:36>> And some of the most interesting things
  374. 12:37I found from this are that the main
  375. 12:39differentiation between There we go. Um
  376. 12:44the main differentiation
  377. 12:46between like a lower capacity model and
  378. 12:48a high capacity model is kind of this
  379. 12:50distillation step. And so, basically,
  380. 12:53higher capacity models have a better
  381. 12:54sense of like what abstraction to save
  382. 12:57to memory that'll be useful later.
  383. 13:00Like they're not just writing a specific
  384. 13:02fact. They're writing how do I how does
  385. 13:03this generalize to future sessions?
  386. 13:05That's kind of the key difference that I
  387. 13:07found that higher capacity models kind
  388. 13:09of have when they're writing memory. So,
  389. 13:11this is a very important thing to keep
  390. 13:12in mind that models are getting better
  391. 13:13better and better at this kind of
  392. 13:14in-band memory writing across model
  393. 13:16generations.
  394. 13:19Now, there's a little trick here which
  395. 13:21is very important. So, we talked about
  396. 13:22kind of in-band memory and we talked
  397. 13:24about dreaming. So, at night I dream and
  398. 13:26I write things to like long-term memory.
  399. 13:28Dreaming is very important because when
  400. 13:30I'm writing memory in band over the
  401. 13:32course of a day over the course of a
  402. 13:33session, sometimes you can write
  403. 13:35incorrect memories.
  404. 13:37And
  405. 13:38or you're writing things that are
  406. 13:40locally optimal, but not globally
  407. 13:42optimal. So, you're writing over the
  408. 13:43course of a task to like kind of help
  409. 13:47you solve that task, but not necessarily
  410. 13:49looking forward to future tasks.
  411. 13:51This is a very important nuance in this
  412. 13:53process of dreaming is is kind of an
  413. 13:55offline or out of band process that
  414. 13:56we've used to consolidate and improve
  415. 13:58memory.
  416. 14:00And I want to show you a fun example
  417. 14:02that I've used dreaming for.
  418. 14:04So, this is again Pokémon.
  419. 14:06And this is I played a lot of games of
  420. 14:09Pokémon with Claude to find this.
  421. 14:12So, this was a very hard one lesson, so
  422. 14:14I hope you appreciate it. Okay, here's
  423. 14:16the point. So, basically what happened
  424. 14:18is
  425. 14:20Claude wrote an incorrect memory, okay?
  426. 14:23And what happened is this incorrect
  427. 14:25memory was related to the location of
  428. 14:27The details don't necessarily matter.
  429. 14:29Uh
  430. 14:30the point is that this incorrect memory
  431. 14:33causes Claude to mislocalize itself or
  432. 14:35Pokémon to mislocalize itself,
  433. 14:38and it falls through this trapdoor,
  434. 14:40okay? That's the key point. So, it
  435. 14:42writes this incorrect memory. This
  436. 14:43incorrect memory causes it to
  437. 14:44mislocalize in the game, and it falls
  438. 14:47down this trap. This is very consistent.
  439. 14:49So, I saw this in five replicates. Five
  440. 14:50out of five replicates with raw memory
  441. 14:52store fell down this trap. With the
  442. 14:55dreaming,
  443. 14:56this error is corrected, and it's able
  444. 14:58to properly localize itself and not fall
  445. 15:02fall down this trap. And I'll show you
  446. 15:03kind of a fun visualization of this.
  447. 15:05This is looking at kind of memory traces
  448. 15:07or basically traces of game progress.
  449. 15:09So, going upwards on the Y axis is
  450. 15:12improvement. That's like moving to the
  451. 15:13next level. Going down is backtracking.
  452. 15:16So, what's interesting here,
  453. 15:18the no memory baseline, which is that
  454. 15:20gray bar, kind of doesn't make much
  455. 15:22progress at all. It just kind of like is
  456. 15:25stuck. It's like a particularly hard
  457. 15:26level, okay?
  458. 15:28The memory, which is the orange,
  459. 15:30actually keeps falling down this
  460. 15:31trapdoor and falls back. So, it
  461. 15:33backtracks. The dreaming traces though
  462. 15:36consistently kind of fix this error in
  463. 15:38its memory and proceed to the next
  464. 15:40level. So, this is a very practical
  465. 15:42example of how dreaming can kind of work
  466. 15:44out of band on your memory store to fix
  467. 15:47corrections. Because what it does is it
  468. 15:49looks at your memory store and it looks
  469. 15:50at all your prior traces or sessions and
  470. 15:53kind of can find and correct errors.
  471. 15:54That's the key point and that's why the
  472. 15:56dreaming process can be very helpful
  473. 15:58because in band, while Claude is writing
  474. 16:00to memory, it can make mistakes. And
  475. 16:02those mistakes get stuck in memory
  476. 16:03unless you have an offline process to
  477. 16:05kind of correct them. That's the key
  478. 16:06intuition.
  479. 16:08Um
  480. 16:09and theme four, and I'll open up for
  481. 16:11questions after this, is what I think um
  482. 16:14is kind of this trend that we're going
  483. 16:15to see moving towards org-level
  484. 16:17harnesses with async agents.
  485. 16:19And
  486. 16:21so, we released Claude Tag
  487. 16:23and a lot of the reaction was like, "Ah,
  488. 16:26Slack bot."
  489. 16:28And like, look, I have actually created
  490. 16:29a lot of Slack bots myself. I understand
  491. 16:31not every Slack bot is particularly
  492. 16:32interesting or great. In fact, I've
  493. 16:33created many Slack bots that are quite
  494. 16:35bad.
  495. 16:36But
  496. 16:37what's interesting about Claude Tag is
  497. 16:39not the fact that it's accessible
  498. 16:41through Slack. What's interesting about
  499. 16:42it is the fact that it has a very very
  500. 16:44rich kind of system underneath it, which
  501. 16:46I want to just cut touch on briefly.
  502. 16:49And in particular, what's interesting
  503. 16:50about it is it represents what I
  504. 16:51consider an org-level harness.
  505. 16:54So, agents historically have been kind
  506. 16:55of single player. So, you have an an
  507. 16:58agent like Claude code on your machine
  508. 16:59with your local context that you've
  509. 17:01tuned and configured for yourself.
  510. 17:03What's interesting about Claude Tag is
  511. 17:05it is a harness that everyone in the
  512. 17:06organization has access to and can use.
  513. 17:09So, it is a multiplayer harness.
  514. 17:11And what's nice about that is it has its
  515. 17:13own identity. Its identity and
  516. 17:14credentials are not tied to a given user
  517. 17:16and has access to organizational level
  518. 17:18context, not just my local context.
  519. 17:20This has many interesting and useful
  520. 17:23implications.
  521. 17:24Including the ability to like check
  522. 17:26others work before you do an experiment,
  523. 17:28the ability to kind of deduplicate
  524. 17:30findings, the ability to like do
  525. 17:31internal research, the ability to give
  526. 17:34everyone access to kind of a a very well
  527. 17:36developed harness on day one, whereas
  528. 17:38when you your own personal harness,
  529. 17:40often new employees takes them weeks or
  530. 17:43maybe even months to kind of ramp up
  531. 17:44fully to configure all the right
  532. 17:46connectors and so forth. So, org level
  533. 17:48harnesses are real leveler of the
  534. 17:50playing field and I think it was
  535. 17:52people kind of saw the Slack bot piece,
  536. 17:53but they didn't really appreciate the
  537. 17:55depth of kind of
  538. 17:56that the depth of benefit you get from
  539. 17:58building out org level harnesses. So, I
  540. 18:00do think that was kind of important
  541. 18:01thing to note.
  542. 18:03And I think we're going to see the rise
  543. 18:05of kind of harnesses that operate across
  544. 18:07orgs, across many different users that
  545. 18:09can operate increasingly on longer async
  546. 18:11async
  547. 18:13um
  548. 18:14async agents that can operate on longer
  549. 18:16time frames. That's kind of one clear
  550. 18:17kind of follow-up that we're I think
  551. 18:19we're going to see from this.
  552. 18:20And another thing where I think we're
  553. 18:22going to see is that
  554. 18:24asynchronous agents um
  555. 18:26are going to be increasingly proactive.
  556. 18:28Uh so, typically with for example like
  557. 18:30local locally scoped agents, they tend
  558. 18:33to be reactive. They're responsive to
  559. 18:34how you steer it. versus async agents
  560. 18:37increasingly have the ability to steer
  561. 18:39proactivity. And that's one very nice
  562. 18:41thing about Cloud Tag.
  563. 18:43Um where basically you can configure it
  564. 18:46to tell you things when like looking at
  565. 18:49this org level context, alert me with
  566. 18:51things I might need to know about.
  567. 18:53Um and this is a very important kind of
  568. 18:55new kind of UX that I think is going to
  569. 18:57be more and more common with async
  570. 18:58agents that kind of access to
  571. 18:59organizational context.
  572. 19:01And of course multiplayer. So, the the
  573. 19:03ability for a single harness to be
  574. 19:05steered by many many different people
  575. 19:06kind of concurrently is an important
  576. 19:08shift in agent UX
  577. 19:10uh that I think uh will be quite
  578. 19:12interesting going forward. So,
  579. 19:15um,
  580. 19:16yeah, let me let me just open up for
  581. 19:17questions and um, thank you for
  582. 19:19listening.
  583. 19:21>> [applause]
  584. 19:26[applause]
  585. 19:30>> Sure.
  586. 19:31>> Uh, one thing that we see sort of
  587. 19:33empirically and also in the benchmarks
  588. 19:35is that the frontier models perform
  589. 19:37better on these long horizon tasks than
  590. 19:39on the ones that have stacking. What
  591. 19:41they need to keep in mind that this is
  592. 19:43not
  593. 19:44What is your view on and like the
  594. 19:47guidance on it and like the right away
  595. 19:48model?
  596. 19:49>> Yeah.
  597. 19:53>> Um, what's your view on why that is and
  598. 19:56then how long the frontier will be able
  599. 19:58to maintain that gap?
  600. 20:00>> I see. So, the question was kind of on
  601. 20:02the the gap between the frontier models
  602. 20:04kind of on for example, a benchmark like
  603. 20:07meter on like long horizon tasks.
  604. 20:09>> Yeah.
  605. 20:09>> Okay. Um,
  606. 20:12so,
  607. 20:13in the latest results that I saw from
  608. 20:15like for example, Codex like 56, I think
  609. 20:17I'd also kind of is in that 12 plus hour
  610. 20:20regime on meter. So, I think you're
  611. 20:21right that like the frontier models like
  612. 20:23the, you know, mythos class models,
  613. 20:25strong models from Open AI kind of are
  614. 20:27in this like 12 plus hour regime.
  615. 20:29Um,
  616. 20:32why is it that
  617. 20:34non-frontier models are not kind of in
  618. 20:36that regime? I am actually not
  619. 20:37necessarily sure. I do think
  620. 20:40I do think um,
  621. 20:42in order to build agents that can
  622. 20:45effectively operate in this regime, it's
  623. 20:48important to note that it's not just the
  624. 20:49model capability. Like for example, with
  625. 20:51a Claude tag product, it's actually a
  626. 20:52combination of improvement in memory
  627. 20:55because memory is very important. If you
  628. 20:57have agents working for for example, 12
  629. 20:58hours, you want to make sure that your
  630. 21:00productivity preferences are well
  631. 21:02encoded in memory. So, it knows when to
  632. 21:04reach out to you if it gets stuck for
  633. 21:05example. So, memory is very important,
  634. 21:07security is very important, so
  635. 21:09resistance to prompt injection. Um
  636. 21:11also like model architecture, like kind
  637. 21:13of the agent architecture is very
  638. 21:14important, that decoupling of brain and
  639. 21:16hand, so it's secure and safe and like
  640. 21:17resistant to failure. So, actually I
  641. 21:19think
  642. 21:21to build real agents that can operate in
  643. 21:23these long time horizons, a bunch of
  644. 21:24things need to come together in terms of
  645. 21:25like architecture,
  646. 21:27infrastructure, security, memory. And
  647. 21:31that might be why and and but Frontier
  648. 21:32Labs invested in all these areas, so
  649. 21:34that might be why it's become a gap in
  650. 21:36terms of like the agent products that
  651. 21:37we've released. And so we spent a lot of
  652. 21:38time, for example, building managed
  653. 21:39agents
  654. 21:41to kind of have these kind of
  655. 21:42considerations baked in.
  656. 21:44Yeah.
  657. 21:45Sure.
  658. 22:05Yeah, okay, this is interesting.
  659. 22:07Um the question was about um kind of
  660. 22:10like the the best memory substrate, so
  661. 22:12like why file systems versus, for
  662. 22:13example, databases.
  663. 22:15Um
  664. 22:17this is kind of a subtle point that
  665. 22:18actually um
  666. 22:21I want to think about carefully. So,
  667. 22:24I don't necessarily think that
  668. 22:28it has to be the case that you use a
  669. 22:29file system for memory. I think what's
  670. 22:31quite important that we've seen is that
  671. 22:33it's you want something that is highly
  672. 22:36programmable
  673. 22:38with simple primitives that the model
  674. 22:39can manipulate to like write, manage its
  675. 22:41own memory. So, for example,
  676. 22:43a database could work fine relative to
  677. 22:46the file system.
  678. 22:47But what I've seen doesn't work is when
  679. 22:50you specify the structure of memory for
  680. 22:53the model very explicitly, whether
  681. 22:55that's in a file system or database or
  682. 22:56whatever. Like a memory schema, I pre I
  683. 22:58kind of pre populate, here's the types
  684. 23:00of memories you need to save.
  685. 23:02Cuz I think that ends up being not very
  686. 23:03bitter lesson pill in the sense that
  687. 23:04models can learn to manage their own
  688. 23:06memory much better than you can intuit
  689. 23:08these memory types for the model ahead
  690. 23:10of time. So, I think what we've seen is
  691. 23:12that very general substrates for memory,
  692. 23:15be it just your database or file system,
  693. 23:17are good because the model can manage
  694. 23:19them freely versus
  695. 23:20a very very kind of like prescriptive
  696. 23:23memory schema they are trying to
  697. 23:25pigeonhole the model into. That's when
  698. 23:27you see performance drop. That's the key
  699. 23:29differentiation.
  700. 23:35Right.
  701. 23:37That's the key point. Let the model
  702. 23:39structure and maintain its own memory.
  703. 23:40Don't give it a prescribed memory
  704. 23:42schema.
  705. 23:43And that's like a common failure because
  706. 23:45models are getting good enough that they
  707. 23:46can manage their own memory much more
  708. 23:47effectively than you can reason about
  709. 23:49types of This is like very classically
  710. 23:51bitter lesson pill, but like
  711. 23:53you can Models can reason about their
  712. 23:54own memory and context structure much
  713. 23:56better than you can prescribe for them a
  714. 23:58way to structure their own memories.
  715. 24:00That's the key observation.
  716. 24:02Yep.
  717. 24:07Yes.
  718. 24:10That's right. Exactly.
  719. 24:12General substrates for for memory
  720. 24:13management.
  721. 24:15Yep.
  722. 24:19Yes.
  723. 24:22Okay. So, that's a good point.
  724. 24:26Basically, the question is
  725. 24:28so, you do this dreaming thing.
  726. 24:30You look at the sessions, you look at
  727. 24:31the memory store, you update the memory
  728. 24:32store. How do you know those are
  729. 24:33correct? Um
  730. 24:36so, evaluations obviously are one way to
  731. 24:37do it. Um this is kind of a fun
  732. 24:40anecdotal example from Pokémon showing
  733. 24:42that like you can do perform corrections
  734. 24:43via dreaming.
  735. 24:45The key point is that we've actually run
  736. 24:46a lot of different evals showing that
  737. 24:47dreaming can indeed improve performance
  738. 24:50for very intuitive reasons as you see
  739. 24:52here. But of course, evals are important
  740. 24:54in like your own context to confirm it's
  741. 24:55actually worth the offline compute.
  742. 24:58Yep. I guess we're done. Thank you all.
  743. 25:16>> [music]

About this transcript

This page contains the full transcript of Claude for Long-Horizon Tasks — Lance Martin, Anthropic by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 4,469 words across 743 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.