Claude for Long-Horizon Tasks — Lance Martin, Anthropic — Transcript
Full transcript
- 0:01[music]
- 0:12>> Good to go.
- 0:14All right. Well,
- 0:16take a quick sip and then let's start.
- 0:19It is great to be here. Um this is like
- 0:21like my third year coming to this
- 0:22conference and I always really enjoy it.
- 0:25And thank you for coming to this
- 0:26workshop. I know there's many
- 0:27interesting talks.
- 0:29Let me talk a little bit about um
- 0:31our view of async agents at Anthropic
- 0:33and some things we've been up to lately.
- 0:36So, this is kind of a way I think about
- 0:38models in product. So, you can think
- 0:40about Claude as a light source and you
- 0:42can think about products as windows that
- 0:44allow the light to pass through.
- 0:46And what's kind of interesting is over
- 0:48time, the window that you need to
- 0:50actually kind of see the light of the
- 0:51model kind of shifts.
- 0:54And we've seen this over the past few
- 0:56years. So, I'm plotting here different
- 0:58Claude models and their task horizon.
- 1:00So, how much autonomous work can they do
- 1:03over time?
- 1:04And you might recall back in like the
- 1:06Opus 3 days, this was kind of like 2024,
- 1:09models could only do, you know, maybe 10
- 1:11to 20 minutes of autonomous work. This
- 1:13is measured by meter.
- 1:15And in that regime, only certain product
- 1:17surfaces made sense. Things like
- 1:18autocomplete, things like chat, where
- 1:20your human is very in the loop cuz the
- 1:22model's really only doing a very short
- 1:23amount of work before you're steering
- 1:25it.
- 1:26Now, the past year we saw the rise of
- 1:28synchronous coding agents like Claude
- 1:29code. And this is, you know, a kind of a
- 1:31shift because then models could do maybe
- 1:33an hour of work.
- 1:35So, it made sense to have them run, but
- 1:37typically locally, where you could still
- 1:39steer them easily.
- 1:41And it's kind of interesting because
- 1:42during this regime, I remember efforts
- 1:44and I was involved in some efforts to
- 1:45build kind of async agents.
- 1:48But when models can only do like an hour
- 1:50of work, async as an experience is kind
- 1:52of bad. Um the model goes off and it
- 1:55like hits an error and it comes back to
- 1:57you over a short period of time.
- 1:59In order to really unlock async, we
- 2:00needed longer task horizons.
- 2:03And so we're starting to see that now.
- 2:06And kind of with this shift in
- 2:08capability and time horizon came a shift
- 2:10in the API surfaces. So if you look at
- 2:12the lower left,
- 2:15Messages API came out like 2 years ago.
- 2:17It's basically prompt response. It's
- 2:19great for building harnesses,
- 2:22but it's a very simple API. Again,
- 2:24you're just passing in a message and you
- 2:25get a response out.
- 2:27There's no sense of deployment with
- 2:29that. So you basically take Messages API
- 2:30and you can roll your own harness, you
- 2:32can deploy that harness, and you have an
- 2:33agent.
- 2:34Now over the past year, we saw the rise
- 2:37of, you know, coding agents in
- 2:38particular. So we released Agent SDK. So
- 2:40that's basically a way to
- 2:41programmatically call Claude code.
- 2:43And that's like basically us giving you
- 2:45a harness.
- 2:47But over the past few months, since
- 2:49April, as we've seen longer longer task
- 2:51horizons, we released a new API called
- 2:53Managed Agents,
- 2:55which basically packages both the
- 2:57harness as well as all the managed
- 2:58deployment infrastructure for you.
- 3:01And I'll talk about some of the themes
- 3:02that underpin this new uh surface,
- 3:04Claude Managed Agents, and and some of
- 3:06the themes that kind of extend beyond
- 3:08just Managed Agents broadly to think
- 3:10about this kind of new type of
- 3:11asynchronous agents, uh which can apply
- 3:13of course to Claude and other types of
- 3:15kind of longer running long horizon
- 3:17agents.
- 3:19So theme one is decoupling the brain
- 3:21from the hands.
- 3:23Um
- 3:24so when we first set out to build
- 3:26Managed Agents, we started with the
- 3:27container.
- 3:28We put the harness in the sandbox in the
- 3:30same container.
- 3:32Now the problem here is, what happens if
- 3:36the harness dies or the container dies?
- 3:39What we saw is you actually lose the
- 3:40session.
- 3:41So basically, this architecture is kind
- 3:44of tricky for long horizon agents
- 3:46because what can happen is your agent's
- 3:48running, and And container dies, you
- 3:51lose everything with it.
- 3:54Also, as models get more capable,
- 3:56putting the credentials in the same
- 3:58container with the agent itself can be
- 4:01problematic.
- 4:02So, for example, giving Claude access to
- 4:03a bunch of your secrets and letting it
- 4:05run for 10 hours and not watching it can
- 4:07be a little bit spooky and have some
- 4:08security concerns, especially as models
- 4:10get extremely capable.
- 4:13So, for this reason, we kind of decouple
- 4:15what we call the brain, that's the
- 4:16harness, from the hands, the execution
- 4:18environments, and manage agents to set
- 4:19up like this.
- 4:21So, the story here is that the harness
- 4:25becomes a stateless process
- 4:27that talks to a session. The session is
- 4:30an append-only event log
- 4:33and that can reach out to hands, which
- 4:34are just containers.
- 4:36So, that's just Those are sandboxes
- 4:38where work is done.
- 4:39And one thing that's interesting is
- 4:41Claude is increasingly capable of
- 4:42managing many hands.
- 4:44So, that is you can give one harness
- 4:46access to many different containers to
- 4:48perform to execution.
- 4:50And Claude can manage this very easily
- 4:52and and and and effectively.
- 4:55If the session, uh sorry, if the harness
- 4:58dies or sandbox dies, it's completely
- 5:01fine because the session is always
- 5:03backed up in this append-only log and
- 5:05credentials are never actually added to
- 5:07the sandbox.
- 5:08They're stored in a separate vault. So,
- 5:10this decoupling actually makes it quite
- 5:12reliable and safe, particularly for
- 5:14long-horizon tasks. And this is kind of
- 5:15one of the core ideas that underpins
- 5:17manage agents architecturally.
- 5:20And I think an interesting thing that
- 5:21falls out of this is related to you guys
- 5:24may have kind of seen or come across the
- 5:26recursive language models work.
- 5:28The session becomes an external context
- 5:30object that the model interrogates.
- 5:33And this has all sorts of benefits for
- 5:35context management. So, you think about
- 5:37it, when you're doing something like
- 5:38compaction, you're choosing some logic
- 5:40to retain some amount of context, and
- 5:43naively in a typical in a kind of a
- 5:45typical step, you're discarding all the
- 5:46context that you didn't compact.
- 5:49In this architecture, and also more
- 5:50broadly with recursive language models,
- 5:52the idea is that the context object is
- 5:54persistent and is unadulterated. So,
- 5:57it's append-only.
- 5:58And the model can always go back and
- 6:01fetch old context. So, basically creates
- 6:02a very very nice architecture for
- 6:04context engineering because the core
- 6:06context object is immutable in the sense
- 6:08that it's it's it's non-destructive
- 6:11and you can only and you only append to
- 6:13it over time.
- 6:14So, we've seen this be to be quite nice
- 6:16in terms of long horizon context
- 6:17engineering as well.
- 6:20So, the second theme is use verifiers.
- 6:25And one of the problems that we've seen
- 6:28with Claude and other models in general
- 6:31is that when you ask them to do a bunch
- 6:32of work and then say, "Okay, grade your
- 6:35work."
- 6:37If that same context is being used to
- 6:39both do the work and grade, you can get
- 6:41lots of odd artifacts and confabulation
- 6:44and and basically odd behavior.
- 6:47For example, this is just an image
- 6:48showing you can think about that that
- 6:50context window is filled with lots of
- 6:51different information and the model is
- 6:53grading itself, often it's not properly
- 6:56tuned to do kind of critical
- 6:58verification.
- 7:00And so, what we found is it's quite
- 7:02effective to separate verification into
- 7:04a separate context window.
- 7:06This is a very general trend.
- 7:08We talked about it in a number of
- 7:09different engineering blogs.
- 7:11Um and the reason is the verifier
- 7:13context can be tuned very specifically
- 7:15for the critical verification task.
- 7:18>> [snorts]
- 7:18>> And so, the way this works in practice
- 7:20is
- 7:21when you build loops,
- 7:23you can have a loop of a build context
- 7:24and a verifier context. And this can be
- 7:25a build agent verifier agent. And what
- 7:27happens is
- 7:28the verifier has some goal
- 7:31or rubric
- 7:32and it's verifying the result or work of
- 7:35the build agent. And this continues in a
- 7:37loop until verification is complete.
- 7:40And this is really the big idea behind
- 7:41this whole loops trend that you might
- 7:42have heard about and we found it to be a
- 7:44kind of a very powerful paradigm
- 7:45especially for working with some of the
- 7:47higher capacity models.
- 7:50And here are some of the primitives. So
- 7:51in Claude code you have goal and manage
- 7:54agents you have outcomes and the
- 7:55principles are really the same.
- 7:58You're setting up a measurable end state
- 7:59in both cases.
- 8:01You're using an independent context
- 8:02model.
- 8:03You're using independent context to
- 8:05grade
- 8:07over the course of this loop.
- 8:09The loop can run and you only exit the
- 8:11loop until this independent verifier has
- 8:13verified that it has the outcomes or
- 8:14outputs that you want. That's the key
- 8:16idea.
- 8:18Now let me tell you a story about how
- 8:20I've used this. So this is a kind of a
- 8:22fun and interesting challenge called
- 8:23parameter golf. It's a benchmark that
- 8:25set up that was put up by Open AI.
- 8:28And it tests models ability to
- 8:30effectively do kind of ML research.
- 8:33So it asks the model to basically take a
- 8:36small um kind of model and train it in
- 8:40with eight with eight 8100 GPUs in less
- 8:44than 10 minutes. And what you see on the
- 8:46Y is basically you can think about it as
- 8:48loss. So lower is better, okay?
- 8:51And what I did was I set up a kind of a
- 8:53verifier loop using manage agents and
- 8:55outcomes
- 8:57to test the ability for Opus 4.7 and one
- 9:00of our frontier models like mythos class
- 9:02models on this task. And what you see is
- 9:06basically I allow the model to continue
- 9:08to iterate until the outcome that I
- 9:10specify is is satisfied which is it
- 9:13finished exactly 20 iterations and kind
- 9:17of it kind of met all the experimental
- 9:18criteria as defined by the benchmark.
- 9:21What you see is
- 9:23the frontier capability models
- 9:25are extremely good with this pattern of
- 9:27kind of loops in software and and kind
- 9:29of verification because what happens is
- 9:33instead of
- 9:35encoding steering me and into like me as
- 9:38the human, you're encoding the signal
- 9:40into the environment. So, the model can
- 9:42self-correct when it receives feedback
- 9:44from, for example, the verifier.
- 9:47And using this kind of paradigm with
- 9:48very high-capacity models, you can get
- 9:50very strong results.
- 9:52So, the main point I'm trying to make
- 9:54here is that this paradigm of loops,
- 9:56which a lot of people been talking about
- 9:57today,
- 9:59paired with very capacity models is a
- 10:01very good general primitive for
- 10:03long-running asynchronous work.
- 10:05That's really the key point here.
- 10:07>> [snorts]
- 10:09>> Now, let me talk another about another
- 10:11theme of self-learning.
- 10:14So, the human brain has two kind of
- 10:18interesting systems for memory.
- 10:21So, one is as you go about your day, the
- 10:22hippocampus is kind of writing traces of
- 10:25kind of short-term, very fast kind of
- 10:26experiential memory. Like, you might
- 10:28remember what you had for lunch today.
- 10:29You had lunch an hour ago, you kind of
- 10:31remember, that's kind of written to
- 10:32short-term memory.
- 10:33When you go to bed at night, though, an
- 10:35offline process or out-of-band process,
- 10:37dreams. And dreaming stores certain
- 10:41important details to long-term memory in
- 10:42the cortex. So, for example, tomorrow,
- 10:44you might not remember what you ate for
- 10:45lunch today. That's kind of a local
- 10:46trace. But, like, if you had a very
- 10:48important experience today, maybe this
- 10:50talk, maybe not this talk, but if you
- 10:53had an interesting experience today,
- 10:54that might be written to long-term
- 10:55memory. That's kind of the point. These
- 10:56two subsystems in human work like this.
- 10:59And we've actually found memory systems
- 11:00with Claude actually can employ these
- 11:02same two principles.
- 11:04So, this is showing Claude's capacity as
- 11:08an in-band memory writer. So, you
- 11:10basically give Claude memory tools. And
- 11:12when I say memory tools, I mean it's
- 11:13basically the ability to write to a file
- 11:15system, that is basically a memory
- 11:16directory. That's really it.
- 11:19Now, this is showing some work on Claude
- 11:21plays Pokémon with Claude's on it 3.5.
- 11:23And here's the key point. When Sonnet
- 11:263.5 is given this access to a memory
- 11:28directory, and it can write memory,
- 11:30{quote} in-band, as it progresses
- 11:32through this game, it's not very good.
- 11:35So, the memories it writes are pretty
- 11:36crappy. It's kind of um
- 11:38it's kind of uh
- 11:40tactical notes. It's it's not very
- 11:42strategic. And the game progress is
- 11:44quite limited.
- 11:46But with more recent models like this is
- 11:48looking at 4 6,
- 11:50the notes are much more strategic
- 11:52and game progress is much further. So,
- 11:53the key point I'm making here is that
- 11:56Claude has gotten much better at this
- 11:57in-band memory writing across model
- 11:59generations.
- 12:02And this is another way to show that
- 12:04same result.
- 12:05So, this is a benchmark that I ran
- 12:07called Continual Learning Bench. It's an
- 12:08open-source benchmark. I took one of the
- 12:10tasks. This is a task that basically
- 12:11asked the model to perform a sequential
- 12:14question answering with a SQL database
- 12:16and it can write memory in between each
- 12:17step.
- 12:19And what you see is basically the
- 12:21performance improves across models.
- 12:25Um
- 12:26and so what this is kind of showing is
- 12:29that models get natively better at this
- 12:32in-band memory writing with respect to
- 12:34model capability.
- 12:35>> [snorts]
- 12:36>> And some of the most interesting things
- 12:37I found from this are that the main
- 12:39differentiation between There we go. Um
- 12:44the main differentiation
- 12:46between like a lower capacity model and
- 12:48a high capacity model is kind of this
- 12:50distillation step. And so, basically,
- 12:53higher capacity models have a better
- 12:54sense of like what abstraction to save
- 12:57to memory that'll be useful later.
- 13:00Like they're not just writing a specific
- 13:02fact. They're writing how do I how does
- 13:03this generalize to future sessions?
- 13:05That's kind of the key difference that I
- 13:07found that higher capacity models kind
- 13:09of have when they're writing memory. So,
- 13:11this is a very important thing to keep
- 13:12in mind that models are getting better
- 13:13better and better at this kind of
- 13:14in-band memory writing across model
- 13:16generations.
- 13:19Now, there's a little trick here which
- 13:21is very important. So, we talked about
- 13:22kind of in-band memory and we talked
- 13:24about dreaming. So, at night I dream and
- 13:26I write things to like long-term memory.
- 13:28Dreaming is very important because when
- 13:30I'm writing memory in band over the
- 13:32course of a day over the course of a
- 13:33session, sometimes you can write
- 13:35incorrect memories.
- 13:37And
- 13:38or you're writing things that are
- 13:40locally optimal, but not globally
- 13:42optimal. So, you're writing over the
- 13:43course of a task to like kind of help
- 13:47you solve that task, but not necessarily
- 13:49looking forward to future tasks.
- 13:51This is a very important nuance in this
- 13:53process of dreaming is is kind of an
- 13:55offline or out of band process that
- 13:56we've used to consolidate and improve
- 13:58memory.
- 14:00And I want to show you a fun example
- 14:02that I've used dreaming for.
- 14:04So, this is again Pokémon.
- 14:06And this is I played a lot of games of
- 14:09Pokémon with Claude to find this.
- 14:12So, this was a very hard one lesson, so
- 14:14I hope you appreciate it. Okay, here's
- 14:16the point. So, basically what happened
- 14:18is
- 14:20Claude wrote an incorrect memory, okay?
- 14:23And what happened is this incorrect
- 14:25memory was related to the location of
- 14:27The details don't necessarily matter.
- 14:29Uh
- 14:30the point is that this incorrect memory
- 14:33causes Claude to mislocalize itself or
- 14:35Pokémon to mislocalize itself,
- 14:38and it falls through this trapdoor,
- 14:40okay? That's the key point. So, it
- 14:42writes this incorrect memory. This
- 14:43incorrect memory causes it to
- 14:44mislocalize in the game, and it falls
- 14:47down this trap. This is very consistent.
- 14:49So, I saw this in five replicates. Five
- 14:50out of five replicates with raw memory
- 14:52store fell down this trap. With the
- 14:55dreaming,
- 14:56this error is corrected, and it's able
- 14:58to properly localize itself and not fall
- 15:02fall down this trap. And I'll show you
- 15:03kind of a fun visualization of this.
- 15:05This is looking at kind of memory traces
- 15:07or basically traces of game progress.
- 15:09So, going upwards on the Y axis is
- 15:12improvement. That's like moving to the
- 15:13next level. Going down is backtracking.
- 15:16So, what's interesting here,
- 15:18the no memory baseline, which is that
- 15:20gray bar, kind of doesn't make much
- 15:22progress at all. It just kind of like is
- 15:25stuck. It's like a particularly hard
- 15:26level, okay?
- 15:28The memory, which is the orange,
- 15:30actually keeps falling down this
- 15:31trapdoor and falls back. So, it
- 15:33backtracks. The dreaming traces though
- 15:36consistently kind of fix this error in
- 15:38its memory and proceed to the next
- 15:40level. So, this is a very practical
- 15:42example of how dreaming can kind of work
- 15:44out of band on your memory store to fix
- 15:47corrections. Because what it does is it
- 15:49looks at your memory store and it looks
- 15:50at all your prior traces or sessions and
- 15:53kind of can find and correct errors.
- 15:54That's the key point and that's why the
- 15:56dreaming process can be very helpful
- 15:58because in band, while Claude is writing
- 16:00to memory, it can make mistakes. And
- 16:02those mistakes get stuck in memory
- 16:03unless you have an offline process to
- 16:05kind of correct them. That's the key
- 16:06intuition.
- 16:08Um
- 16:09and theme four, and I'll open up for
- 16:11questions after this, is what I think um
- 16:14is kind of this trend that we're going
- 16:15to see moving towards org-level
- 16:17harnesses with async agents.
- 16:19And
- 16:21so, we released Claude Tag
- 16:23and a lot of the reaction was like, "Ah,
- 16:26Slack bot."
- 16:28And like, look, I have actually created
- 16:29a lot of Slack bots myself. I understand
- 16:31not every Slack bot is particularly
- 16:32interesting or great. In fact, I've
- 16:33created many Slack bots that are quite
- 16:35bad.
- 16:36But
- 16:37what's interesting about Claude Tag is
- 16:39not the fact that it's accessible
- 16:41through Slack. What's interesting about
- 16:42it is the fact that it has a very very
- 16:44rich kind of system underneath it, which
- 16:46I want to just cut touch on briefly.
- 16:49And in particular, what's interesting
- 16:50about it is it represents what I
- 16:51consider an org-level harness.
- 16:54So, agents historically have been kind
- 16:55of single player. So, you have an an
- 16:58agent like Claude code on your machine
- 16:59with your local context that you've
- 17:01tuned and configured for yourself.
- 17:03What's interesting about Claude Tag is
- 17:05it is a harness that everyone in the
- 17:06organization has access to and can use.
- 17:09So, it is a multiplayer harness.
- 17:11And what's nice about that is it has its
- 17:13own identity. Its identity and
- 17:14credentials are not tied to a given user
- 17:16and has access to organizational level
- 17:18context, not just my local context.
- 17:20This has many interesting and useful
- 17:23implications.
- 17:24Including the ability to like check
- 17:26others work before you do an experiment,
- 17:28the ability to kind of deduplicate
- 17:30findings, the ability to like do
- 17:31internal research, the ability to give
- 17:34everyone access to kind of a a very well
- 17:36developed harness on day one, whereas
- 17:38when you your own personal harness,
- 17:40often new employees takes them weeks or
- 17:43maybe even months to kind of ramp up
- 17:44fully to configure all the right
- 17:46connectors and so forth. So, org level
- 17:48harnesses are real leveler of the
- 17:50playing field and I think it was
- 17:52people kind of saw the Slack bot piece,
- 17:53but they didn't really appreciate the
- 17:55depth of kind of
- 17:56that the depth of benefit you get from
- 17:58building out org level harnesses. So, I
- 18:00do think that was kind of important
- 18:01thing to note.
- 18:03And I think we're going to see the rise
- 18:05of kind of harnesses that operate across
- 18:07orgs, across many different users that
- 18:09can operate increasingly on longer async
- 18:11async
- 18:13um
- 18:14async agents that can operate on longer
- 18:16time frames. That's kind of one clear
- 18:17kind of follow-up that we're I think
- 18:19we're going to see from this.
- 18:20And another thing where I think we're
- 18:22going to see is that
- 18:24asynchronous agents um
- 18:26are going to be increasingly proactive.
- 18:28Uh so, typically with for example like
- 18:30local locally scoped agents, they tend
- 18:33to be reactive. They're responsive to
- 18:34how you steer it. versus async agents
- 18:37increasingly have the ability to steer
- 18:39proactivity. And that's one very nice
- 18:41thing about Cloud Tag.
- 18:43Um where basically you can configure it
- 18:46to tell you things when like looking at
- 18:49this org level context, alert me with
- 18:51things I might need to know about.
- 18:53Um and this is a very important kind of
- 18:55new kind of UX that I think is going to
- 18:57be more and more common with async
- 18:58agents that kind of access to
- 18:59organizational context.
- 19:01And of course multiplayer. So, the the
- 19:03ability for a single harness to be
- 19:05steered by many many different people
- 19:06kind of concurrently is an important
- 19:08shift in agent UX
- 19:10uh that I think uh will be quite
- 19:12interesting going forward. So,
- 19:15um,
- 19:16yeah, let me let me just open up for
- 19:17questions and um, thank you for
- 19:19listening.
- 19:21>> [applause]
- 19:26[applause]
- 19:30>> Sure.
- 19:31>> Uh, one thing that we see sort of
- 19:33empirically and also in the benchmarks
- 19:35is that the frontier models perform
- 19:37better on these long horizon tasks than
- 19:39on the ones that have stacking. What
- 19:41they need to keep in mind that this is
- 19:43not
- 19:44What is your view on and like the
- 19:47guidance on it and like the right away
- 19:48model?
- 19:49>> Yeah.
- 19:53>> Um, what's your view on why that is and
- 19:56then how long the frontier will be able
- 19:58to maintain that gap?
- 20:00>> I see. So, the question was kind of on
- 20:02the the gap between the frontier models
- 20:04kind of on for example, a benchmark like
- 20:07meter on like long horizon tasks.
- 20:09>> Yeah.
- 20:09>> Okay. Um,
- 20:12so,
- 20:13in the latest results that I saw from
- 20:15like for example, Codex like 56, I think
- 20:17I'd also kind of is in that 12 plus hour
- 20:20regime on meter. So, I think you're
- 20:21right that like the frontier models like
- 20:23the, you know, mythos class models,
- 20:25strong models from Open AI kind of are
- 20:27in this like 12 plus hour regime.
- 20:29Um,
- 20:32why is it that
- 20:34non-frontier models are not kind of in
- 20:36that regime? I am actually not
- 20:37necessarily sure. I do think
- 20:40I do think um,
- 20:42in order to build agents that can
- 20:45effectively operate in this regime, it's
- 20:48important to note that it's not just the
- 20:49model capability. Like for example, with
- 20:51a Claude tag product, it's actually a
- 20:52combination of improvement in memory
- 20:55because memory is very important. If you
- 20:57have agents working for for example, 12
- 20:58hours, you want to make sure that your
- 21:00productivity preferences are well
- 21:02encoded in memory. So, it knows when to
- 21:04reach out to you if it gets stuck for
- 21:05example. So, memory is very important,
- 21:07security is very important, so
- 21:09resistance to prompt injection. Um
- 21:11also like model architecture, like kind
- 21:13of the agent architecture is very
- 21:14important, that decoupling of brain and
- 21:16hand, so it's secure and safe and like
- 21:17resistant to failure. So, actually I
- 21:19think
- 21:21to build real agents that can operate in
- 21:23these long time horizons, a bunch of
- 21:24things need to come together in terms of
- 21:25like architecture,
- 21:27infrastructure, security, memory. And
- 21:31that might be why and and but Frontier
- 21:32Labs invested in all these areas, so
- 21:34that might be why it's become a gap in
- 21:36terms of like the agent products that
- 21:37we've released. And so we spent a lot of
- 21:38time, for example, building managed
- 21:39agents
- 21:41to kind of have these kind of
- 21:42considerations baked in.
- 21:44Yeah.
- 21:45Sure.
- 22:05Yeah, okay, this is interesting.
- 22:07Um the question was about um kind of
- 22:10like the the best memory substrate, so
- 22:12like why file systems versus, for
- 22:13example, databases.
- 22:15Um
- 22:17this is kind of a subtle point that
- 22:18actually um
- 22:21I want to think about carefully. So,
- 22:24I don't necessarily think that
- 22:28it has to be the case that you use a
- 22:29file system for memory. I think what's
- 22:31quite important that we've seen is that
- 22:33it's you want something that is highly
- 22:36programmable
- 22:38with simple primitives that the model
- 22:39can manipulate to like write, manage its
- 22:41own memory. So, for example,
- 22:43a database could work fine relative to
- 22:46the file system.
- 22:47But what I've seen doesn't work is when
- 22:50you specify the structure of memory for
- 22:53the model very explicitly, whether
- 22:55that's in a file system or database or
- 22:56whatever. Like a memory schema, I pre I
- 22:58kind of pre populate, here's the types
- 23:00of memories you need to save.
- 23:02Cuz I think that ends up being not very
- 23:03bitter lesson pill in the sense that
- 23:04models can learn to manage their own
- 23:06memory much better than you can intuit
- 23:08these memory types for the model ahead
- 23:10of time. So, I think what we've seen is
- 23:12that very general substrates for memory,
- 23:15be it just your database or file system,
- 23:17are good because the model can manage
- 23:19them freely versus
- 23:20a very very kind of like prescriptive
- 23:23memory schema they are trying to
- 23:25pigeonhole the model into. That's when
- 23:27you see performance drop. That's the key
- 23:29differentiation.
- 23:35Right.
- 23:37That's the key point. Let the model
- 23:39structure and maintain its own memory.
- 23:40Don't give it a prescribed memory
- 23:42schema.
- 23:43And that's like a common failure because
- 23:45models are getting good enough that they
- 23:46can manage their own memory much more
- 23:47effectively than you can reason about
- 23:49types of This is like very classically
- 23:51bitter lesson pill, but like
- 23:53you can Models can reason about their
- 23:54own memory and context structure much
- 23:56better than you can prescribe for them a
- 23:58way to structure their own memories.
- 24:00That's the key observation.
- 24:02Yep.
- 24:07Yes.
- 24:10That's right. Exactly.
- 24:12General substrates for for memory
- 24:13management.
- 24:15Yep.
- 24:19Yes.
- 24:22Okay. So, that's a good point.
- 24:26Basically, the question is
- 24:28so, you do this dreaming thing.
- 24:30You look at the sessions, you look at
- 24:31the memory store, you update the memory
- 24:32store. How do you know those are
- 24:33correct? Um
- 24:36so, evaluations obviously are one way to
- 24:37do it. Um this is kind of a fun
- 24:40anecdotal example from Pokémon showing
- 24:42that like you can do perform corrections
- 24:43via dreaming.
- 24:45The key point is that we've actually run
- 24:46a lot of different evals showing that
- 24:47dreaming can indeed improve performance
- 24:50for very intuitive reasons as you see
- 24:52here. But of course, evals are important
- 24:54in like your own context to confirm it's
- 24:55actually worth the offline compute.
- 24:58Yep. I guess we're done. Thank you all.
- 25:16>> [music]
About this transcript
This page contains the full transcript of Claude for Long-Horizon Tasks — Lance Martin, Anthropic by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 4,469 words across 743 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.