The prompting playbook — Transcript
Full transcript
- 0:20Hello everyone. Um, thank you so much
- 0:23for joining me this afternoon in the
- 0:25breakout room. The last session today of
- 0:28Code with Claude. I hope you all had a
- 0:30fantastic day so far. My name is Margo
- 0:32van Laar. I am an applied AI engineer at
- 0:35Anthropic here in London. And this
- 0:38afternoon we're going to be talking
- 0:40about the prompting playbook. And
- 0:43prompting is arguably one of the first
- 0:45skills, if not the first skill, that we
- 0:48had to learn as engineers when we first
- 0:50started to work with LLMs. And even now
- 0:54it continues to be one of the most
- 0:56critical, um, skills to building
- 0:59effective AI systems.
- 1:01So, today we're going to discuss some
- 1:03best practices,
- 1:05um, in the context of
- 1:08two practical scenarios that you're
- 1:09probably encountering at work. The first
- 1:12is where you have an existing prompt in
- 1:15production that you've been maintaining
- 1:17for some time, um,
- 1:20and possibly you're migrating it to a
- 1:22new model or making a change to the
- 1:24architecture. And for some reason it's
- 1:26no longer working as well.
- 1:28The second scenario is where we're
- 1:29building an entirely new agentic use
- 1:31case from the ground up, and we need to
- 1:33build the prompt from zero to one.
- 1:36Now,
- 1:38in order to illustrate these best
- 1:40practices, I don't just want to give you
- 1:43a list of do's and don'ts. I want to
- 1:45walk through a practical example that's
- 1:47been inspired by real prompts that, um,
- 1:51I've seen some of our customers work
- 1:53with who are building on Claude.
- 1:56So,
- 1:57the prompt that we'll look at today is a
- 1:59miniaturized example. The prompts that
- 2:02you're working with are probably a lot
- 2:04longer and more complex than the one
- 2:06we'll see today. Um but it's
- 2:08representative of some common problems
- 2:11that you might encounter when
- 2:13maintaining a prompt.
- 2:14So, imagine that we have a prompt that
- 2:18multiple people have been collaborating
- 2:20on, contributing to. There's no clear
- 2:23owner. It covers a lot of different
- 2:25areas like policy, like tone, processes.
- 2:30Um we have some patches
- 2:32for kind of previous models that we've
- 2:34migrated to all mixed together.
- 2:37Um it's built up and it's complex.
- 2:40And when we're migrating to a new model,
- 2:42we're finding that suddenly a lot of our
- 2:45test cases are no longer working as well
- 2:48as we expected. So, what's actually
- 2:50going on here? Well, in order to start
- 2:53unpacking that question, um we need a
- 2:55starting point, and that starting point
- 2:58is evaluations. We need evaluations to
- 3:01provide that rigor um to understand
- 3:05whether a change to our prompt is
- 3:07actually correlating to an improvement
- 3:09in its performance.
- 3:11And we have different models which have
- 3:14different capabilities and different
- 3:16behaviors. And when you migrate to a
- 3:20different model, it could be that your
- 3:22system is no longer working as well for
- 3:24two reasons. First of all, if uh
- 3:28the new model might be capable, but it's
- 3:30behaving differently. And therefore, we
- 3:32can tune our prompting to fix that
- 3:34behavior.
- 3:36The second case is where actually the
- 3:38model that we're changing to isn't as
- 3:40capable, and no amount of prompting is
- 3:43going to fix that. So, we need to have
- 3:45an eval suite to act as a way of testing
- 3:50that regression so that we can apply our
- 3:52prompting best practices to that.
- 3:56So, in the example that we're going to
- 3:58be looking at today, as I said, it's
- 4:01going to be a miniaturized example.
- 4:02We'll have five test cases in our eval.
- 4:07In reality, you'll have a lot more test
- 4:09cases in your eval suite, but the key
- 4:12thing here is that it's representative
- 4:14of three key cases that we need to
- 4:16cover.
- 4:17The three key cases include having a
- 4:20control case, which is a case which
- 4:22should always pass. It's something that
- 4:24them we know the model handles well.
- 4:26It's unambiguous.
- 4:28The second is
- 4:30edge cases, and these are cases where
- 4:33we've seen the model fail before. And by
- 4:36including instructions into the prompt,
- 4:39we're making sure that same behavior
- 4:41doesn't slip through again in the
- 4:43future.
- 4:44And finally and critically, we need to
- 4:46make sure that the model has a good
- 4:48understanding
- 4:50of the extent of its capabilities, where
- 4:53it should be handing off to a human, or
- 4:55where actually maybe it should be
- 4:57point-blank refusing to answer a
- 5:00request.
- 5:02So, in the example that we're going to
- 5:04be looking at today, um we'll be using a
- 5:08um prompt for a customer support bot for
- 5:11a Telco company called Meridian Mobile.
- 5:14And these are the five test cases that
- 5:17we are going to be looking at today. We
- 5:19have a simple control case looking at
- 5:24you know, what's the data limit in the
- 5:25basic plan.
- 5:26Uh we're also looking at
- 5:29edge cases such as its ability to do
- 5:32calculations, such as calculating
- 5:34proration bills. If I switch my bill
- 5:37halfway through the month, uh or if I
- 5:39switch my plan halfway through the
- 5:40month, what will my bill look like?
- 5:43We want to check that it's accurately
- 5:45addressing key questions which are
- 5:47covered by our policy.
- 5:49Um we need to make sure that it's
- 5:51escalating to a human whenever there is
- 5:54um a billing error.
- 5:57Um and finally, we want to make sure
- 5:58that our model isn't withholding any
- 6:01information that it has access to which
- 6:03it should be handing over to the
- 6:05customer.
- 6:07So, what we're going to do in this
- 6:08process is we'll take our prompt
- 6:11and we'll run it on our V0
- 6:14um
- 6:15of of the eval. And we'll see what our
- 6:18failure modes are and systematically
- 6:20target those failure modes one at a time
- 6:22to see if we can resolve those failure
- 6:24modes by prompting. And along the way,
- 6:26we'll learn a little bit more about the
- 6:27kind of antipatterns
- 6:29um and traps to avoid.
- 6:32And this is representative of how we
- 6:35would apply these best prompting
- 6:37techniques in practice, right? We are
- 6:39rarely writing a prompt from scratch.
- 6:41We're often debugging an existing
- 6:44prompt.
- 6:46And best practice before we start
- 6:48targeting those failure modes
- 6:50specifically is to kind of apply our
- 6:52general prompting 101 best practices
- 6:55applying general hygiene to clean up
- 6:58before we do the eval run.
- 7:01So, let's have a little look at the
- 7:03example that we're going to be using.
- 7:05So, what we're looking at here first of
- 7:07all before we look at the prompt is just
- 7:09this five-coded web app that I've made
- 7:11for the presentation today so that we
- 7:13can look at how we're iterating on the
- 7:15prompt together um in this page here. I
- 7:18can easily run my evals on all five test
- 7:23cases and inspect the results in a
- 7:25little bit more detail.
- 7:27So, before we have a look at the prompt,
- 7:30I'm just going to run the evals in the
- 7:31background.
- 7:34This is a pretty good first pass at a
- 7:37prompt. When we look at this, we've
- 7:40defined the bot's role at the top.
- 7:44When we scroll down, we've given it some
- 7:47data.
- 7:49We've given it some information on how
- 7:51to reason over
- 7:53the answers that it should be giving to
- 7:55the customer. It's giving some critical
- 7:57instructions around the tone it should
- 8:00use,
- 8:01how to do calculations, etc. And then
- 8:04finally, we're passing in our customer
- 8:06account context and our user message.
- 8:10So, let's have a look at how our first
- 8:13pass at the evals did. So, we can see as
- 8:15we expect our control case, all of our
- 8:18test cases have passed. This is what we
- 8:21expect for this unambiguous test case.
- 8:23But, it's performing pretty poorly in
- 8:27these other areas.
- 8:29Now, before we zoom in on those specific
- 8:32failure modes here, let's do some
- 8:34general clean up of our prompt.
- 8:43So, as we mentioned, when we look
- 8:44through this prompt, there's a couple
- 8:47oddities here already. So, for example,
- 8:50first one is we're telling the bot that
- 8:52it's
- 8:53a human, which just isn't true, right?
- 8:56We can see as we scroll down, there's
- 8:57clearly some information here that has
- 8:59been copied directly from a website. So,
- 9:02the key giveaway here is a reference to
- 9:04a hero image. There's even some
- 9:06references to cookies
- 9:08at the bottom.
- 9:10So, we need to remove a bit of redundant
- 9:12information.
- 9:15When we look at the instructions here,
- 9:17they're all grouped into one big
- 9:19paragraph. So, we've got some reasoning
- 9:21here, we've got instructions about the
- 9:22role, some critical instructions as
- 9:26well, without a real way of unpacking
- 9:30policy from guidelines, from tone, etc.
- 9:34So,
- 9:37let me just
- 9:41I preempted some changes we want to make
- 9:43to this prompt, and this is just a diff
- 9:44view of some of those changes. So, what
- 9:47we've done is first of all added some
- 9:49structure. So, you can see that we've
- 9:51added XML tags here to define the role,
- 9:54to separate general guidelines,
- 9:58to separate policy,
- 10:00to separate tone of voice,
- 10:03um etc.
- 10:08So, if we run that eval then
- 10:11on this new updated prompt, we should
- 10:13hopefully see an improvement in the
- 10:16output as is.
- 10:27So,
- 10:28we can see just by clearing up the
- 10:29prompt, we've already improved the
- 10:31model's performance on this prepaid
- 10:34scenario.
- 10:36There's an interesting regression there
- 10:37in that fifth hotspot case, and I don't
- 10:40want to worry too much about that now.
- 10:42There's going to be some natural level
- 10:44of variance in the different runs of the
- 10:47eval, and we'll come back to that case
- 10:49specifically to see if we can make the
- 10:51prompt consistently better in that area.
- 10:54So,
- 10:55what did we learn from this then? Um
- 10:58simply clearing up the prompt
- 11:00with a better structure, with a better
- 11:02role description has improved the
- 11:04performance, and this is a best practice
- 11:06that you can return to at any stage of
- 11:09writing and maintaining your prompt,
- 11:10especially as your prompts get more
- 11:12detailed and more complex.
- 11:14A general rule of thumb that I like to
- 11:16follow is if you're reading a prompt and
- 11:19you can't tell guidelines from policy
- 11:22from data, most likely the model isn't
- 11:24able to either.
- 11:27So,
- 11:28before looking at some of those cases in
- 11:30more detail, there's a little bit more
- 11:32general cleanup we can do. Um Um
- 11:34specifically here looking at creating an
- 11:37output contract. This is a key best
- 11:40practice to follow if you're struggling
- 11:43with your output format consistency.
- 11:45Now, in this case we have a customer
- 11:47support bot. We want it to reply in a
- 11:49conversational tone. So, it's unlikely
- 11:51to be a big issue in this case, but it's
- 11:54something to bear in mind if you're
- 11:56you're dealing with more complex output
- 11:57structures like nested JSONs, for
- 12:00example.
- 12:04So, again, if we go back to the prompts
- 12:07and see what fixes we can apply here.
- 12:10First of all,
- 12:12we've added a section
- 12:14um at the end where we've defined an
- 12:15output format for the model telling it
- 12:18to use um XML tags to output the
- 12:22response. But, the prompt is not always
- 12:26the most effective way of handling
- 12:29issues. We can also change things in the
- 12:31harness to ensure consistency to a
- 12:34higher degree.
- 12:35So, what we've added here
- 12:37to the API call is a stop sequence,
- 12:40which is going to detect that closing
- 12:43XML tag and tell the model to stop
- 12:46generating a response at that point.
- 12:49Now, when I run the eval here, I don't
- 12:52necessarily expect to see any clear
- 12:55improvement in performance,
- 12:57um but it's a general best practice that
- 12:59we should be following and as I said, is
- 13:00something that we should remember in
- 13:02particular when we have more complex
- 13:04output schemas.
- 13:09One thing to point out here as well, if
- 13:11you do have a more complex output
- 13:13schema, something like structured
- 13:14outputs can be incredibly helpful to
- 13:16ensure that consistency in a more
- 13:18programmatic way.
- 13:21Okay, so after the cleanup then, we can
- 13:24see that we now have two test cases
- 13:27which are consistently passing, but we
- 13:29have three key failure modes, the
- 13:31proration, the billing error, and the
- 13:33hotspot. So,
- 13:35let's isolate these one by one
- 13:38uh to iterate on the prompt and and see
- 13:40the effect of that.
- 13:50First of all, then, the hotspot
- 13:51question. So, the question is how much
- 13:53hotspot data is on my unlimited plan?
- 13:57What we expect the model to do is state
- 13:59directly the amount of hotspot data that
- 14:02the customer has.
- 14:04And the reason this is a slightly
- 14:06complex case is because the customer
- 14:08test case that we're dealing with is on
- 14:10a legacy plan. So, actually, the current
- 14:12policy doesn't apply to them. So, if we
- 14:16see what's going on in the actual test
- 14:18case here,
- 14:20the customer data which we are feeding
- 14:23uh
- 14:23to the prompt includes the amount of
- 14:26hotspot data that customer has. They
- 14:28have 5 GB, right? But, they also have a
- 14:30grandfathered plan.
- 14:32So, what
- 14:34we're seeing uh the model is actually
- 14:37telling the customer is
- 14:40the general uh
- 14:41the unlimited plan includes 4 GB, um but
- 14:45since you're on a legacy plan, you
- 14:46should go check this out yourself.
- 14:55So, let's have a look at the prompt then
- 14:57to see
- 14:58um why the model
- 15:01is deflecting this question to the
- 15:05customer account URL rather than
- 15:06actually giving the information itself.
- 15:10Now, if we read this prompt, originally,
- 15:12it said, "We changed our plans recently,
- 15:14and the policy doc shows the current
- 15:15plan data, and customers on
- 15:17grandfathered plan have different rates.
- 15:19Never give a customer the wrong plan
- 15:22details. Instead, point them to the URL.
- 15:25So,
- 15:27it's clear that this instruction, this
- 15:29latter one, never give customer the
- 15:31wrong information, is the instruction
- 15:33that the bot has been optimizing for.
- 15:36And
- 15:37you might recognize this as being very
- 15:40similar to a patch that you might have
- 15:42introduced in a previous model that you
- 15:45were using to avoid where the model was
- 15:48giving the customer the wrong
- 15:49information about that plan. Now, as our
- 15:52models have evolved, they've gotten much
- 15:55better at instruction following. So,
- 15:56it's likely that instructions like these
- 15:58have now become redundant and are
- 16:00actually being overfitted to.
- 16:06So, what we're going to tell the model
- 16:07instead is give this balanced view uh um
- 16:11where it says, you know, customers on
- 16:13grandfather's plan have different
- 16:14allowances, but it's captured in the
- 16:16customer information that's given, and
- 16:18that is the accurate source of truth.
- 16:26So, running the eval here, we should
- 16:28hopefully be addressing uh um all of the
- 16:31test cases for the hotspot case. Now, I
- 16:34am running this live, so there could be
- 16:36some variability here, but we see here
- 16:38that now clearly all of our test cases
- 16:40our are passing.
- 16:44So, what did we learn from this? Well,
- 16:46we worry a lot about hallucinations
- 16:49or the invention of facts and numbers,
- 16:53but actually the opposite can also
- 16:55happen.
- 16:57The model can withhold information that
- 16:59it actually has access to.
- 17:02Now, we saw here that this is likely a
- 17:03result of a patch that we introduced for
- 17:05a previous model, and a best practice
- 17:08that we could follow here is actually
- 17:10using version control. Where wherever we
- 17:13are making defensive changes in the
- 17:16prompt,
- 17:17We are tracking the reason why we've
- 17:20introduced these. Sometimes they're
- 17:22necessary, but in the future, these kind
- 17:25of changes can produce unwanted effects,
- 17:27so that we can backtrack on them.
- 17:31So,
- 17:32the next failing test case then is this
- 17:34proration calculation where a customer
- 17:38asks,
- 17:39"What if I upgrade to the 30 GB plan?
- 17:42What will my next bill be?" And what we
- 17:44want the model to do is to perform some
- 17:46calculation and return exactly
- 17:50what the next bill would be, rather than
- 17:53giving some sort of vague output, which
- 17:55is what we can see it's doing right now.
- 17:59If we look at what the model is
- 18:01returning,
- 18:02it's clearly reasoning through it. It's
- 18:05doing a little bit of mental math here
- 18:07and there, but it's not really giving
- 18:08the customer a concrete answer, and I
- 18:10wouldn't rely on this as being able to
- 18:12accurately give the customer a response.
- 18:20So,
- 18:21if we look at the prompt then to see how
- 18:23we can fix this.
- 18:29In the original prompt,
- 18:33we can see that all the instructions
- 18:35that were given to it is telling it,
- 18:37"Don't ever give a customer a vague
- 18:39answer.
- 18:41Critical. Always calculate any prorated
- 18:44amounts correctly."
- 18:46Now, telling the model to do a good job
- 18:49isn't particularly helpful when we don't
- 18:52give the model the capability to
- 18:54actually do a good job.
- 18:56We want to avoid the model doing mental
- 18:58math. So, what we're going to introduce
- 19:00is give the model a tool. So, we're
- 19:01saying in the prompt, "Whenever you're
- 19:03doing any calculations, please use the
- 19:06calculate proration tool to do so."
- 19:10In order to introduce that tool, we need
- 19:12to introduce it into the API to tell the
- 19:16model you have access to this tool. We
- 19:18need to define
- 19:21the tool schema, which tells the model
- 19:23what this tool does and when to use it.
- 19:26And then finally, we need to actually
- 19:28implement the tool, which is the maths
- 19:30behind how it should be doing that
- 19:32calculation.
- 19:38So, running that eval then for another
- 19:40pass,
- 19:48we can see that all the test cases are
- 19:50now passing. It's clearly done
- 19:53the maths using the tool in the
- 19:55background and returning the correct
- 19:57response.
- 20:00So, the key lesson to take away here is
- 20:03instructions don't add capability.
- 20:06Telling the model it's critical to do a
- 20:08calculation right doesn't make it better
- 20:12at mental maths. So, the correct
- 20:13approach was to give it a tool. Overall,
- 20:16giving it the ability to reason over
- 20:19harder problems and using tools to
- 20:22actually execute them reliably.
- 20:25So, now we have one final failing test
- 20:28case, which we need to address, which is
- 20:30the billing error here.
- 20:34In this scenario, there is a billing
- 20:36conflict, and what we really want is the
- 20:39agent to escalate this to a human. And
- 20:44what we're seeing it doing instead is
- 20:46it's trying to
- 20:48explain to the customer what the reason
- 20:51behind it might be.
- 20:54And it's trying to kind of diagnose the
- 20:56problem itself.
- 21:04So, in order to
- 21:06fix this behavior, let's again have a
- 21:09look what it was told in the prompt.
- 21:15We see in the initial instructions it
- 21:17was giving it says avoid escalating or
- 21:19transferring to a care specialist unless
- 21:22absolutely necessary as it cost
- 21:24approximately $8 and accounts against
- 21:27our team's fast contract resolution.
- 21:29Now, this is only giving one side of the
- 21:32story, right? We're telling it what the
- 21:34cost is to escalating the benefits which
- 21:38means it's going to overfit again to not
- 21:41escalating the scenario. And second of
- 21:44all, we've got this clear conflict
- 21:46between what we've defined
- 21:49in the eval in terms of what we want the
- 21:51model to do to do this escalation versus
- 21:54what we're actually telling it to do.
- 21:56And the fix that's relevant here is to
- 21:58give it both sides of the story by
- 22:00saying it costs $8
- 22:04to escalate the case, but actually if
- 22:07you get this wrong
- 22:10then
- 22:11it's going to cost you a refund as well
- 22:14as customer trust.
- 22:16Again here we observed how the model
- 22:20optimizes for a goal and this kind of
- 22:22instruction is a common instruction to
- 22:24give it's quite similar to the one we
- 22:27saw earlier where we didn't want it to
- 22:29overfit to a certain type of behavior,
- 22:31but it's the kind of instruction that
- 22:33can be followed quite differently by
- 22:35different generations of models and
- 22:38specifically as models become more
- 22:40intelligent, we need to remember to
- 22:42state both sides of the trade-offs
- 22:44because our models are
- 22:46becoming better themselves at making
- 22:49those trade-offs themselves.
- 22:57So,
- 22:59if we just go back to our eval then and
- 23:02uh um
- 23:03run our final test case.
- 23:06We should see that all of our evals are
- 23:09now passing correctly.
- 23:12So, overall, we looked at applying
- 23:14general hygiene principles, how that can
- 23:16provide an initial uplift to the prompt,
- 23:18making sure we're removing any redundant
- 23:21instructions which were initially
- 23:23intended as patches for previous models
- 23:26behavior, making sure we're giving it
- 23:28tools to do certain tasks reliably.
- 23:34Now, there's one other scenario that we
- 23:38uh introduced at the start, which is one
- 23:39that you might also encounter in your
- 23:42work, which is where we're building a
- 23:44new agent from scratch. And the example
- 23:47that we'll look at here is um an agent
- 23:51whose purpose it is to create a
- 23:53week-long retail staff schedule based on
- 23:56employee availability and other
- 23:58constraints.
- 24:00And when we're building
- 24:02a new agent from scratch,
- 24:05we need to consider not just the prompt,
- 24:08but also the model that we're using and
- 24:10the harness that we're using. So,
- 24:13in this next example, we're going to
- 24:15compare a number of approaches to
- 24:17explore the impact of those three
- 24:19different areas.
- 24:24So, again, I've just five coded up this
- 24:27web app so that we can walk through this
- 24:29problem um
- 24:31in this demo. Here, I've just laid out
- 24:33what the problem is that we're
- 24:34addressing. We have our eight employees.
- 24:37Um on the right, we have this schedule
- 24:40that we need to staff with the head
- 24:41count. And we have our constraints that
- 24:45must be satisfied in every scenario.
- 24:50Now, because we have these hard rules,
- 24:52rather than using an LLM judge like we
- 24:55did in the previous case to do the
- 24:57grading, we can actually use a just a
- 25:01Python function which programmatically
- 25:03checks for every schedule that's
- 25:05generated, how many violations were
- 25:08made.
- 25:13So, to begin with,
- 25:15in a
- 25:16we want to start simple. We're going to
- 25:18use a simple prompt. We're going to use
- 25:20the bare bones that we think we'll need
- 25:21with a model Sonnet 4.6 to see how it
- 25:24performs and how we're going to hill
- 25:26climb against that.
- 25:28Um so, here is our baseline prompt.
- 25:30We've already applied some of that
- 25:32general hygiene and those best practices
- 25:35that we saw earlier on using XML tags to
- 25:37structure the prompt. We've given it an
- 25:39output format as well now that we're
- 25:41giving a schedule. Uh we're asking it to
- 25:45output a JSON which,
- 25:47if we don't
- 25:49give that output structure, might lead
- 25:51to parsing errors uh downstream.
- 25:59When we run the
- 26:02simple model on a first iteration of the
- 26:05evals, all cases fail. Now, just what
- 26:09we're looking at here is in our test
- 26:11set, we're essentially repeating
- 26:14uh um
- 26:15we're doing five trials here uh and
- 26:18these numbers are showing how many
- 26:19violations were made in each trial.
- 26:25In the output, we can see that it's made
- 26:28a decent attempt at reasoning through
- 26:31the problem,
- 26:33but
- 26:34it's burning a lot of tokens uh and
- 26:38clearly not checking its work uh um as
- 26:40it's not getting to the right impact.
- 26:43So,
- 26:44let's try a larger model, uh a model
- 26:47which we know is better at reasoning.
- 26:49So, we're going to run it through
- 26:52uh Opus 4.7 instead.
- 26:57Keeping everything else the same. Now,
- 27:00interestingly, whilst all test cases are
- 27:03still failing, you can see that the
- 27:05overall number of violations that Opus
- 27:08has made has reduced significantly from
- 27:12Sonnet 4.6. So,
- 27:15we're possibly onto something here,
- 27:17right? This isn't good enough to ship
- 27:20because it's still failing, but clearly
- 27:22giving it more reasoning capability is
- 27:24helping drive it towards a better
- 27:26result.
- 27:27So, what we're going to try next is
- 27:30using Opus with adaptive thinking
- 27:33instead. So, it can decide for itself
- 27:38how much thinking it needs how much
- 27:40reasoning it needs to use to solve this
- 27:44issue. So, no change to the prompt,
- 27:46really just uh a change to the API here.
- 27:54So, this now seems to reliably generate
- 27:58compliance schedules, but it requires a
- 28:01lot more tokens.
- 28:04We're tripling essentially in the number
- 28:06of tokens that we're using here, and
- 28:08we're tripling the latency.
- 28:10So, we want to try and see if we can
- 28:12optimize that cost latency trade-off a
- 28:14little bit more. Um
- 28:16this is latency 100 seconds. Obviously,
- 28:19I'm running this one async uh for the
- 28:22purposes of time. Opus 4.7 hasn't
- 28:24magically gotten much faster since the
- 28:26last time uh you used it. Um
- 28:31So,
- 28:33let's see if we can optimize a little
- 28:35bit more for all the token latency
- 28:37trade-off.
- 28:39What we haven't tried yet is using
- 28:41Sonnet 4.6, so a smaller model, but with
- 28:43a better prompt. We looked a lot at the
- 28:45prompt optimization uh um in that last
- 28:48section. So, I've added a couple details
- 28:51to the prompt here. We're We're
- 28:54in particular how to reason through this
- 28:58problem. And most critically telling it
- 29:01to check its work before outputting it.
- 29:10So, when I ran that eval, we see that it
- 29:13passes in two out of the five cases.
- 29:17Now, the failure modes that we're seeing
- 29:20is actually not violations
- 29:23of
- 29:24the scheduling requirements, but the
- 29:27model hasn't been able to finish the
- 29:30tasks within the output limit that we
- 29:34set. So,
- 29:36whilst we could increase the max tokens
- 29:39that this model is able to use to get
- 29:42all five test cases passing,
- 29:45we see here that we're using even more
- 29:47tokens, and this run has an even higher
- 29:49latency. So, this is probably not the
- 29:51route that we want to go down.
- 29:54Now, as a final pass then, we want to
- 29:58look at doing this a little bit more
- 30:01agentically.
- 30:03So, we're going to use this generate
- 30:05evaluate repair loop, where essentially
- 30:09the generator now creates a first draft
- 30:12of the schedule. And then we have a
- 30:14separate prompt,
- 30:16which reports any specific violations
- 30:19that it made. So, not programmatically
- 30:21checking it, but checking it with an
- 30:23LLM.
- 30:24So, we're checking for every rule, and
- 30:27we're providing evidence of every
- 30:28violation.
- 30:30And we then have a third
- 30:33repair prompt, which
- 30:36receives any violations that were made,
- 30:39and tries to make targeted fixes to it.
- 30:43So,
- 30:45we have three very simple prompts, but
- 30:46they're now running independently rather
- 30:49than trying to do everything in one
- 30:51large prompt.
- 30:57So, we can see in this case our
- 31:01agentic approach has solved all of our
- 31:04test cases
- 31:06uh with
- 31:07a much lower number of tokens and with a
- 31:10lower latency than trying Sonnet 4 6
- 31:14with a better prompt.
- 31:16So, going forward, it seems like there's
- 31:17two
- 31:19appropriate approaches to take here.
- 31:21Using Opus 4 7 with adaptive thinking or
- 31:24using this agentic loop. Now, moving
- 31:27forward, we'd probably want to do a
- 31:29little bit more optimization on this
- 31:32loop to try and get it to be more
- 31:34efficient, but there's one key benefit
- 31:36as well from using this generate
- 31:38evaluate repair loop. And that is that
- 31:41you can put in soft requirements at
- 31:44runtime. So, in the evaluation prompt,
- 31:47we can say Harry doesn't like working
- 31:49with Sally. So, as much as possible, try
- 31:52and separate them from working together.
- 31:54Or, we need a third shift uh um on
- 31:57Wednesday, for example. So, it means
- 31:59that you're not having to make changes
- 32:01to the um Python function, which is
- 32:04doing the evaluation in the back end
- 32:05every time to satisfy for any soft
- 32:08constraints, which might depend just on
- 32:11a case-by-case basis.
- 32:16So, to wrap up then,
- 32:19pulling all of those learnings together,
- 32:21what did we see? Well, we looked at two
- 32:23scenarios, two scenarios which I as an
- 32:25engineer see
- 32:27most in my day-to-day, which is where
- 32:29we're maintaining a prompt, we're
- 32:31migrating to a new model which has some
- 32:33different behaviors,
- 32:35and we're and building a new use case
- 32:37from scratch.
- 32:39We saw that general hygiene principles,
- 32:42following those can immediately uplift
- 32:46the performance against
- 32:48a set of evals, and that we need those
- 32:50evals to be able to rigorously see any
- 32:53impacts of changing our prompt on the
- 32:56output.
- 32:57Then we saw this process of targeting
- 32:59failure modes one by one,
- 33:02adding structure, avoiding long ban
- 33:04lists, etc. were all things that helped
- 33:07push our model to the correct behavior.
- 33:10And then finally with
- 33:13our new agentic bot that we were
- 33:15building, we saw the impact of splitting
- 33:20into three separate prompt systems. So,
- 33:22rather than using one
- 33:25prompt to address everything, we're
- 33:27actually isolating different tasks where
- 33:29it's easy and repeatable to separate out
- 33:32the steps that it needs to take every
- 33:34time.
- 33:36Thank you so much for attending this
- 33:38afternoon. I hope you have a fantastic
- 33:40rest of your day.
About this transcript
This page contains the full transcript of The prompting playbook by Claude, generated from the public captions YouTube serves with the video. The transcript has 4,733 words across 779 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.