YouTube2Text

The prompting playbook — Transcript

by Claude · 4,733 words · 779 segments · language en · Watch on YouTube

Full transcript

  1. 0:20Hello everyone. Um, thank you so much
  2. 0:23for joining me this afternoon in the
  3. 0:25breakout room. The last session today of
  4. 0:28Code with Claude. I hope you all had a
  5. 0:30fantastic day so far. My name is Margo
  6. 0:32van Laar. I am an applied AI engineer at
  7. 0:35Anthropic here in London. And this
  8. 0:38afternoon we're going to be talking
  9. 0:40about the prompting playbook. And
  10. 0:43prompting is arguably one of the first
  11. 0:45skills, if not the first skill, that we
  12. 0:48had to learn as engineers when we first
  13. 0:50started to work with LLMs. And even now
  14. 0:54it continues to be one of the most
  15. 0:56critical, um, skills to building
  16. 0:59effective AI systems.
  17. 1:01So, today we're going to discuss some
  18. 1:03best practices,
  19. 1:05um, in the context of
  20. 1:08two practical scenarios that you're
  21. 1:09probably encountering at work. The first
  22. 1:12is where you have an existing prompt in
  23. 1:15production that you've been maintaining
  24. 1:17for some time, um,
  25. 1:20and possibly you're migrating it to a
  26. 1:22new model or making a change to the
  27. 1:24architecture. And for some reason it's
  28. 1:26no longer working as well.
  29. 1:28The second scenario is where we're
  30. 1:29building an entirely new agentic use
  31. 1:31case from the ground up, and we need to
  32. 1:33build the prompt from zero to one.
  33. 1:36Now,
  34. 1:38in order to illustrate these best
  35. 1:40practices, I don't just want to give you
  36. 1:43a list of do's and don'ts. I want to
  37. 1:45walk through a practical example that's
  38. 1:47been inspired by real prompts that, um,
  39. 1:51I've seen some of our customers work
  40. 1:53with who are building on Claude.
  41. 1:56So,
  42. 1:57the prompt that we'll look at today is a
  43. 1:59miniaturized example. The prompts that
  44. 2:02you're working with are probably a lot
  45. 2:04longer and more complex than the one
  46. 2:06we'll see today. Um but it's
  47. 2:08representative of some common problems
  48. 2:11that you might encounter when
  49. 2:13maintaining a prompt.
  50. 2:14So, imagine that we have a prompt that
  51. 2:18multiple people have been collaborating
  52. 2:20on, contributing to. There's no clear
  53. 2:23owner. It covers a lot of different
  54. 2:25areas like policy, like tone, processes.
  55. 2:30Um we have some patches
  56. 2:32for kind of previous models that we've
  57. 2:34migrated to all mixed together.
  58. 2:37Um it's built up and it's complex.
  59. 2:40And when we're migrating to a new model,
  60. 2:42we're finding that suddenly a lot of our
  61. 2:45test cases are no longer working as well
  62. 2:48as we expected. So, what's actually
  63. 2:50going on here? Well, in order to start
  64. 2:53unpacking that question, um we need a
  65. 2:55starting point, and that starting point
  66. 2:58is evaluations. We need evaluations to
  67. 3:01provide that rigor um to understand
  68. 3:05whether a change to our prompt is
  69. 3:07actually correlating to an improvement
  70. 3:09in its performance.
  71. 3:11And we have different models which have
  72. 3:14different capabilities and different
  73. 3:16behaviors. And when you migrate to a
  74. 3:20different model, it could be that your
  75. 3:22system is no longer working as well for
  76. 3:24two reasons. First of all, if uh
  77. 3:28the new model might be capable, but it's
  78. 3:30behaving differently. And therefore, we
  79. 3:32can tune our prompting to fix that
  80. 3:34behavior.
  81. 3:36The second case is where actually the
  82. 3:38model that we're changing to isn't as
  83. 3:40capable, and no amount of prompting is
  84. 3:43going to fix that. So, we need to have
  85. 3:45an eval suite to act as a way of testing
  86. 3:50that regression so that we can apply our
  87. 3:52prompting best practices to that.
  88. 3:56So, in the example that we're going to
  89. 3:58be looking at today, as I said, it's
  90. 4:01going to be a miniaturized example.
  91. 4:02We'll have five test cases in our eval.
  92. 4:07In reality, you'll have a lot more test
  93. 4:09cases in your eval suite, but the key
  94. 4:12thing here is that it's representative
  95. 4:14of three key cases that we need to
  96. 4:16cover.
  97. 4:17The three key cases include having a
  98. 4:20control case, which is a case which
  99. 4:22should always pass. It's something that
  100. 4:24them we know the model handles well.
  101. 4:26It's unambiguous.
  102. 4:28The second is
  103. 4:30edge cases, and these are cases where
  104. 4:33we've seen the model fail before. And by
  105. 4:36including instructions into the prompt,
  106. 4:39we're making sure that same behavior
  107. 4:41doesn't slip through again in the
  108. 4:43future.
  109. 4:44And finally and critically, we need to
  110. 4:46make sure that the model has a good
  111. 4:48understanding
  112. 4:50of the extent of its capabilities, where
  113. 4:53it should be handing off to a human, or
  114. 4:55where actually maybe it should be
  115. 4:57point-blank refusing to answer a
  116. 5:00request.
  117. 5:02So, in the example that we're going to
  118. 5:04be looking at today, um we'll be using a
  119. 5:08um prompt for a customer support bot for
  120. 5:11a Telco company called Meridian Mobile.
  121. 5:14And these are the five test cases that
  122. 5:17we are going to be looking at today. We
  123. 5:19have a simple control case looking at
  124. 5:24you know, what's the data limit in the
  125. 5:25basic plan.
  126. 5:26Uh we're also looking at
  127. 5:29edge cases such as its ability to do
  128. 5:32calculations, such as calculating
  129. 5:34proration bills. If I switch my bill
  130. 5:37halfway through the month, uh or if I
  131. 5:39switch my plan halfway through the
  132. 5:40month, what will my bill look like?
  133. 5:43We want to check that it's accurately
  134. 5:45addressing key questions which are
  135. 5:47covered by our policy.
  136. 5:49Um we need to make sure that it's
  137. 5:51escalating to a human whenever there is
  138. 5:54um a billing error.
  139. 5:57Um and finally, we want to make sure
  140. 5:58that our model isn't withholding any
  141. 6:01information that it has access to which
  142. 6:03it should be handing over to the
  143. 6:05customer.
  144. 6:07So, what we're going to do in this
  145. 6:08process is we'll take our prompt
  146. 6:11and we'll run it on our V0
  147. 6:14um
  148. 6:15of of the eval. And we'll see what our
  149. 6:18failure modes are and systematically
  150. 6:20target those failure modes one at a time
  151. 6:22to see if we can resolve those failure
  152. 6:24modes by prompting. And along the way,
  153. 6:26we'll learn a little bit more about the
  154. 6:27kind of antipatterns
  155. 6:29um and traps to avoid.
  156. 6:32And this is representative of how we
  157. 6:35would apply these best prompting
  158. 6:37techniques in practice, right? We are
  159. 6:39rarely writing a prompt from scratch.
  160. 6:41We're often debugging an existing
  161. 6:44prompt.
  162. 6:46And best practice before we start
  163. 6:48targeting those failure modes
  164. 6:50specifically is to kind of apply our
  165. 6:52general prompting 101 best practices
  166. 6:55applying general hygiene to clean up
  167. 6:58before we do the eval run.
  168. 7:01So, let's have a little look at the
  169. 7:03example that we're going to be using.
  170. 7:05So, what we're looking at here first of
  171. 7:07all before we look at the prompt is just
  172. 7:09this five-coded web app that I've made
  173. 7:11for the presentation today so that we
  174. 7:13can look at how we're iterating on the
  175. 7:15prompt together um in this page here. I
  176. 7:18can easily run my evals on all five test
  177. 7:23cases and inspect the results in a
  178. 7:25little bit more detail.
  179. 7:27So, before we have a look at the prompt,
  180. 7:30I'm just going to run the evals in the
  181. 7:31background.
  182. 7:34This is a pretty good first pass at a
  183. 7:37prompt. When we look at this, we've
  184. 7:40defined the bot's role at the top.
  185. 7:44When we scroll down, we've given it some
  186. 7:47data.
  187. 7:49We've given it some information on how
  188. 7:51to reason over
  189. 7:53the answers that it should be giving to
  190. 7:55the customer. It's giving some critical
  191. 7:57instructions around the tone it should
  192. 8:00use,
  193. 8:01how to do calculations, etc. And then
  194. 8:04finally, we're passing in our customer
  195. 8:06account context and our user message.
  196. 8:10So, let's have a look at how our first
  197. 8:13pass at the evals did. So, we can see as
  198. 8:15we expect our control case, all of our
  199. 8:18test cases have passed. This is what we
  200. 8:21expect for this unambiguous test case.
  201. 8:23But, it's performing pretty poorly in
  202. 8:27these other areas.
  203. 8:29Now, before we zoom in on those specific
  204. 8:32failure modes here, let's do some
  205. 8:34general clean up of our prompt.
  206. 8:43So, as we mentioned, when we look
  207. 8:44through this prompt, there's a couple
  208. 8:47oddities here already. So, for example,
  209. 8:50first one is we're telling the bot that
  210. 8:52it's
  211. 8:53a human, which just isn't true, right?
  212. 8:56We can see as we scroll down, there's
  213. 8:57clearly some information here that has
  214. 8:59been copied directly from a website. So,
  215. 9:02the key giveaway here is a reference to
  216. 9:04a hero image. There's even some
  217. 9:06references to cookies
  218. 9:08at the bottom.
  219. 9:10So, we need to remove a bit of redundant
  220. 9:12information.
  221. 9:15When we look at the instructions here,
  222. 9:17they're all grouped into one big
  223. 9:19paragraph. So, we've got some reasoning
  224. 9:21here, we've got instructions about the
  225. 9:22role, some critical instructions as
  226. 9:26well, without a real way of unpacking
  227. 9:30policy from guidelines, from tone, etc.
  228. 9:34So,
  229. 9:37let me just
  230. 9:41I preempted some changes we want to make
  231. 9:43to this prompt, and this is just a diff
  232. 9:44view of some of those changes. So, what
  233. 9:47we've done is first of all added some
  234. 9:49structure. So, you can see that we've
  235. 9:51added XML tags here to define the role,
  236. 9:54to separate general guidelines,
  237. 9:58to separate policy,
  238. 10:00to separate tone of voice,
  239. 10:03um etc.
  240. 10:08So, if we run that eval then
  241. 10:11on this new updated prompt, we should
  242. 10:13hopefully see an improvement in the
  243. 10:16output as is.
  244. 10:27So,
  245. 10:28we can see just by clearing up the
  246. 10:29prompt, we've already improved the
  247. 10:31model's performance on this prepaid
  248. 10:34scenario.
  249. 10:36There's an interesting regression there
  250. 10:37in that fifth hotspot case, and I don't
  251. 10:40want to worry too much about that now.
  252. 10:42There's going to be some natural level
  253. 10:44of variance in the different runs of the
  254. 10:47eval, and we'll come back to that case
  255. 10:49specifically to see if we can make the
  256. 10:51prompt consistently better in that area.
  257. 10:54So,
  258. 10:55what did we learn from this then? Um
  259. 10:58simply clearing up the prompt
  260. 11:00with a better structure, with a better
  261. 11:02role description has improved the
  262. 11:04performance, and this is a best practice
  263. 11:06that you can return to at any stage of
  264. 11:09writing and maintaining your prompt,
  265. 11:10especially as your prompts get more
  266. 11:12detailed and more complex.
  267. 11:14A general rule of thumb that I like to
  268. 11:16follow is if you're reading a prompt and
  269. 11:19you can't tell guidelines from policy
  270. 11:22from data, most likely the model isn't
  271. 11:24able to either.
  272. 11:27So,
  273. 11:28before looking at some of those cases in
  274. 11:30more detail, there's a little bit more
  275. 11:32general cleanup we can do. Um Um
  276. 11:34specifically here looking at creating an
  277. 11:37output contract. This is a key best
  278. 11:40practice to follow if you're struggling
  279. 11:43with your output format consistency.
  280. 11:45Now, in this case we have a customer
  281. 11:47support bot. We want it to reply in a
  282. 11:49conversational tone. So, it's unlikely
  283. 11:51to be a big issue in this case, but it's
  284. 11:54something to bear in mind if you're
  285. 11:56you're dealing with more complex output
  286. 11:57structures like nested JSONs, for
  287. 12:00example.
  288. 12:04So, again, if we go back to the prompts
  289. 12:07and see what fixes we can apply here.
  290. 12:10First of all,
  291. 12:12we've added a section
  292. 12:14um at the end where we've defined an
  293. 12:15output format for the model telling it
  294. 12:18to use um XML tags to output the
  295. 12:22response. But, the prompt is not always
  296. 12:26the most effective way of handling
  297. 12:29issues. We can also change things in the
  298. 12:31harness to ensure consistency to a
  299. 12:34higher degree.
  300. 12:35So, what we've added here
  301. 12:37to the API call is a stop sequence,
  302. 12:40which is going to detect that closing
  303. 12:43XML tag and tell the model to stop
  304. 12:46generating a response at that point.
  305. 12:49Now, when I run the eval here, I don't
  306. 12:52necessarily expect to see any clear
  307. 12:55improvement in performance,
  308. 12:57um but it's a general best practice that
  309. 12:59we should be following and as I said, is
  310. 13:00something that we should remember in
  311. 13:02particular when we have more complex
  312. 13:04output schemas.
  313. 13:09One thing to point out here as well, if
  314. 13:11you do have a more complex output
  315. 13:13schema, something like structured
  316. 13:14outputs can be incredibly helpful to
  317. 13:16ensure that consistency in a more
  318. 13:18programmatic way.
  319. 13:21Okay, so after the cleanup then, we can
  320. 13:24see that we now have two test cases
  321. 13:27which are consistently passing, but we
  322. 13:29have three key failure modes, the
  323. 13:31proration, the billing error, and the
  324. 13:33hotspot. So,
  325. 13:35let's isolate these one by one
  326. 13:38uh to iterate on the prompt and and see
  327. 13:40the effect of that.
  328. 13:50First of all, then, the hotspot
  329. 13:51question. So, the question is how much
  330. 13:53hotspot data is on my unlimited plan?
  331. 13:57What we expect the model to do is state
  332. 13:59directly the amount of hotspot data that
  333. 14:02the customer has.
  334. 14:04And the reason this is a slightly
  335. 14:06complex case is because the customer
  336. 14:08test case that we're dealing with is on
  337. 14:10a legacy plan. So, actually, the current
  338. 14:12policy doesn't apply to them. So, if we
  339. 14:16see what's going on in the actual test
  340. 14:18case here,
  341. 14:20the customer data which we are feeding
  342. 14:23uh
  343. 14:23to the prompt includes the amount of
  344. 14:26hotspot data that customer has. They
  345. 14:28have 5 GB, right? But, they also have a
  346. 14:30grandfathered plan.
  347. 14:32So, what
  348. 14:34we're seeing uh the model is actually
  349. 14:37telling the customer is
  350. 14:40the general uh
  351. 14:41the unlimited plan includes 4 GB, um but
  352. 14:45since you're on a legacy plan, you
  353. 14:46should go check this out yourself.
  354. 14:55So, let's have a look at the prompt then
  355. 14:57to see
  356. 14:58um why the model
  357. 15:01is deflecting this question to the
  358. 15:05customer account URL rather than
  359. 15:06actually giving the information itself.
  360. 15:10Now, if we read this prompt, originally,
  361. 15:12it said, "We changed our plans recently,
  362. 15:14and the policy doc shows the current
  363. 15:15plan data, and customers on
  364. 15:17grandfathered plan have different rates.
  365. 15:19Never give a customer the wrong plan
  366. 15:22details. Instead, point them to the URL.
  367. 15:25So,
  368. 15:27it's clear that this instruction, this
  369. 15:29latter one, never give customer the
  370. 15:31wrong information, is the instruction
  371. 15:33that the bot has been optimizing for.
  372. 15:36And
  373. 15:37you might recognize this as being very
  374. 15:40similar to a patch that you might have
  375. 15:42introduced in a previous model that you
  376. 15:45were using to avoid where the model was
  377. 15:48giving the customer the wrong
  378. 15:49information about that plan. Now, as our
  379. 15:52models have evolved, they've gotten much
  380. 15:55better at instruction following. So,
  381. 15:56it's likely that instructions like these
  382. 15:58have now become redundant and are
  383. 16:00actually being overfitted to.
  384. 16:06So, what we're going to tell the model
  385. 16:07instead is give this balanced view uh um
  386. 16:11where it says, you know, customers on
  387. 16:13grandfather's plan have different
  388. 16:14allowances, but it's captured in the
  389. 16:16customer information that's given, and
  390. 16:18that is the accurate source of truth.
  391. 16:26So, running the eval here, we should
  392. 16:28hopefully be addressing uh um all of the
  393. 16:31test cases for the hotspot case. Now, I
  394. 16:34am running this live, so there could be
  395. 16:36some variability here, but we see here
  396. 16:38that now clearly all of our test cases
  397. 16:40our are passing.
  398. 16:44So, what did we learn from this? Well,
  399. 16:46we worry a lot about hallucinations
  400. 16:49or the invention of facts and numbers,
  401. 16:53but actually the opposite can also
  402. 16:55happen.
  403. 16:57The model can withhold information that
  404. 16:59it actually has access to.
  405. 17:02Now, we saw here that this is likely a
  406. 17:03result of a patch that we introduced for
  407. 17:05a previous model, and a best practice
  408. 17:08that we could follow here is actually
  409. 17:10using version control. Where wherever we
  410. 17:13are making defensive changes in the
  411. 17:16prompt,
  412. 17:17We are tracking the reason why we've
  413. 17:20introduced these. Sometimes they're
  414. 17:22necessary, but in the future, these kind
  415. 17:25of changes can produce unwanted effects,
  416. 17:27so that we can backtrack on them.
  417. 17:31So,
  418. 17:32the next failing test case then is this
  419. 17:34proration calculation where a customer
  420. 17:38asks,
  421. 17:39"What if I upgrade to the 30 GB plan?
  422. 17:42What will my next bill be?" And what we
  423. 17:44want the model to do is to perform some
  424. 17:46calculation and return exactly
  425. 17:50what the next bill would be, rather than
  426. 17:53giving some sort of vague output, which
  427. 17:55is what we can see it's doing right now.
  428. 17:59If we look at what the model is
  429. 18:01returning,
  430. 18:02it's clearly reasoning through it. It's
  431. 18:05doing a little bit of mental math here
  432. 18:07and there, but it's not really giving
  433. 18:08the customer a concrete answer, and I
  434. 18:10wouldn't rely on this as being able to
  435. 18:12accurately give the customer a response.
  436. 18:20So,
  437. 18:21if we look at the prompt then to see how
  438. 18:23we can fix this.
  439. 18:29In the original prompt,
  440. 18:33we can see that all the instructions
  441. 18:35that were given to it is telling it,
  442. 18:37"Don't ever give a customer a vague
  443. 18:39answer.
  444. 18:41Critical. Always calculate any prorated
  445. 18:44amounts correctly."
  446. 18:46Now, telling the model to do a good job
  447. 18:49isn't particularly helpful when we don't
  448. 18:52give the model the capability to
  449. 18:54actually do a good job.
  450. 18:56We want to avoid the model doing mental
  451. 18:58math. So, what we're going to introduce
  452. 19:00is give the model a tool. So, we're
  453. 19:01saying in the prompt, "Whenever you're
  454. 19:03doing any calculations, please use the
  455. 19:06calculate proration tool to do so."
  456. 19:10In order to introduce that tool, we need
  457. 19:12to introduce it into the API to tell the
  458. 19:16model you have access to this tool. We
  459. 19:18need to define
  460. 19:21the tool schema, which tells the model
  461. 19:23what this tool does and when to use it.
  462. 19:26And then finally, we need to actually
  463. 19:28implement the tool, which is the maths
  464. 19:30behind how it should be doing that
  465. 19:32calculation.
  466. 19:38So, running that eval then for another
  467. 19:40pass,
  468. 19:48we can see that all the test cases are
  469. 19:50now passing. It's clearly done
  470. 19:53the maths using the tool in the
  471. 19:55background and returning the correct
  472. 19:57response.
  473. 20:00So, the key lesson to take away here is
  474. 20:03instructions don't add capability.
  475. 20:06Telling the model it's critical to do a
  476. 20:08calculation right doesn't make it better
  477. 20:12at mental maths. So, the correct
  478. 20:13approach was to give it a tool. Overall,
  479. 20:16giving it the ability to reason over
  480. 20:19harder problems and using tools to
  481. 20:22actually execute them reliably.
  482. 20:25So, now we have one final failing test
  483. 20:28case, which we need to address, which is
  484. 20:30the billing error here.
  485. 20:34In this scenario, there is a billing
  486. 20:36conflict, and what we really want is the
  487. 20:39agent to escalate this to a human. And
  488. 20:44what we're seeing it doing instead is
  489. 20:46it's trying to
  490. 20:48explain to the customer what the reason
  491. 20:51behind it might be.
  492. 20:54And it's trying to kind of diagnose the
  493. 20:56problem itself.
  494. 21:04So, in order to
  495. 21:06fix this behavior, let's again have a
  496. 21:09look what it was told in the prompt.
  497. 21:15We see in the initial instructions it
  498. 21:17was giving it says avoid escalating or
  499. 21:19transferring to a care specialist unless
  500. 21:22absolutely necessary as it cost
  501. 21:24approximately $8 and accounts against
  502. 21:27our team's fast contract resolution.
  503. 21:29Now, this is only giving one side of the
  504. 21:32story, right? We're telling it what the
  505. 21:34cost is to escalating the benefits which
  506. 21:38means it's going to overfit again to not
  507. 21:41escalating the scenario. And second of
  508. 21:44all, we've got this clear conflict
  509. 21:46between what we've defined
  510. 21:49in the eval in terms of what we want the
  511. 21:51model to do to do this escalation versus
  512. 21:54what we're actually telling it to do.
  513. 21:56And the fix that's relevant here is to
  514. 21:58give it both sides of the story by
  515. 22:00saying it costs $8
  516. 22:04to escalate the case, but actually if
  517. 22:07you get this wrong
  518. 22:10then
  519. 22:11it's going to cost you a refund as well
  520. 22:14as customer trust.
  521. 22:16Again here we observed how the model
  522. 22:20optimizes for a goal and this kind of
  523. 22:22instruction is a common instruction to
  524. 22:24give it's quite similar to the one we
  525. 22:27saw earlier where we didn't want it to
  526. 22:29overfit to a certain type of behavior,
  527. 22:31but it's the kind of instruction that
  528. 22:33can be followed quite differently by
  529. 22:35different generations of models and
  530. 22:38specifically as models become more
  531. 22:40intelligent, we need to remember to
  532. 22:42state both sides of the trade-offs
  533. 22:44because our models are
  534. 22:46becoming better themselves at making
  535. 22:49those trade-offs themselves.
  536. 22:57So,
  537. 22:59if we just go back to our eval then and
  538. 23:02uh um
  539. 23:03run our final test case.
  540. 23:06We should see that all of our evals are
  541. 23:09now passing correctly.
  542. 23:12So, overall, we looked at applying
  543. 23:14general hygiene principles, how that can
  544. 23:16provide an initial uplift to the prompt,
  545. 23:18making sure we're removing any redundant
  546. 23:21instructions which were initially
  547. 23:23intended as patches for previous models
  548. 23:26behavior, making sure we're giving it
  549. 23:28tools to do certain tasks reliably.
  550. 23:34Now, there's one other scenario that we
  551. 23:38uh introduced at the start, which is one
  552. 23:39that you might also encounter in your
  553. 23:42work, which is where we're building a
  554. 23:44new agent from scratch. And the example
  555. 23:47that we'll look at here is um an agent
  556. 23:51whose purpose it is to create a
  557. 23:53week-long retail staff schedule based on
  558. 23:56employee availability and other
  559. 23:58constraints.
  560. 24:00And when we're building
  561. 24:02a new agent from scratch,
  562. 24:05we need to consider not just the prompt,
  563. 24:08but also the model that we're using and
  564. 24:10the harness that we're using. So,
  565. 24:13in this next example, we're going to
  566. 24:15compare a number of approaches to
  567. 24:17explore the impact of those three
  568. 24:19different areas.
  569. 24:24So, again, I've just five coded up this
  570. 24:27web app so that we can walk through this
  571. 24:29problem um
  572. 24:31in this demo. Here, I've just laid out
  573. 24:33what the problem is that we're
  574. 24:34addressing. We have our eight employees.
  575. 24:37Um on the right, we have this schedule
  576. 24:40that we need to staff with the head
  577. 24:41count. And we have our constraints that
  578. 24:45must be satisfied in every scenario.
  579. 24:50Now, because we have these hard rules,
  580. 24:52rather than using an LLM judge like we
  581. 24:55did in the previous case to do the
  582. 24:57grading, we can actually use a just a
  583. 25:01Python function which programmatically
  584. 25:03checks for every schedule that's
  585. 25:05generated, how many violations were
  586. 25:08made.
  587. 25:13So, to begin with,
  588. 25:15in a
  589. 25:16we want to start simple. We're going to
  590. 25:18use a simple prompt. We're going to use
  591. 25:20the bare bones that we think we'll need
  592. 25:21with a model Sonnet 4.6 to see how it
  593. 25:24performs and how we're going to hill
  594. 25:26climb against that.
  595. 25:28Um so, here is our baseline prompt.
  596. 25:30We've already applied some of that
  597. 25:32general hygiene and those best practices
  598. 25:35that we saw earlier on using XML tags to
  599. 25:37structure the prompt. We've given it an
  600. 25:39output format as well now that we're
  601. 25:41giving a schedule. Uh we're asking it to
  602. 25:45output a JSON which,
  603. 25:47if we don't
  604. 25:49give that output structure, might lead
  605. 25:51to parsing errors uh downstream.
  606. 25:59When we run the
  607. 26:02simple model on a first iteration of the
  608. 26:05evals, all cases fail. Now, just what
  609. 26:09we're looking at here is in our test
  610. 26:11set, we're essentially repeating
  611. 26:14uh um
  612. 26:15we're doing five trials here uh and
  613. 26:18these numbers are showing how many
  614. 26:19violations were made in each trial.
  615. 26:25In the output, we can see that it's made
  616. 26:28a decent attempt at reasoning through
  617. 26:31the problem,
  618. 26:33but
  619. 26:34it's burning a lot of tokens uh and
  620. 26:38clearly not checking its work uh um as
  621. 26:40it's not getting to the right impact.
  622. 26:43So,
  623. 26:44let's try a larger model, uh a model
  624. 26:47which we know is better at reasoning.
  625. 26:49So, we're going to run it through
  626. 26:52uh Opus 4.7 instead.
  627. 26:57Keeping everything else the same. Now,
  628. 27:00interestingly, whilst all test cases are
  629. 27:03still failing, you can see that the
  630. 27:05overall number of violations that Opus
  631. 27:08has made has reduced significantly from
  632. 27:12Sonnet 4.6. So,
  633. 27:15we're possibly onto something here,
  634. 27:17right? This isn't good enough to ship
  635. 27:20because it's still failing, but clearly
  636. 27:22giving it more reasoning capability is
  637. 27:24helping drive it towards a better
  638. 27:26result.
  639. 27:27So, what we're going to try next is
  640. 27:30using Opus with adaptive thinking
  641. 27:33instead. So, it can decide for itself
  642. 27:38how much thinking it needs how much
  643. 27:40reasoning it needs to use to solve this
  644. 27:44issue. So, no change to the prompt,
  645. 27:46really just uh a change to the API here.
  646. 27:54So, this now seems to reliably generate
  647. 27:58compliance schedules, but it requires a
  648. 28:01lot more tokens.
  649. 28:04We're tripling essentially in the number
  650. 28:06of tokens that we're using here, and
  651. 28:08we're tripling the latency.
  652. 28:10So, we want to try and see if we can
  653. 28:12optimize that cost latency trade-off a
  654. 28:14little bit more. Um
  655. 28:16this is latency 100 seconds. Obviously,
  656. 28:19I'm running this one async uh for the
  657. 28:22purposes of time. Opus 4.7 hasn't
  658. 28:24magically gotten much faster since the
  659. 28:26last time uh you used it. Um
  660. 28:31So,
  661. 28:33let's see if we can optimize a little
  662. 28:35bit more for all the token latency
  663. 28:37trade-off.
  664. 28:39What we haven't tried yet is using
  665. 28:41Sonnet 4.6, so a smaller model, but with
  666. 28:43a better prompt. We looked a lot at the
  667. 28:45prompt optimization uh um in that last
  668. 28:48section. So, I've added a couple details
  669. 28:51to the prompt here. We're We're
  670. 28:54in particular how to reason through this
  671. 28:58problem. And most critically telling it
  672. 29:01to check its work before outputting it.
  673. 29:10So, when I ran that eval, we see that it
  674. 29:13passes in two out of the five cases.
  675. 29:17Now, the failure modes that we're seeing
  676. 29:20is actually not violations
  677. 29:23of
  678. 29:24the scheduling requirements, but the
  679. 29:27model hasn't been able to finish the
  680. 29:30tasks within the output limit that we
  681. 29:34set. So,
  682. 29:36whilst we could increase the max tokens
  683. 29:39that this model is able to use to get
  684. 29:42all five test cases passing,
  685. 29:45we see here that we're using even more
  686. 29:47tokens, and this run has an even higher
  687. 29:49latency. So, this is probably not the
  688. 29:51route that we want to go down.
  689. 29:54Now, as a final pass then, we want to
  690. 29:58look at doing this a little bit more
  691. 30:01agentically.
  692. 30:03So, we're going to use this generate
  693. 30:05evaluate repair loop, where essentially
  694. 30:09the generator now creates a first draft
  695. 30:12of the schedule. And then we have a
  696. 30:14separate prompt,
  697. 30:16which reports any specific violations
  698. 30:19that it made. So, not programmatically
  699. 30:21checking it, but checking it with an
  700. 30:23LLM.
  701. 30:24So, we're checking for every rule, and
  702. 30:27we're providing evidence of every
  703. 30:28violation.
  704. 30:30And we then have a third
  705. 30:33repair prompt, which
  706. 30:36receives any violations that were made,
  707. 30:39and tries to make targeted fixes to it.
  708. 30:43So,
  709. 30:45we have three very simple prompts, but
  710. 30:46they're now running independently rather
  711. 30:49than trying to do everything in one
  712. 30:51large prompt.
  713. 30:57So, we can see in this case our
  714. 31:01agentic approach has solved all of our
  715. 31:04test cases
  716. 31:06uh with
  717. 31:07a much lower number of tokens and with a
  718. 31:10lower latency than trying Sonnet 4 6
  719. 31:14with a better prompt.
  720. 31:16So, going forward, it seems like there's
  721. 31:17two
  722. 31:19appropriate approaches to take here.
  723. 31:21Using Opus 4 7 with adaptive thinking or
  724. 31:24using this agentic loop. Now, moving
  725. 31:27forward, we'd probably want to do a
  726. 31:29little bit more optimization on this
  727. 31:32loop to try and get it to be more
  728. 31:34efficient, but there's one key benefit
  729. 31:36as well from using this generate
  730. 31:38evaluate repair loop. And that is that
  731. 31:41you can put in soft requirements at
  732. 31:44runtime. So, in the evaluation prompt,
  733. 31:47we can say Harry doesn't like working
  734. 31:49with Sally. So, as much as possible, try
  735. 31:52and separate them from working together.
  736. 31:54Or, we need a third shift uh um on
  737. 31:57Wednesday, for example. So, it means
  738. 31:59that you're not having to make changes
  739. 32:01to the um Python function, which is
  740. 32:04doing the evaluation in the back end
  741. 32:05every time to satisfy for any soft
  742. 32:08constraints, which might depend just on
  743. 32:11a case-by-case basis.
  744. 32:16So, to wrap up then,
  745. 32:19pulling all of those learnings together,
  746. 32:21what did we see? Well, we looked at two
  747. 32:23scenarios, two scenarios which I as an
  748. 32:25engineer see
  749. 32:27most in my day-to-day, which is where
  750. 32:29we're maintaining a prompt, we're
  751. 32:31migrating to a new model which has some
  752. 32:33different behaviors,
  753. 32:35and we're and building a new use case
  754. 32:37from scratch.
  755. 32:39We saw that general hygiene principles,
  756. 32:42following those can immediately uplift
  757. 32:46the performance against
  758. 32:48a set of evals, and that we need those
  759. 32:50evals to be able to rigorously see any
  760. 32:53impacts of changing our prompt on the
  761. 32:56output.
  762. 32:57Then we saw this process of targeting
  763. 32:59failure modes one by one,
  764. 33:02adding structure, avoiding long ban
  765. 33:04lists, etc. were all things that helped
  766. 33:07push our model to the correct behavior.
  767. 33:10And then finally with
  768. 33:13our new agentic bot that we were
  769. 33:15building, we saw the impact of splitting
  770. 33:20into three separate prompt systems. So,
  771. 33:22rather than using one
  772. 33:25prompt to address everything, we're
  773. 33:27actually isolating different tasks where
  774. 33:29it's easy and repeatable to separate out
  775. 33:32the steps that it needs to take every
  776. 33:34time.
  777. 33:36Thank you so much for attending this
  778. 33:38afternoon. I hope you have a fantastic
  779. 33:40rest of your day.

About this transcript

This page contains the full transcript of The prompting playbook by Claude, generated from the public captions YouTube serves with the video. The transcript has 4,733 words across 779 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.