YouTube2Text

Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley — Transcript

by AI Engineer · 2,938 words · 508 segments · language en · Watch on YouTube

Full transcript

  1. 0:01[music]
  2. 0:12>> Okay, I'm getting rolling and uh welcome
  3. 0:14aboard. We just had a little technical
  4. 0:16issues,
  5. 0:17but uh we resolved them. So, my name is
  6. 0:19Frank Coyle.
  7. 0:20Uh I am a computer science guy. I've
  8. 0:23been teaching computer science for over
  9. 0:2530 years,
  10. 0:26and I'm now teaching at Berkeley. And
  11. 0:29one of the problems that uh all my
  12. 0:30students,
  13. 0:32past and present, are having is AI,
  14. 0:34because computer science is no longer
  15. 0:37the magic pathway to a job. So, I've
  16. 0:41been trying to figure out ways to uh
  17. 0:43help them come up with schemes to help
  18. 0:46them get ready for this world of agentic
  19. 0:48AI. And one of the things that sort of
  20. 0:51uh
  21. 0:51dropped into my uh plate was the
  22. 0:55something called the Claude Certified
  23. 0:57Architect exam, which I will be talking
  24. 0:59about today, and it has um a number of
  25. 1:03aspects to it. And I think if you're
  26. 1:04interested in a career in agentic AI,
  27. 1:07then certainly take a look at least what
  28. 1:09the exam is about, because I feel that
  29. 1:12um Anthropic knows how people are using
  30. 1:16their system and what the issues are
  31. 1:18going to be.
  32. 1:19So, before we jump into that, I want to
  33. 1:21give a little bit of my
  34. 1:23uh
  35. 1:23my philosophy.
  36. 1:26bop bop bop bop
  37. 1:33May have to do this manually, getting
  38. 1:34stuck.
  39. 1:36So,
  40. 1:37this is a quote from uh
  41. 1:40a woman named Sister Corita Kent.
  42. 1:42Nothing is a mistake. There's no win and
  43. 1:45no fail. There's only make.
  44. 1:48Bottom line here is experiment,
  45. 1:50experiment, experiment. Not only should
  46. 1:53you read, but you should do. You should
  47. 1:55make stuff. Now, what happens when you
  48. 1:58make stuff? A lot of times things don't
  49. 2:01work.
  50. 2:03Thomas Edison said, "I have not failed.
  51. 2:07I've only found 10,000 ways
  52. 2:09that don't work."
  53. 2:11And
  54. 2:13what I want to emphasize here is that
  55. 2:15what this shows us are something that in
  56. 2:18the design patterns movement, which came
  57. 2:20around in the early 1990s with
  58. 2:22object-oriented programming, we had
  59. 2:24patterns for objects. We now have
  60. 2:27patterns for agents, but there's also
  61. 2:30anti-patterns. And I think anti-patterns
  62. 2:32are a key
  63. 2:34to understanding what you should not do
  64. 2:37because understanding what you should
  65. 2:38not do is the key to leading you to what
  66. 2:41you should do.
  67. 2:44So, a little bit about the Claude
  68. 2:46Certified Exam, released in March, so
  69. 2:49it's brand new.
  70. 2:50It is uh
  71. 2:52it is
  72. 2:53based on scenarios. It is timed. It is
  73. 2:56proctored.
  74. 2:57It is available to companies in the
  75. 3:01Claude ecosystem, the Anthropic
  76. 3:03ecosystem, but individuals can pay $99
  77. 3:06and take the exam once every once every
  78. 3:096 months.
  79. 3:11And it's not just
  80. 3:13multiple-choice questions. It is
  81. 3:15multiple-choice, but they're
  82. 3:17they are based on
  83. 3:19uh realistic constraints and realistic
  84. 3:22scenarios.
  85. 3:24The five domains.
  86. 3:26There are five domains that are covered
  87. 3:28and they give you the percentages of
  88. 3:29each. So, agentic architecture, 27%.
  89. 3:33Claude code, how to configure the Claude
  90. 3:35code system and workflow, 20%. How to
  91. 3:40doing prompt engineering, structuring
  92. 3:42your output, using JSON all over the
  93. 3:46place.
  94. 3:47Tool design. Model context protocol
  95. 3:50integration. These are topics that you
  96. 3:52should understand and know whether
  97. 3:54you're going to take the exam or not.
  98. 3:56This is going to help you get ready for
  99. 3:58whatever
  100. 4:00the agentic world is going to throw at
  101. 4:02you. And then there's going to be
  102. 4:03contact management and reliability. So
  103. 4:06these are the
  104. 4:07areas of of the kind of questions you're
  105. 4:10going to run into.
  106. 4:13Then there are and they they provide you
  107. 4:16with six production scenarios and your
  108. 4:20the exam will randomly choose four and
  109. 4:24all the questions will be centered
  110. 4:26around the four that they choose.
  111. 4:29And what I'm going to do is walk you
  112. 4:31through
  113. 4:32um
  114. 4:34the production scenarios and give you
  115. 4:36some anti-patterns to be aware of
  116. 4:38because there's a number of ways you can
  117. 4:40solve the problem but one of the big
  118. 4:41things is what not to do and that often
  119. 4:44can be the key to getting these
  120. 4:46questions right. So, number one customer
  121. 4:49support resolution agent. So we have
  122. 4:51agentic loops, control, something called
  123. 4:54stop reason which is
  124. 4:56uh what Cloud Code has. Every time
  125. 4:58something happens, there's a stop reason
  126. 5:01and you need to take a look at that
  127. 5:02because that can give you a lot of
  128. 5:03information about what's going on.
  129. 5:05Uh scenario two, code generation.
  130. 5:08Three, multi-agent research system which
  131. 5:11we'll look at. How do you How do you
  132. 5:14distribute your agents? Hub and spoke.
  133. 5:17Who's the orchestrator? How much
  134. 5:18information should they know? All these
  135. 5:21are important factors. Um
  136. 5:23scenario four, developer
  137. 5:26productivity with code. So how do you do
  138. 5:28subtask isolation? Keep your tasks in
  139. 5:31their little universes. And this
  140. 5:33hearkens back to what we learn in
  141. 5:35computer science from doing
  142. 5:36multi-threaded programming.
  143. 5:39When you have multiple threads operating
  144. 5:40and sharing memory, then you get into
  145. 5:43issues with synchronization. You You to
  146. 5:45put locks
  147. 5:47Keep the little threads independent.
  148. 5:50Keep your agents independent.
  149. 5:52Um
  150. 5:54and then some cloud code for continuous
  151. 5:56integration.
  152. 5:58And then we'll look at some patterns for
  153. 6:00structured data extraction. Okay, that's
  154. 6:04kind of where we're going to go.
  155. 6:06Now, here's something that I I I like to
  156. 6:09point out. Everybody's talking about
  157. 6:11loops, right? Every The loop is the new
  158. 6:13thing.
  159. 6:14Um
  160. 6:16uh Boris Cherney says he doesn't write
  161. 6:19code, but his job is to write loops.
  162. 6:22And Peter Steinberger
  163. 6:24master of Open Claw says, "I don't I
  164. 6:26don't uh I don't code anymore. I just
  165. 6:28design loops
  166. 6:30that prompt your agents."
  167. 6:32So, loops are the new big thing, right?
  168. 6:34Well, no, they're not. Okay? Um
  169. 6:38back in the day
  170. 6:40uh early days of computing, we had
  171. 6:43programming languages were exploding. We
  172. 6:45had Fortran, we had COBOL, and there
  173. 6:47were big fights. My program My
  174. 6:50programming language is better than
  175. 6:52yours. It can do more. No, it can't. We
  176. 6:55can do this.
  177. 6:56Böhm and Jacopini, 1966
  178. 6:59proved that if you want a language to be
  179. 7:02Turing complete, which means can compute
  180. 7:05anything that computers are possibly
  181. 7:08able to compute, then you need only
  182. 7:11three things.
  183. 7:13The ability to
  184. 7:14to to write statements sequentially,
  185. 7:17okay?
  186. 7:18To have if-then conditionals, and the
  187. 7:21third piece is the loop.
  188. 7:24If you add the loop,
  189. 7:26you have Turing computability. And now
  190. 7:29we are seeing this being resurrected in
  191. 7:32the agentic world with the focus on
  192. 7:35loops, cuz up to now we've had sort of
  193. 7:37sequences. You have prompts, you have
  194. 7:39maybe if-then, but now we have a loop.
  195. 7:42And now this is what's giving us the
  196. 7:43power. This is where the agentic stuff
  197. 7:46is getting very exciting.
  198. 7:48Okay.
  199. 7:50I'm start with uh
  200. 7:52with scenario one, customer support
  201. 7:54resolution.
  202. 7:56So here we have
  203. 7:59a loop operating and
  204. 8:01the I'm going to jump to the
  205. 8:03anti-pattern. What you don't want is
  206. 8:05just to let the agent go and do
  207. 8:07something and get the response back and
  208. 8:11use it, okay? What you want to do is you
  209. 8:13want to loop with something called the
  210. 8:15stop reason. So I'm going to show you a
  211. 8:17little code here.
  212. 8:19So here we have while loop. It's a while
  213. 8:21true, it's a loop. We're looping right
  214. 8:23here, okay? So the first little block is
  215. 8:26where we call uh we call the model,
  216. 8:29okay? And we pass it the messages. The
  217. 8:31messages are essentially the sequence of
  218. 8:34prompts that exist in the context
  219. 8:37window, okay? And we are asking the and
  220. 8:42we have a we have a prompt and we have
  221. 8:45we have the context and we have a tool.
  222. 8:47And we're asking the LLM
  223. 8:50to do something with this tool and help
  224. 8:52us out. The problem is the LLM can't do
  225. 8:56anything. It is just a probabilistic
  226. 8:59next word predictor.
  227. 9:01It can't execute tools. So what it does
  228. 9:04though is it can figure out
  229. 9:08if you point it to a tool, it can figure
  230. 9:11out how to set things up so that you or
  231. 9:14your code can execute it. So it's
  232. 9:17important to understand that the LLM is
  233. 9:18not executing these tools. It can't do
  234. 9:20anything except talk back to you, very
  235. 9:23intelligently sometimes, but all it can
  236. 9:25do is talk back to you. So
  237. 9:28when it finishes
  238. 9:29this
  239. 9:31task and has a result which is basically
  240. 9:36here is I've I know what you want. I
  241. 9:39know what the tool can do. Here's how I
  242. 9:42It sets up the parameters that can then
  243. 9:45be or that then used to actually execute
  244. 9:48the tool. So, the second block you see
  245. 9:51why did
  246. 9:53the LLM come back to us? That's our stop
  247. 9:56reason.
  248. 9:57Tool use. Oh, okay. We've stopped
  249. 9:59because
  250. 10:00the LLM it wants to use the tool.
  251. 10:03So, let's just run the tool. So, that's
  252. 10:05what the second block is. Run tool, the
  253. 10:08response is what the LLM said, and it's
  254. 10:10basically the parameters that it has
  255. 10:13extracted from the data that you
  256. 10:15provided it.
  257. 10:17Okay? Then it executes that.
  258. 10:19Then it goes back.
  259. 10:20That then it continues. Continues means
  260. 10:23the LLM sees it and says, "Oh,
  261. 10:25successful run. So, okay."
  262. 10:28Come back down.
  263. 10:31We're not running a tool anymore. We're
  264. 10:32end the end of our loop. Bingo.
  265. 10:35Now,
  266. 10:36then we take the answer, and this is an
  267. 10:38opportunity for you to
  268. 10:39have a human in the loop potentially.
  269. 10:43You check the confidence. If it looks
  270. 10:45good, you keep it. If you don't, then
  271. 10:47you escalate to a human.
  272. 10:49So, now there's another reason why you
  273. 10:52need to make sure you check your stop
  274. 10:54reason. One of the stop reasons may be
  275. 10:57you have run out of tokens, and this
  276. 11:00response is based on partial when the
  277. 11:04LLM had to stop.
  278. 11:06And it's going to give you a response,
  279. 11:08but if you have run out of tokens, then
  280. 11:10you need to take action.
  281. 11:12Okay.
  282. 11:13Um
  283. 11:15Next scenario.
  284. 11:17Uh code generation with Claude. So,
  285. 11:19Claude code has this has this concept of
  286. 11:22the Claude MD file, a markdown file,
  287. 11:24where you put all the things you wanted
  288. 11:26to know.
  289. 11:27What Anthropic recommends is you have
  290. 11:31three levels of Claude.
  291. 11:34One
  292. 11:36that you have at the top level of your
  293. 11:37project,
  294. 11:39the other that you have in inside your
  295. 11:41sort of the project folder, and then
  296. 11:45within directories you can also specify.
  297. 11:48So, the idea is to have a hierarchical
  298. 11:50set of rules that that can then control
  299. 11:55how the system is going to respond.
  300. 11:58Okay.
  301. 12:00Moving right along,
  302. 12:02uh we have a multi-agent research
  303. 12:04system. So, here we're going to have uh
  304. 12:08the problem is
  305. 12:10how do I how do I get my agents to to go
  306. 12:12off and do stuff and bring the answers
  307. 12:14back in a reasonable way? The
  308. 12:16anti-pattern
  309. 12:18you
  310. 12:19have one agent and you load it up with
  311. 12:21tools, all right? So, I like to think
  312. 12:23about you
  313. 12:24you know, you hire somebody to come to
  314. 12:25your house, you hire a carpenter to come
  315. 12:27to the house, and the guy shows up with
  316. 12:30uh
  317. 12:31plumbing tools, carpenter tools,
  318. 12:33electrical tools. He says, "I can do
  319. 12:35anything." Well, maybe you don't want
  320. 12:36this guy, maybe you want a a
  321. 12:38professional carpenter. So, that's the
  322. 12:40kind of idea. And this kind of back
  323. 12:42takes us back to some of the the
  324. 12:44functional programming
  325. 12:46uh
  326. 12:47ideas that functions should be do one
  327. 12:50thing. And if you can get your agents to
  328. 12:53do one thing,
  329. 12:55you with maybe one or two tools
  330. 12:58available to it, then that's going to be
  331. 13:01a win, and that's going to help you with
  332. 13:02this exam. So, specialize,
  333. 13:05don't overload.
  334. 13:07The other part of this is
  335. 13:09don't let your agents
  336. 13:11context spill over into the main context
  337. 13:16because context means tokens, tokens
  338. 13:19mean money,
  339. 13:21and the more context you have, the more
  340. 13:23confused the LLM is going to be in
  341. 13:26giving you an answer. So, even though
  342. 13:28oh, a million token context window, I
  343. 13:31can put everything in there. No, no,
  344. 13:32don't put everything in there.
  345. 13:34Limit what's going to go in there
  346. 13:35because then you're going to get
  347. 13:37a much more accurate system.
  348. 13:41So, here's a
  349. 13:44Here's an example of a specialized sub
  350. 13:47agents.
  351. 13:48You're giving it
  352. 13:50So, this would be the critic. So, let's
  353. 13:52say you've run some stuff. Now, you want
  354. 13:54to get an agent to look at what's
  355. 13:57happened. What you want to do is just
  356. 13:59give it what it needs to solve that
  357. 14:02critic problem. I'm only giving it here
  358. 14:05the
  359. 14:07we're passing it
  360. 14:08the claim and the evidence. So, this is
  361. 14:11your claim is sort of how we're going to
  362. 14:13solve the problem. Here's Here's the
  363. 14:14evidence, but we're not giving it the
  364. 14:18the thought processes that went in to
  365. 14:22creating this claim. Why?
  366. 14:25When you
  367. 14:27When you get a bunch of agents together
  368. 14:29collaborating and talking to each other,
  369. 14:32there's a tendency to have group think.
  370. 14:35And
  371. 14:36all the agents seem to kind of devolve
  372. 14:39into one idea. I mean, it's it's like,
  373. 14:42you know, you're in a group, you know,
  374. 14:43you're at a party, and everybody wants
  375. 14:46pizza except you, but then people talk
  376. 14:49you into
  377. 14:50you you know, you don't want to be uh
  378. 14:53you don't want to spoil the party, so
  379. 14:54you'll go along. And it seems that
  380. 14:55agents kind of work in the same way.
  381. 14:58So, you're going to return
  382. 15:00Basically, you're going to give each
  383. 15:02agent only a slice. I didn't think about
  384. 15:05the pizza analogy, but yes. Every agent
  385. 15:08gets its own slice, and and it it should
  386. 15:11come through.
  387. 15:12Okay.
  388. 15:17Fourth scenario,
  389. 15:19developer productivity. So, the
  390. 15:22anti-pattern.
  391. 15:25Let every subtask dump its full output
  392. 15:27into the primary thread, crowding out
  393. 15:29the context. Again, this is what we're I
  394. 15:31was just talking about. This is bad. Let
  395. 15:34the context grow unbounded. Bad, right?
  396. 15:38For the reasons we just talked about.
  397. 15:40You want to isolate your subtask output,
  398. 15:43and you want to compact
  399. 15:46long sessions. I'm going to take a
  400. 15:48second to talk about that. So, here's
  401. 15:51here's a
  402. 15:52an example of a pattern.
  403. 15:54Uh
  404. 15:55you want to have your agent
  405. 15:59uh
  406. 16:00look at the logs and create a summary
  407. 16:04of where the problems are in the log.
  408. 16:06So, here's your task, scan all the logs
  409. 16:09for error.
  410. 16:10Context fork. So, you're forking the
  411. 16:13agent into a like a separate thread
  412. 16:16where
  413. 16:17whatever the agent does and thinks and
  414. 16:20adds tokens to does not come back and
  415. 16:23pollute the main
  416. 16:25uh
  417. 16:26the main context.
  418. 16:28Now,
  419. 16:30you see here what happens, then you take
  420. 16:32this
  421. 16:33summation, and then you add that
  422. 16:35summation without all the other stuff
  423. 16:38into the overriding context. Now, this
  424. 16:42last little block is kind of
  425. 16:43interesting, I think. Because
  426. 16:46you can check your token count,
  427. 16:49and you can determine how big the token
  428. 16:51count is.
  429. 16:53And
  430. 16:55if you can set some limit and you know,
  431. 16:57if if you have more than 150,000 tokens,
  432. 16:59then what you want to do is you can run
  433. 17:01a compact. So, Anthropic and Claude have
  434. 17:04these compaction algorithms
  435. 17:08that take this giant context and and
  436. 17:10compact it in some way, shape, or form.
  437. 17:12Not quite sure how the implementation is
  438. 17:15of that, but there is compaction. Now, a
  439. 17:18little side effect a little side channel
  440. 17:21I've been walking around when you walk
  441. 17:23outside, you see see these guys handing
  442. 17:24out these books.
  443. 17:26Okay? Anybody see these guys handing out
  444. 17:28these but take them. This is this is
  445. 17:30actually a pretty good little book. In
  446. 17:32fact, I was looking at it last night and
  447. 17:35one of the things it had in it was this
  448. 17:37is by this guy Sam
  449. 17:39Sam Bagwell. I have no connection I
  450. 17:41didn't even know Sam, but it there's a
  451. 17:44online page 32.
  452. 17:46It says
  453. 17:47uh his company provides custom logic for
  454. 17:50compression of context. So, he's got an
  455. 17:54and you can write your own. He's got a
  456. 17:56he's got he you can extend his base
  457. 17:57class and have your own
  458. 18:00compression of your data, whatever you
  459. 18:01think is important. So, I think that's
  460. 18:03kind of an interesting spin on this
  461. 18:06whole thing.
  462. 18:07Okay.
  463. 18:09Cloud code for
  464. 18:12uh uh continuous integration
  465. 18:15uh anti-pattern
  466. 18:18Always have interactive modes in a
  467. 18:19pipeline. Well, no no no cuz interactive
  468. 18:22modes mean uh
  469. 18:25Cloud will stop and ask you, "You want
  470. 18:27to do this? You want to do that? Can I
  471. 18:28have permission for that?" So, there are
  472. 18:29ways to set it up so that it'll just run
  473. 18:32straight through, okay?
  474. 18:34The other
  475. 18:36uh
  476. 18:37the other tip that I'll give you here
  477. 18:41is there's something called
  478. 18:43the uh
  479. 18:45the batch. So, you can take your
  480. 18:47prompts, you can take your work, and you
  481. 18:50can put them in a batch and for 50%
  482. 18:54fewer token cost you will get the result
  483. 18:57they promise in at at least 24 hours.
  484. 19:00So, if you're going to go take a nap,
  485. 19:01you're going to go on vacation, you're
  486. 19:03going to go out, take a a day off, run
  487. 19:05your stuff in batch mode, and you're
  488. 19:07going to have a a
  489. 19:09less to pay.
  490. 19:13Where am I here?
  491. 19:15All right, I've only got a few few
  492. 19:17minutes left, few seconds left, but I
  493. 19:20want to conclude with this.
  494. 19:22Remember, nothing is a mistake. There's
  495. 19:25no win, there's no fail, there's no
  496. 19:26exam,
  497. 19:28only make. You do it and you make it and
  498. 19:31you're going to succeed. If you want to
  499. 19:33reach out to me, reach out to me uh coil
  500. 19:35at Berkeley, look at my websites. I got
  501. 19:38a website co-supreme AI. I'm a big jazz
  502. 19:41fan and I named this website after John
  503. 19:43Coltrane, Love Supreme, if you know that
  504. 19:44song, great. Anyway, that's my story and
  505. 19:47I'm sticking to it and I'm about to zero
  506. 19:49time. Okay,
  507. 19:50>> [applause]
  508. 19:51>> thank you.

About this transcript

This page contains the full transcript of Anthropic's CCA Exam as a Field-Guide for Agentic Engineering — Frank Coyle, UC Berkeley by AI Engineer, generated from the public captions YouTube serves with the video. The transcript has 2,938 words across 508 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.