YouTube2Text

Fable 5 And GPT-5.6 Don't Need Better Prompts. They Need A Clean Setup. — Transcript

by AI News & Strategy Daily | Nate B Jones · 3,224 words · 460 segments · language en · Watch on YouTube

Full transcript

  1. 0:00I overbuilt my harness for Fable 5 and
  2. 0:02Chat GPT 5.6 and I did it before they
  3. 0:05came out. I was so excited. I've been
  4. 0:07using AI a lot. I've been adding new
  5. 0:09rules and new instructions and new
  6. 0:10skills and I've been enthusiastic and
  7. 0:12you know what that means? It means my
  8. 0:14everything around it, all of the skills,
  9. 0:16all of the old system prompts, all of
  10. 0:18that was bloat that kept my AI models
  11. 0:22from performing well. Every time my AI
  12. 0:24missed something, I added another rule.
  13. 0:26A lot of us do, right? I would tell it
  14. 0:27use these sources correctly or write in
  15. 0:29my voice or check the length in this way
  16. 0:31or read this file first. Never make this
  17. 0:33mistake again. Make no mistakes, right?
  18. 0:35Every rule ended up fixing a real
  19. 0:38problem at the time, but over time, I
  20. 0:40didn't realize how bloated my harness
  21. 0:43had become. If you don't know what a
  22. 0:45harness is, that's part of the problem,
  23. 0:47too, because we haven't been clear about
  24. 0:49that as a community. A harness is
  25. 0:51everything wrapped around the model.
  26. 0:54So, it's your custom instructions. It's
  27. 0:56your project files. It's your saved
  28. 0:58prompts. It's your memory. Your your
  29. 1:00skills, your tools.
  30. 1:02It's your permissions, any checks that
  31. 1:04you run, right? It shapes the answer
  32. 1:06before you type anything into the prompt
  33. 1:08window because it comes with that prompt
  34. 1:10and helps the model understand what to
  35. 1:12respond with. Most of us never set out
  36. 1:14to build a harness. We built it
  37. 1:16accidentally, one correction at a time.
  38. 1:18I had no screen that showed me the full
  39. 1:20extent of the harness I built in Codex
  40. 1:22and Claude. I could see the latest rule
  41. 1:24I added if I looked for it, but I
  42. 1:26couldn't see the whole system in one
  43. 1:27place.
  44. 1:29When the AI started behaving strangely,
  45. 1:31I would often blame the model and I know
  46. 1:33a lot of people do and then I would add
  47. 1:34another rule and that would make it even
  48. 1:36more bloated, right? So, I built a skill
  49. 1:38that makes the harness visible. I'll
  50. 1:40show you six principles for building a
  51. 1:42more stable harness, including what
  52. 1:45works differently for Fable 5 and the
  53. 1:47Chat GPT 5.6 family of models. Then
  54. 1:50we'll show you how you can run the same
  55. 1:52cleaner on your setup. The goal is very
  56. 1:54simple. The models keep evolving, right?
  57. 1:57And that's not going to change. Now,
  58. 1:58when I ran the first inventory on my
  59. 2:00setup, it found 66 reusable skills and
  60. 2:04172 instruction-related files. I've been
  61. 2:06very busy. One normal writing job could
  62. 2:09pull in an 18,000-word file before it
  63. 2:12even adjusted the prompt. But, the
  64. 2:14number wasn't the real problem. Some of
  65. 2:16those instructions actually protect
  66. 2:17really important work. They tell the AI
  67. 2:20which sources matter when I'm doing
  68. 2:21research. They stop it from inventing an
  69. 2:24opinion that I genuinely don't have. I
  70. 2:26don't like it when it does that. They
  71. 2:28define what it can do without asking me
  72. 2:30and what it needs to ask me about.
  73. 2:32Problem was, I couldn't tell with the
  74. 2:34new intelligence, the new models, which
  75. 2:37rules were protecting the work and which
  76. 2:39were just copies or overlaps or arriving
  77. 2:42too early or unnecessary with today's
  78. 2:44models or or overlapping in a way that
  79. 2:46was confusing to the model. It's like
  80. 2:47everything around the engine in the car
  81. 2:49all the way to what makes the wheels go,
  82. 2:50right? The drive shaft is an example.
  83. 2:52The the full chassis. All of the parts
  84. 2:54of a car that are not the engine that
  85. 2:56are required to transfer force from the
  86. 3:00engine to the wheels. That's what a
  87. 3:02harness is for an AI model. It actually
  88. 3:04makes the work possible and we should
  89. 3:06definitely design it intentionally and
  90. 3:08not just guess and throw bolts into the
  91. 3:10car. The thing is, when the model
  92. 3:12changes, the harness doesn't magically
  93. 3:14rebuild itself. Sometimes you choose the
  94. 3:16new model. Sometimes ChatGPT or Claude
  95. 3:18will actually change your default for
  96. 3:20you and retire an older model and it
  97. 3:22will route a a difficult request
  98. 3:24somewhere new. And sometimes it will
  99. 3:25even switch models in the middle of the
  100. 3:27conversation.
  101. 3:29You may not even know it happened,
  102. 3:31right? And the old setup, the old
  103. 3:33harness, the old chassis for that car
  104. 3:35stays in place. The new model may behave
  105. 3:38very, very differently as a result and
  106. 3:40you may wonder why. You may blame the
  107. 3:41model, right? The experience could get
  108. 3:43worse. So, you add another instruction.
  109. 3:45I pointed at the setup I already use and
  110. 3:47I ask it to clean the harness that it
  111. 3:49can see. I don't gather every prompt by
  112. 3:51hand. I don't decide what should be
  113. 3:52deleted before it starts. The skill
  114. 3:55begins by making a map of my harness.
  115. 3:57That leads to the first rule, right? You
  116. 3:59map the harness before you clean it.
  117. 4:01It's actually a principle, right? The
  118. 4:02skill does it, but it's because it's a
  119. 4:03good idea. The map gives every important
  120. 4:06control that you have in your harness a
  121. 4:08separate row and asks, "Where does this
  122. 4:10control live? When does it load? What
  123. 4:13job does it do? Who owns this? Is there
  124. 4:16any evidence that it still helps? What
  125. 4:18problem can it create if it's misused?"
  126. 4:21And this was the first time I saw my
  127. 4:24whole harness in one place. And I don't
  128. 4:25know about you, but that's It was really
  129. 4:27illuminating how much junk there was
  130. 4:29there. It also exposed a difference that
  131. 4:32chat bots normally hides. Other controls
  132. 4:34are actual locks. They have teeth to
  133. 4:36them, right? A permission can block an
  134. 4:38action, a schema can reject a broken
  135. 4:41JSON file, a task can refuse a bad
  136. 4:44result if it fails. Those are not the
  137. 4:46same kind of instruction, right? Until
  138. 4:49you map the harness, it can all look
  139. 4:51like text to you, but it doesn't read
  140. 4:53the same way to an AI. The second
  141. 4:55principle that I learned as I built this
  142. 4:57is to blame the right layer.
  143. 4:59Uh I tested the same underlying job with
  144. 5:01Fable 5 in two different setups. The
  145. 5:04compact setup gave Fable the goal, the
  146. 5:07facts, the permission boundary that I
  147. 5:09had for it, and the finish line. The
  148. 5:11thicker setup, the thicker skill, gave
  149. 5:14it all of that plus the full method,
  150. 5:16plus a scoring system, an eval plan, and
  151. 5:19classification scheme. And the thicker
  152. 5:21version did do the job more completely.
  153. 5:23The analysis that I got was much richer,
  154. 5:26but it also failed actual delivery
  155. 5:28requirements twice.
  156. 5:31One result broke the JSON, another broke
  157. 5:33the word limit. The compact setup
  158. 5:35finished correctly three times out of
  159. 5:37three. Like all three times it finished
  160. 5:39right. That doesn't mean that short
  161. 5:41prompts always win. I want to be clear
  162. 5:42about that. It means that the model and
  163. 5:44the harness produce the work that we
  164. 5:47care about together.
  165. 5:49If you blame the model for everything,
  166. 5:51you will keep adding instructions to
  167. 5:53solve problems created by instructions,
  168. 5:57not by the model. And so, the cleaner
  169. 5:59skill force is a much more useful
  170. 6:00question for all of us. Did the model
  171. 6:02fail or did the surrounding setup fail?
  172. 6:04Right? Did the harness fail? Let's jump
  173. 6:05to the third. The third rule is one
  174. 6:08rule, one home, one owner. Right? It's
  175. 6:11easy to remember. My setup had versions
  176. 6:14of the same authorship and source rule
  177. 6:17in 15 different top-level skills,
  178. 6:19indicating I cared about it a lot, but I
  179. 6:21was not getting traction or progress in
  180. 6:23any of the individual skills, so I kept
  181. 6:24working on it. For me, the critical
  182. 6:27thing that I care about is that I don't
  183. 6:29want AI putting words in my mouth. I
  184. 6:31don't want it to come back with research
  185. 6:32and pretend it's my opinion. And so, I
  186. 6:35have 15 different versions of a skill
  187. 6:37that stops AI from reading a pile of
  188. 6:39research and coming back in my voice,
  189. 6:41because I really, really care about
  190. 6:43understanding exactly what was
  191. 6:45researched and getting the citation
  192. 6:47correct. So, I want to keep that
  193. 6:48quality, but I sure don't need 15
  194. 6:51different files pretending to do that
  195. 6:53once, and clearly none of them quite did
  196. 6:55it right because I kept adding to it.
  197. 6:56Every copy is another place where the
  198. 6:58rule can drift. One version gets fixed
  199. 7:00after a failure, uh 14 don't, now it's
  200. 7:03out of sync, now the model has several
  201. 7:05versions of the truth. You see where
  202. 7:06this is going. This is just bad. The
  203. 7:08cleaner doesn't ask whether an
  204. 7:09instruction is too long. It's smarter
  205. 7:11than that. It asks what job the
  206. 7:14instruction does, where that job should
  207. 7:17live, and who should update it the next
  208. 7:20time something changes. A lot of our
  209. 7:22theme along the way here has been about
  210. 7:24making sure we load the right
  211. 7:26information at the right time. And the
  212. 7:27fourth rule really codes that clearly.
  213. 7:29It says we should load specialist
  214. 7:31knowledge when the work actually needs
  215. 7:33it. I had six different editorial guides
  216. 7:37that were loading whenever one writing
  217. 7:39skill ran. So, for example, the source
  218. 7:40guide matters a lot when I'm trying to
  219. 7:42do the research. And if I'm looking for
  220. 7:44examples for YouTube, I need to be
  221. 7:46finding that at the time when I'm
  222. 7:48wrestling with the YouTube script for
  223. 7:49you guys. All of that is great, but if
  224. 7:52you load all of it at the beginning,
  225. 7:53you're just loading a bunch of crud into
  226. 7:55your AI that is likely to produce worse
  227. 7:57results overall. Your research is worse
  228. 7:59because it's thinking about YouTube
  229. 8:01examples when it really should just be
  230. 8:03thinking about research. So in this
  231. 8:04world, the cleaner skill keeps the
  232. 8:06library, right? It keeps that library of
  233. 8:08specialist skills that's useful. It just
  234. 8:10changes when each part of the library
  235. 8:12appears. And this is where I think a lot
  236. 8:14of cleaning prompts don't really
  237. 8:17fully understand what they're doing. The
  238. 8:19goal is not to throw away any useful
  239. 8:22context. Your library can actually be
  240. 8:24quite large. That's not bad. The fifth
  241. 8:25rule is that if you have hard
  242. 8:27requirements, they need hard checks. And
  243. 8:29so there's a difference between telling
  244. 8:31the model that you have a point of view
  245. 8:32and asking it to wrestle with that point
  246. 8:34of view with you. And there's a simple
  247. 8:35rule here that the cleaner skill will
  248. 8:37check for. If you have those kinds of
  249. 8:39rules, 50 word type rules where it's yes
  250. 8:41or no answers that the model can test
  251. 8:43for, those should be put into a schema
  252. 8:47that the model can test against. And the
  253. 8:48cleaner skill can help with that. The
  254. 8:50key principle here is that we let the
  255. 8:52system enforce the parts that a machine
  256. 8:54can verify, and that makes the harness
  257. 8:56lighter and safer at the same time. Now
  258. 8:58the sixth and final rule is to build for
  259. 9:00the model and the product actually doing
  260. 9:03the work. We want to be sophisticated
  261. 9:04enough that we can actually
  262. 9:06differentiate Fable 5 from Chat GPT 5.6.
  263. 9:09If you use AI a lot, you understand that
  264. 9:10Fable 5 in Claude.ai doesn't have the
  265. 9:12same harness as Fable 5 in Claude Code
  266. 9:15or via the API. Chat GPT 5.6 in Chat GPT
  267. 9:19Work isn't the same as in Codex and
  268. 9:21doesn't expose the same controls as it
  269. 9:23does via the API. So the model matters,
  270. 9:25but the product around it determines how
  271. 9:27skills load, which tools exist, what can
  272. 9:30be checked, and what proof comes back.
  273. 9:32What facts does the model need? What is
  274. 9:34it allowed to do? What must be true
  275. 9:37before the work is finished? What's our
  276. 9:38eval effect? And those core rules don't
  277. 9:40change just because I switched from
  278. 9:42Claude to chat GPT cuz the value on the
  279. 9:45work is the same, right? We're testing
  280. 9:46the same work across both harnesses.
  281. 9:48That setup passed every single delivery
  282. 9:50requirement three times out of three,
  283. 9:52right? Like it actually delivered what
  284. 9:54it needed to because it was suited to
  285. 9:56Fable. And as I discussed the richer
  286. 9:57skills, the heavier harness, they
  287. 9:59produced richer analysis, but Fable was
  288. 10:02unable to meet my output constraints.
  289. 10:04And so there were issues that Fable had
  290. 10:05just because it was struggling with the
  291. 10:07thicker skill. And so what does this
  292. 10:08teach us about how Fable works, right?
  293. 10:10You give it the real outcome, you give
  294. 10:11it context that it can't infer, you give
  295. 10:14it room to inspect the problem, and you
  296. 10:16let it plan its own approach usefully.
  297. 10:19And then you bring in specialist
  298. 10:20material when the work reaches that
  299. 10:23phase, which the cleaner skill can help
  300. 10:25you so that you invoke it at the right
  301. 10:26time, or you can let Fable invoke it at
  302. 10:28the right time. And what failed was
  303. 10:30narrating the entire method in that big
  304. 10:32thick skill file before Fable had even
  305. 10:35seen the job, and expecting all of that
  306. 10:37prompt prose to enforce like JSON and
  307. 10:40word count requirements. That's just not
  308. 10:41going to work really well, even if
  309. 10:43Fable's very good at rule following. And
  310. 10:46this is not about starving Fable of
  311. 10:47context, by the way. The full method
  312. 10:49helped it to notice more stuff, but the
  313. 10:52depth should arrive when the work needs
  314. 10:54it, not all up front in a way that
  315. 10:56confuses the model, right? So here we
  316. 10:58have the audit results. 66 skill roots
  317. 11:00and 172 instruction assets. 27,000
  318. 11:04description characters against 8,000
  319. 11:06Codex Discovery budget. That is a big
  320. 11:10problem. It's a big problem because it
  321. 11:11means Codex can't read it. There was an
  322. 11:1318,000
  323. 11:15word content route or skill with a
  324. 11:18minimum fan in. Basically, it was a
  325. 11:19skill leading into skills that 18,000
  326. 11:21words.
  327. 11:23Uh provenance governance, how you handle
  328. 11:25sources distributed across 15 different
  329. 11:27skills, that's what I talk about. And
  330. 11:29only six of the 66 root skills had a
  331. 11:32detected local eval, right? I had evals
  332. 11:36on six of them, but I didn't have them
  333. 11:37on the other 60, and I needed to improve
  334. 11:41that. And it's all about consistency
  335. 11:43with Codex, right? Schemas, tool
  336. 11:44restrictions, file checks, and a run
  337. 11:46receipt, they all carry the exact same
  338. 11:48requirements from skill to skill so that
  339. 11:50Codex is able to operate very
  340. 11:52consistently. So, the Codex chat GPT
  341. 11:54cleanup is not simply make the prompt
  342. 11:56shorter. It's make the right route to
  343. 11:58the skill easy to find, then load the
  344. 12:01depth at the right point. And the
  345. 12:03receipt that you have, which is
  346. 12:04important for Codex, it records the
  347. 12:06model, the reasoning setting, the tools,
  348. 12:08the skills, the fallbacks, and the
  349. 12:09checks that ran, so that you can get a
  350. 12:11sense over time of where there are
  351. 12:13problems. It's like a diagnostic for
  352. 12:14your engine, right? You can figure out
  353. 12:16what's going on. If we compare the two
  354. 12:17models together, Fable's failure mode
  355. 12:19showed up after the method became too
  356. 12:22heavy for the delivery job, whereas chat
  357. 12:24GPT 5.6's failure mode in Codex showed
  358. 12:27up much earlier, while the system was
  359. 12:29still trying to find the right method at
  360. 12:31all, and it was just having trouble
  361. 12:32routing across a really huge harness
  362. 12:33layer.
  363. 12:34Both models benefit from selective
  364. 12:37loading, right? Both benefit from
  365. 12:38including those hard checks I talked
  366. 12:40about, but for very different immediate
  367. 12:42reasons. And that has to do with how
  368. 12:45these models work. Like Fable 5 sorting
  369. 12:47through a lot of context and trying to
  370. 12:49figure out how to do the best job and
  371. 12:51kind of overloading itself is a Fable 5
  372. 12:53way to fail. Harnesses grow barnacles
  373. 12:56like ships. Ultimately, what I want is a
  374. 12:58system that helps us work well. I want
  375. 13:00the harnesses to evolve and stay clean
  376. 13:02and not just collect these barnacles of
  377. 13:04extra point solutions, extra additions,
  378. 13:06extra text files, extra skills that I
  379. 13:09randomly find on the internet because I
  380. 13:10found a tweet somewhere, add it in, and
  381. 13:13now I don't know what I'm doing because,
  382. 13:15you know, chat GPT is overwhelmed or
  383. 13:17Fable 5 gets too fat a skill, or
  384. 13:18whatever the failure mode may be. I want
  385. 13:20a clean harness that lets me get work
  386. 13:22done in a way that's efficient. That's
  387. 13:24why I built this cleaner skill. And this
  388. 13:26matters, right? If you're a product
  389. 13:27manager, it means that you have like a
  390. 13:29clean, simple note that loads the right
  391. 13:31PRD skills at the right time. If If
  392. 13:33you're a developer, it means having one
  393. 13:35source of truth instead of several
  394. 13:36instruction files arguing about how the
  395. 13:38repository work. But we need to make
  396. 13:41sure that they add into the system we
  397. 13:42have in a way that's efficient. And if
  398. 13:43you're just using ChatGPT or Claude at
  399. 13:45home or at work, it means that your old
  400. 13:48memories, your project files, your
  401. 13:49examples, your corrections stop quietly
  402. 13:52shaping answers in ways you don't
  403. 13:54necessarily intend. You may have given
  404. 13:56Claude or ChatGPT a correction 6 months
  405. 14:00ago that it's remembering and over
  406. 14:02applying now that you have a new model.
  407. 14:04So, for example, you and may have told
  408. 14:06it a few months ago, "You have to show
  409. 14:08me the step-by-step because back in
  410. 14:10November that was really important." But
  411. 14:12now you don't need to do that cuz the
  412. 14:13models are good, right? And so, those
  413. 14:15are the kinds of things the cleaner
  414. 14:16scope catches. Ultimately, the goal is
  415. 14:18simple. Your AI should become easier to
  416. 14:20use to do useful work because the system
  417. 14:23around it should become easier to
  418. 14:24understand. Once you can see that
  419. 14:27system, you can actually own it.
  420. 14:29You can tell which controls protect you,
  421. 14:31which need one single home, right? Which
  422. 14:34should arrive later in the process,
  423. 14:36which should become real locks instead
  424. 14:38of just polite reminders. I put the
  425. 14:40complete cleaner over on the Substack,
  426. 14:42you can find it. You can install it
  427. 14:43once, you can point it at the AI setup
  428. 14:45or project of your choosing, and you can
  429. 14:48just let it show everything that's
  430. 14:49accumulated. It gives you a map of the
  431. 14:52harness and the cleaning decisions. It's
  432. 14:54going to give you a plain English before
  433. 14:56and after and a receipt showing what ran
  434. 14:58when you allow it to run and and and run
  435. 15:00those corrections. You can review the
  436. 15:02changes before anything important moves.
  437. 15:04I don't want this to be something that
  438. 15:05surprises you in any way. I have a
  439. 15:07feeling that a lot of us are going to
  440. 15:08discover that we built a lot more than
  441. 15:09we realized. I know I certainly did, and
  442. 15:11so I wanted to showcase that here and
  443. 15:13show you all how bloated my harness had
  444. 15:15become and how much I needed to clean. I
  445. 15:18hope this has been useful for you. I'm
  446. 15:20going to keep drilling in on these new
  447. 15:21models. I'll keep updating the cleaner
  448. 15:23skill. Please let me know in the
  449. 15:24comments one how it works for you, but
  450. 15:26also let me know what other models you'd
  451. 15:29like me to optimize the cleaner skill
  452. 15:30for because of course there are plenty
  453. 15:32of other models besides Claude and chat
  454. 15:33GPT and I'm happy to start working on
  455. 15:36that as well. Would love to kind of
  456. 15:38expand this and make this more useful
  457. 15:39for the community over time and I look
  458. 15:41forward to showing you what I've built
  459. 15:43with my new and improved harness on
  460. 15:45Friday.

About this transcript

This page contains the full transcript of Fable 5 And GPT-5.6 Don't Need Better Prompts. They Need A Clean Setup. by AI News & Strategy Daily | Nate B Jones, generated from the public captions YouTube serves with the video. The transcript has 3,224 words across 460 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.