Fable 5 And GPT-5.6 Don't Need Better Prompts. They Need A Clean Setup. — Transcript
Full transcript
- 0:00I overbuilt my harness for Fable 5 and
- 0:02Chat GPT 5.6 and I did it before they
- 0:05came out. I was so excited. I've been
- 0:07using AI a lot. I've been adding new
- 0:09rules and new instructions and new
- 0:10skills and I've been enthusiastic and
- 0:12you know what that means? It means my
- 0:14everything around it, all of the skills,
- 0:16all of the old system prompts, all of
- 0:18that was bloat that kept my AI models
- 0:22from performing well. Every time my AI
- 0:24missed something, I added another rule.
- 0:26A lot of us do, right? I would tell it
- 0:27use these sources correctly or write in
- 0:29my voice or check the length in this way
- 0:31or read this file first. Never make this
- 0:33mistake again. Make no mistakes, right?
- 0:35Every rule ended up fixing a real
- 0:38problem at the time, but over time, I
- 0:40didn't realize how bloated my harness
- 0:43had become. If you don't know what a
- 0:45harness is, that's part of the problem,
- 0:47too, because we haven't been clear about
- 0:49that as a community. A harness is
- 0:51everything wrapped around the model.
- 0:54So, it's your custom instructions. It's
- 0:56your project files. It's your saved
- 0:58prompts. It's your memory. Your your
- 1:00skills, your tools.
- 1:02It's your permissions, any checks that
- 1:04you run, right? It shapes the answer
- 1:06before you type anything into the prompt
- 1:08window because it comes with that prompt
- 1:10and helps the model understand what to
- 1:12respond with. Most of us never set out
- 1:14to build a harness. We built it
- 1:16accidentally, one correction at a time.
- 1:18I had no screen that showed me the full
- 1:20extent of the harness I built in Codex
- 1:22and Claude. I could see the latest rule
- 1:24I added if I looked for it, but I
- 1:26couldn't see the whole system in one
- 1:27place.
- 1:29When the AI started behaving strangely,
- 1:31I would often blame the model and I know
- 1:33a lot of people do and then I would add
- 1:34another rule and that would make it even
- 1:36more bloated, right? So, I built a skill
- 1:38that makes the harness visible. I'll
- 1:40show you six principles for building a
- 1:42more stable harness, including what
- 1:45works differently for Fable 5 and the
- 1:47Chat GPT 5.6 family of models. Then
- 1:50we'll show you how you can run the same
- 1:52cleaner on your setup. The goal is very
- 1:54simple. The models keep evolving, right?
- 1:57And that's not going to change. Now,
- 1:58when I ran the first inventory on my
- 2:00setup, it found 66 reusable skills and
- 2:04172 instruction-related files. I've been
- 2:06very busy. One normal writing job could
- 2:09pull in an 18,000-word file before it
- 2:12even adjusted the prompt. But, the
- 2:14number wasn't the real problem. Some of
- 2:16those instructions actually protect
- 2:17really important work. They tell the AI
- 2:20which sources matter when I'm doing
- 2:21research. They stop it from inventing an
- 2:24opinion that I genuinely don't have. I
- 2:26don't like it when it does that. They
- 2:28define what it can do without asking me
- 2:30and what it needs to ask me about.
- 2:32Problem was, I couldn't tell with the
- 2:34new intelligence, the new models, which
- 2:37rules were protecting the work and which
- 2:39were just copies or overlaps or arriving
- 2:42too early or unnecessary with today's
- 2:44models or or overlapping in a way that
- 2:46was confusing to the model. It's like
- 2:47everything around the engine in the car
- 2:49all the way to what makes the wheels go,
- 2:50right? The drive shaft is an example.
- 2:52The the full chassis. All of the parts
- 2:54of a car that are not the engine that
- 2:56are required to transfer force from the
- 3:00engine to the wheels. That's what a
- 3:02harness is for an AI model. It actually
- 3:04makes the work possible and we should
- 3:06definitely design it intentionally and
- 3:08not just guess and throw bolts into the
- 3:10car. The thing is, when the model
- 3:12changes, the harness doesn't magically
- 3:14rebuild itself. Sometimes you choose the
- 3:16new model. Sometimes ChatGPT or Claude
- 3:18will actually change your default for
- 3:20you and retire an older model and it
- 3:22will route a a difficult request
- 3:24somewhere new. And sometimes it will
- 3:25even switch models in the middle of the
- 3:27conversation.
- 3:29You may not even know it happened,
- 3:31right? And the old setup, the old
- 3:33harness, the old chassis for that car
- 3:35stays in place. The new model may behave
- 3:38very, very differently as a result and
- 3:40you may wonder why. You may blame the
- 3:41model, right? The experience could get
- 3:43worse. So, you add another instruction.
- 3:45I pointed at the setup I already use and
- 3:47I ask it to clean the harness that it
- 3:49can see. I don't gather every prompt by
- 3:51hand. I don't decide what should be
- 3:52deleted before it starts. The skill
- 3:55begins by making a map of my harness.
- 3:57That leads to the first rule, right? You
- 3:59map the harness before you clean it.
- 4:01It's actually a principle, right? The
- 4:02skill does it, but it's because it's a
- 4:03good idea. The map gives every important
- 4:06control that you have in your harness a
- 4:08separate row and asks, "Where does this
- 4:10control live? When does it load? What
- 4:13job does it do? Who owns this? Is there
- 4:16any evidence that it still helps? What
- 4:18problem can it create if it's misused?"
- 4:21And this was the first time I saw my
- 4:24whole harness in one place. And I don't
- 4:25know about you, but that's It was really
- 4:27illuminating how much junk there was
- 4:29there. It also exposed a difference that
- 4:32chat bots normally hides. Other controls
- 4:34are actual locks. They have teeth to
- 4:36them, right? A permission can block an
- 4:38action, a schema can reject a broken
- 4:41JSON file, a task can refuse a bad
- 4:44result if it fails. Those are not the
- 4:46same kind of instruction, right? Until
- 4:49you map the harness, it can all look
- 4:51like text to you, but it doesn't read
- 4:53the same way to an AI. The second
- 4:55principle that I learned as I built this
- 4:57is to blame the right layer.
- 4:59Uh I tested the same underlying job with
- 5:01Fable 5 in two different setups. The
- 5:04compact setup gave Fable the goal, the
- 5:07facts, the permission boundary that I
- 5:09had for it, and the finish line. The
- 5:11thicker setup, the thicker skill, gave
- 5:14it all of that plus the full method,
- 5:16plus a scoring system, an eval plan, and
- 5:19classification scheme. And the thicker
- 5:21version did do the job more completely.
- 5:23The analysis that I got was much richer,
- 5:26but it also failed actual delivery
- 5:28requirements twice.
- 5:31One result broke the JSON, another broke
- 5:33the word limit. The compact setup
- 5:35finished correctly three times out of
- 5:37three. Like all three times it finished
- 5:39right. That doesn't mean that short
- 5:41prompts always win. I want to be clear
- 5:42about that. It means that the model and
- 5:44the harness produce the work that we
- 5:47care about together.
- 5:49If you blame the model for everything,
- 5:51you will keep adding instructions to
- 5:53solve problems created by instructions,
- 5:57not by the model. And so, the cleaner
- 5:59skill force is a much more useful
- 6:00question for all of us. Did the model
- 6:02fail or did the surrounding setup fail?
- 6:04Right? Did the harness fail? Let's jump
- 6:05to the third. The third rule is one
- 6:08rule, one home, one owner. Right? It's
- 6:11easy to remember. My setup had versions
- 6:14of the same authorship and source rule
- 6:17in 15 different top-level skills,
- 6:19indicating I cared about it a lot, but I
- 6:21was not getting traction or progress in
- 6:23any of the individual skills, so I kept
- 6:24working on it. For me, the critical
- 6:27thing that I care about is that I don't
- 6:29want AI putting words in my mouth. I
- 6:31don't want it to come back with research
- 6:32and pretend it's my opinion. And so, I
- 6:35have 15 different versions of a skill
- 6:37that stops AI from reading a pile of
- 6:39research and coming back in my voice,
- 6:41because I really, really care about
- 6:43understanding exactly what was
- 6:45researched and getting the citation
- 6:47correct. So, I want to keep that
- 6:48quality, but I sure don't need 15
- 6:51different files pretending to do that
- 6:53once, and clearly none of them quite did
- 6:55it right because I kept adding to it.
- 6:56Every copy is another place where the
- 6:58rule can drift. One version gets fixed
- 7:00after a failure, uh 14 don't, now it's
- 7:03out of sync, now the model has several
- 7:05versions of the truth. You see where
- 7:06this is going. This is just bad. The
- 7:08cleaner doesn't ask whether an
- 7:09instruction is too long. It's smarter
- 7:11than that. It asks what job the
- 7:14instruction does, where that job should
- 7:17live, and who should update it the next
- 7:20time something changes. A lot of our
- 7:22theme along the way here has been about
- 7:24making sure we load the right
- 7:26information at the right time. And the
- 7:27fourth rule really codes that clearly.
- 7:29It says we should load specialist
- 7:31knowledge when the work actually needs
- 7:33it. I had six different editorial guides
- 7:37that were loading whenever one writing
- 7:39skill ran. So, for example, the source
- 7:40guide matters a lot when I'm trying to
- 7:42do the research. And if I'm looking for
- 7:44examples for YouTube, I need to be
- 7:46finding that at the time when I'm
- 7:48wrestling with the YouTube script for
- 7:49you guys. All of that is great, but if
- 7:52you load all of it at the beginning,
- 7:53you're just loading a bunch of crud into
- 7:55your AI that is likely to produce worse
- 7:57results overall. Your research is worse
- 7:59because it's thinking about YouTube
- 8:01examples when it really should just be
- 8:03thinking about research. So in this
- 8:04world, the cleaner skill keeps the
- 8:06library, right? It keeps that library of
- 8:08specialist skills that's useful. It just
- 8:10changes when each part of the library
- 8:12appears. And this is where I think a lot
- 8:14of cleaning prompts don't really
- 8:17fully understand what they're doing. The
- 8:19goal is not to throw away any useful
- 8:22context. Your library can actually be
- 8:24quite large. That's not bad. The fifth
- 8:25rule is that if you have hard
- 8:27requirements, they need hard checks. And
- 8:29so there's a difference between telling
- 8:31the model that you have a point of view
- 8:32and asking it to wrestle with that point
- 8:34of view with you. And there's a simple
- 8:35rule here that the cleaner skill will
- 8:37check for. If you have those kinds of
- 8:39rules, 50 word type rules where it's yes
- 8:41or no answers that the model can test
- 8:43for, those should be put into a schema
- 8:47that the model can test against. And the
- 8:48cleaner skill can help with that. The
- 8:50key principle here is that we let the
- 8:52system enforce the parts that a machine
- 8:54can verify, and that makes the harness
- 8:56lighter and safer at the same time. Now
- 8:58the sixth and final rule is to build for
- 9:00the model and the product actually doing
- 9:03the work. We want to be sophisticated
- 9:04enough that we can actually
- 9:06differentiate Fable 5 from Chat GPT 5.6.
- 9:09If you use AI a lot, you understand that
- 9:10Fable 5 in Claude.ai doesn't have the
- 9:12same harness as Fable 5 in Claude Code
- 9:15or via the API. Chat GPT 5.6 in Chat GPT
- 9:19Work isn't the same as in Codex and
- 9:21doesn't expose the same controls as it
- 9:23does via the API. So the model matters,
- 9:25but the product around it determines how
- 9:27skills load, which tools exist, what can
- 9:30be checked, and what proof comes back.
- 9:32What facts does the model need? What is
- 9:34it allowed to do? What must be true
- 9:37before the work is finished? What's our
- 9:38eval effect? And those core rules don't
- 9:40change just because I switched from
- 9:42Claude to chat GPT cuz the value on the
- 9:45work is the same, right? We're testing
- 9:46the same work across both harnesses.
- 9:48That setup passed every single delivery
- 9:50requirement three times out of three,
- 9:52right? Like it actually delivered what
- 9:54it needed to because it was suited to
- 9:56Fable. And as I discussed the richer
- 9:57skills, the heavier harness, they
- 9:59produced richer analysis, but Fable was
- 10:02unable to meet my output constraints.
- 10:04And so there were issues that Fable had
- 10:05just because it was struggling with the
- 10:07thicker skill. And so what does this
- 10:08teach us about how Fable works, right?
- 10:10You give it the real outcome, you give
- 10:11it context that it can't infer, you give
- 10:14it room to inspect the problem, and you
- 10:16let it plan its own approach usefully.
- 10:19And then you bring in specialist
- 10:20material when the work reaches that
- 10:23phase, which the cleaner skill can help
- 10:25you so that you invoke it at the right
- 10:26time, or you can let Fable invoke it at
- 10:28the right time. And what failed was
- 10:30narrating the entire method in that big
- 10:32thick skill file before Fable had even
- 10:35seen the job, and expecting all of that
- 10:37prompt prose to enforce like JSON and
- 10:40word count requirements. That's just not
- 10:41going to work really well, even if
- 10:43Fable's very good at rule following. And
- 10:46this is not about starving Fable of
- 10:47context, by the way. The full method
- 10:49helped it to notice more stuff, but the
- 10:52depth should arrive when the work needs
- 10:54it, not all up front in a way that
- 10:56confuses the model, right? So here we
- 10:58have the audit results. 66 skill roots
- 11:00and 172 instruction assets. 27,000
- 11:04description characters against 8,000
- 11:06Codex Discovery budget. That is a big
- 11:10problem. It's a big problem because it
- 11:11means Codex can't read it. There was an
- 11:1318,000
- 11:15word content route or skill with a
- 11:18minimum fan in. Basically, it was a
- 11:19skill leading into skills that 18,000
- 11:21words.
- 11:23Uh provenance governance, how you handle
- 11:25sources distributed across 15 different
- 11:27skills, that's what I talk about. And
- 11:29only six of the 66 root skills had a
- 11:32detected local eval, right? I had evals
- 11:36on six of them, but I didn't have them
- 11:37on the other 60, and I needed to improve
- 11:41that. And it's all about consistency
- 11:43with Codex, right? Schemas, tool
- 11:44restrictions, file checks, and a run
- 11:46receipt, they all carry the exact same
- 11:48requirements from skill to skill so that
- 11:50Codex is able to operate very
- 11:52consistently. So, the Codex chat GPT
- 11:54cleanup is not simply make the prompt
- 11:56shorter. It's make the right route to
- 11:58the skill easy to find, then load the
- 12:01depth at the right point. And the
- 12:03receipt that you have, which is
- 12:04important for Codex, it records the
- 12:06model, the reasoning setting, the tools,
- 12:08the skills, the fallbacks, and the
- 12:09checks that ran, so that you can get a
- 12:11sense over time of where there are
- 12:13problems. It's like a diagnostic for
- 12:14your engine, right? You can figure out
- 12:16what's going on. If we compare the two
- 12:17models together, Fable's failure mode
- 12:19showed up after the method became too
- 12:22heavy for the delivery job, whereas chat
- 12:24GPT 5.6's failure mode in Codex showed
- 12:27up much earlier, while the system was
- 12:29still trying to find the right method at
- 12:31all, and it was just having trouble
- 12:32routing across a really huge harness
- 12:33layer.
- 12:34Both models benefit from selective
- 12:37loading, right? Both benefit from
- 12:38including those hard checks I talked
- 12:40about, but for very different immediate
- 12:42reasons. And that has to do with how
- 12:45these models work. Like Fable 5 sorting
- 12:47through a lot of context and trying to
- 12:49figure out how to do the best job and
- 12:51kind of overloading itself is a Fable 5
- 12:53way to fail. Harnesses grow barnacles
- 12:56like ships. Ultimately, what I want is a
- 12:58system that helps us work well. I want
- 13:00the harnesses to evolve and stay clean
- 13:02and not just collect these barnacles of
- 13:04extra point solutions, extra additions,
- 13:06extra text files, extra skills that I
- 13:09randomly find on the internet because I
- 13:10found a tweet somewhere, add it in, and
- 13:13now I don't know what I'm doing because,
- 13:15you know, chat GPT is overwhelmed or
- 13:17Fable 5 gets too fat a skill, or
- 13:18whatever the failure mode may be. I want
- 13:20a clean harness that lets me get work
- 13:22done in a way that's efficient. That's
- 13:24why I built this cleaner skill. And this
- 13:26matters, right? If you're a product
- 13:27manager, it means that you have like a
- 13:29clean, simple note that loads the right
- 13:31PRD skills at the right time. If If
- 13:33you're a developer, it means having one
- 13:35source of truth instead of several
- 13:36instruction files arguing about how the
- 13:38repository work. But we need to make
- 13:41sure that they add into the system we
- 13:42have in a way that's efficient. And if
- 13:43you're just using ChatGPT or Claude at
- 13:45home or at work, it means that your old
- 13:48memories, your project files, your
- 13:49examples, your corrections stop quietly
- 13:52shaping answers in ways you don't
- 13:54necessarily intend. You may have given
- 13:56Claude or ChatGPT a correction 6 months
- 14:00ago that it's remembering and over
- 14:02applying now that you have a new model.
- 14:04So, for example, you and may have told
- 14:06it a few months ago, "You have to show
- 14:08me the step-by-step because back in
- 14:10November that was really important." But
- 14:12now you don't need to do that cuz the
- 14:13models are good, right? And so, those
- 14:15are the kinds of things the cleaner
- 14:16scope catches. Ultimately, the goal is
- 14:18simple. Your AI should become easier to
- 14:20use to do useful work because the system
- 14:23around it should become easier to
- 14:24understand. Once you can see that
- 14:27system, you can actually own it.
- 14:29You can tell which controls protect you,
- 14:31which need one single home, right? Which
- 14:34should arrive later in the process,
- 14:36which should become real locks instead
- 14:38of just polite reminders. I put the
- 14:40complete cleaner over on the Substack,
- 14:42you can find it. You can install it
- 14:43once, you can point it at the AI setup
- 14:45or project of your choosing, and you can
- 14:48just let it show everything that's
- 14:49accumulated. It gives you a map of the
- 14:52harness and the cleaning decisions. It's
- 14:54going to give you a plain English before
- 14:56and after and a receipt showing what ran
- 14:58when you allow it to run and and and run
- 15:00those corrections. You can review the
- 15:02changes before anything important moves.
- 15:04I don't want this to be something that
- 15:05surprises you in any way. I have a
- 15:07feeling that a lot of us are going to
- 15:08discover that we built a lot more than
- 15:09we realized. I know I certainly did, and
- 15:11so I wanted to showcase that here and
- 15:13show you all how bloated my harness had
- 15:15become and how much I needed to clean. I
- 15:18hope this has been useful for you. I'm
- 15:20going to keep drilling in on these new
- 15:21models. I'll keep updating the cleaner
- 15:23skill. Please let me know in the
- 15:24comments one how it works for you, but
- 15:26also let me know what other models you'd
- 15:29like me to optimize the cleaner skill
- 15:30for because of course there are plenty
- 15:32of other models besides Claude and chat
- 15:33GPT and I'm happy to start working on
- 15:36that as well. Would love to kind of
- 15:38expand this and make this more useful
- 15:39for the community over time and I look
- 15:41forward to showing you what I've built
- 15:43with my new and improved harness on
- 15:45Friday.
About this transcript
This page contains the full transcript of Fable 5 And GPT-5.6 Don't Need Better Prompts. They Need A Clean Setup. by AI News & Strategy Daily | Nate B Jones, generated from the public captions YouTube serves with the video. The transcript has 3,224 words across 460 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.