Building a $1,500 Lead Artifact: Opus 5 vs GPT-5.6 Sol — Transcript
Full transcript
- 0:00Today, we're putting Opus's 5
- 0:02head-to-head with Chat GPT's 5.6 Soul.
- 0:05Now, this isn't some random
- 0:06demonstration of some code base I have
- 0:08zero intention on using. I'm actually in
- 0:10the market for this. So, I figured what
- 0:12better way to test these models than to
- 0:15test it for something I need to build.
- 0:17So, without further ado, let's build.
- 0:20All right. So, as you can see here on
- 0:21the screen, we have two fresh terminals
- 0:23open. I just said hi on both to get the
- 0:26meter logged and start working. What's
- 0:29interesting is Opus 5 on a high cost 30
- 0:32cents to say hi. Uh whereas GPT 5.6 Soul
- 0:36only cost us just about 5 cents to say
- 0:39hello.
- 0:40>> [laughter]
- 0:41>> Uh that's funny. Now, here's the prompt
- 0:43for today.
- 0:44You are building a front-end only
- 0:46prototype of a client-facing AI ops
- 0:48audit report. This is not a production
- 0:50application, no real authentication, no
- 0:53real back-end, no real database, no
- 0:55payment processing, no scheduling
- 0:56integration, no PDF generation service,
- 0:59no multi- tenant logic. All data is
- 1:02synthetic and hard-coded or mocked in
- 1:04the front end. All export and send
- 1:06actions are non-functional UI states and
- 1:08must be visibly non-functional, not
- 1:11misleadingly styled as real. Now, before
- 1:14I continue reading the rest of this,
- 1:15let's copy this prompt and let's get
- 1:18started with this demonstration here.
- 1:24Bam. Bam.
- 1:26And I'm going to switch this to auto
- 1:27mode.
- 1:30We already have permission set to
- 1:32approve for me.
- 1:33Um so, we're just going to let those get
- 1:35to work and we'll continue reading the
- 1:37prompt. So, stack React plus Vite
- 1:39Tailwind CSS
- 1:41Shadcn UI components, Recharts for any
- 1:44data visualization, ship a single
- 1:47deployable front end, no server
- 1:48required. Build these sections as one
- 1:50flowing client portal.
- 1:53So, we have a lot of specifics here. I'm
- 1:55not going to read these line for line,
- 1:57but if you do want a copy this, just
- 1:58screenshot the page here.
- 2:01Synthetic client
- 2:03placeholder
- 2:06what we heard source material,
- 2:07opportunity candidates to score, right?
- 2:09So, we're going to have a certain
- 2:11scoring chart so to speak so that the
- 2:13customer or client can visualize the
- 2:16opportunities based off of these
- 2:19specific metrics and then a few things
- 2:22to note at the end intentionally not
- 2:23specified visual design direction
- 2:26information architecture copy tone chart
- 2:29types section order beyond the list
- 2:30above
- 2:32and how confidently to score each
- 2:33opportunity.
- 2:35That judgment is what's being tested,
- 2:37right? So,
- 2:39while I gave it a rather detailed
- 2:41prompt,
- 2:42the reason I'm giving it this detailed
- 2:44prompt is I actually want to use one of
- 2:46these code bases moving forward. I have
- 2:47a need for this. This is not me building
- 2:49something I don't need. Uh hence the
- 2:51purpose of this demonstration here. So,
- 2:54I really wanted to be able to judge.
- 2:56We're not going to factor in speed that
- 2:58much. They should finish around the same
- 3:00time I'd imagine.
- 3:02But, I want to test its judgment in
- 3:04visual design, how it's going to
- 3:06architect this portal, and the overall
- 3:09experience end to end, right? Uh that
- 3:12way I can pick a code base I'm happy
- 3:14with and then actually build the thing
- 3:16on audit.clearmud.ai.
- 3:20So, let's get back to the demo here.
- 3:22And hey, they're both getting started.
- 3:24So, what we're going to do
- 3:26is whoop, let's also scroll down. I have
- 3:28a lot of meters open. So, these are the
- 3:30two meters associated with the project.
- 3:32We can also confirm here by looking as
- 3:34you can see audit app.opus all sessions.
- 3:38That's what this one's called here.
- 3:40The same goes for GPT.
- 3:44audit.app.GPT
- 3:45all sessions
- 3:48And what we're going to do is we're just
- 3:50going to fast forward through the boring
- 3:51part.
- 3:55All right. So, I wanted to end that time
- 3:58lapse here. As you can see, GPT finished
- 4:01in lightning speed.
- 4:03Uh it finished first.
- 4:06And it's already come in a lot cheaper.
- 4:09I'm actually very shocked by this price.
- 4:12Uh it even did some QA and it it I
- 4:14thought it was going to be local, but it
- 4:16deployed it to Vercel.
- 4:18Very interesting. While Opus is still
- 4:20working. So, before we check that out,
- 4:23we are going to kind of prepare, copy
- 4:26this,
- 4:28add in a browser tab to the right,
- 4:31paste that in there.
- 4:32I'm going to hit enter, but I'm going to
- 4:33immediately tab over because I do not
- 4:35want to
- 4:37um
- 4:38cheat by looking at it first.
- 4:42Wow.
- 4:45It's a big price discrepancy there. So,
- 4:47looks like this was 1.95
- 4:50million input.
- 4:52Okay.
- 4:53I'll
- 4:54I'll have to cross-reference. I think
- 4:56that's accurate, right? So, very
- 4:58interesting. Very interesting. So, we're
- 5:00just going to fast forward through the
- 5:01boring part until Opus finishes.
- 5:06All right. So,
- 5:09>> [laughter]
- 5:10>> I was not expecting this.
- 5:13Opus 5 appears to be insanely token
- 5:17hungry.
- 5:19It's so absurd that I'm going to have to
- 5:22go back at the end of this demo to make
- 5:25sure that my my API meter is accurate.
- 5:28Um what?
- 5:33>> [laughter]
- 5:34>> Uh
- 5:35so,
- 5:36that took
- 5:3727 minutes
- 5:39on Opus, 11 minutes on ChatGPT. The
- 5:43price kind of shows you can see the
- 5:46tokens. I mean,
- 5:48wow.
- 5:49Now, one undesirable thing here I'm
- 5:51already seeing, it's
- 5:54it's sending us to an artifact through
- 5:57our account.
- 5:58I honestly thought they were going to
- 6:00both spin up a local preview for me.
- 6:03Uh chat GPT went above and beyond uh and
- 6:06deployed the Versel. I know for a fact
- 6:09my Claude CLI has access to that. So,
- 6:11I'm I'm a bit curious like, okay, Claude
- 6:14defaults to sending you to their
- 6:16platform, whereas it looks like Codex
- 6:18defaults to maybe
- 6:21the last method of deployment uh because
- 6:24I have been deploying a lot with
- 6:26Codex. Um wow, okay.
- 6:29Well,
- 6:32let's go ahead and open this up. I might
- 6:34have to log in to
- 6:36Oh, no, I don't. Okay, great.
- 6:39Do not want to share. All right. So,
- 6:42since GPT finished first,
- 6:45we're going to go with GPT's demo first.
- 6:48Now, let's go full screen here.
- 6:53And we'll bring that down.
- 6:56All right, nice split screen. I love how
- 6:58this looks, super professional.
- 6:59Obviously, it's not our brand color, but
- 7:01I didn't give it our palette. I didn't
- 7:03give it uh a clear method front-end
- 7:05design skill to follow. I wanted it to
- 7:07use its best judgment.
- 7:09Continue with Google. Okay.
- 7:11Final report. So, this is just a
- 7:13one-page scroll from what I can see.
- 7:14We're We're scrolling down
- 7:16and it allows you to kind of navigate
- 7:18between Cool. This is a great-looking
- 7:20brief.
- 7:22Very professional.
- 7:32Key clear pain points we heard in the
- 7:34meeting.
- 7:36Opportunity map.
- 7:38Chart is not necessarily
- 7:41Is it working as as as intended? Kind
- 7:44of. I would have loved to see a little
- 7:46bit more, but I get what it's doing with
- 7:48the dots. Very impressive. Look at these
- 7:51Look at this functionality. It's pretty
- 7:53cool. I would make those a little bit
- 7:54larger, right?
- 7:57So, five for efficiency, revenue upside,
- 8:00revenue protection, effort. Okay, great.
- 8:05Feel like we would want to add an
- 8:06overall score baseline as well.
- 8:09Build the operating layer in the right
- 8:10order.
- 8:13Okay.
- 8:15It's pretty solid.
- 8:18From audit to operating rhythm, kind of
- 8:20gives a timeline.
- 8:23One thing it's missing though that I
- 8:25realized I forgot to is like how would
- 8:27we build it? Yes, this is an audit. This
- 8:29is where what we would build,
- 8:32um but I would want a tools section
- 8:34here. And I realize this is my fault for
- 8:36not including that, but I was hoping
- 8:38that one of the models would decide uh
- 8:40as far as giving certain recommendations
- 8:43per uh pain point. But this is okay.
- 8:47Listen, needs some work. Needs some
- 8:49work, but solid solid output. Okay. Now,
- 8:53let's go over to
- 8:57Claude, shall we?
- 8:58Very similar split screen approach.
- 9:02Again, I would want to give it some some
- 9:04coloring here.
- 9:06Wow, very
- 9:09very basic on the coloring side.
- 9:20Okay, what we heard.
- 9:26Okay, I would want to kind of change the
- 9:28visual. I guess depending on the target
- 9:30audience, this might work.
- 9:32But I do love the idea of having that
- 9:34vertical menu bar rather than this,
- 9:37right?
- 9:38Um so interesting difference that Opus
- 9:40went with.
- 9:42Cool little accordion effect, pain
- 9:44points,
- 9:45business consequence, okay. Now, again,
- 9:49I was hoping one of these models would
- 9:50have chosen. I guess it was something I
- 9:52had to get specific with as far as what
- 9:54tools or what solutions we would
- 9:56recommend based on those pain points.
- 10:00Very similar opportunity map.
- 10:04Decided to go with this kind of
- 10:06I forget what you call this chart, but
- 10:08this approach.
- 10:10Okay.
- 10:16Feels more detailed, for sure.
- 10:27But again,
- 10:28okay.
- 10:45Okay.
- 10:49Interesting, eh? You know, I'm going to
- 10:51be honest with you. While
- 10:55I do like Opus's output, the fact that
- 10:58it almost took three times as long and
- 11:00three times the cost for that matter, um
- 11:06I would have to go with GPT. It just it
- 11:09was so much more efficient.
- 11:11Uh obviously, I'm going to have to
- 11:13improve on this a bit. This is not for
- 11:16the public. This is something I'm going
- 11:18to use internally once we perform the
- 11:20audit. I'm also going to be building
- 11:22um an audit
- 11:24kind of dashboard for me and the
- 11:26prospect who signs up and and and
- 11:29schedules that booking to walk through
- 11:31together.
- 11:33Um but more importantly, right? We're
- 11:35going to have a very fluid natural
- 11:37conversation.
- 11:38And I'm just going to be listening first
- 11:40and foremost, and then I'm going to have
- 11:42an input to where I can take that
- 11:43transcript, that audio file from the
- 11:45meeting, and input it into my platform
- 11:48so that I can have Claude or Codex or
- 11:51both
- 11:52analyze that transcript, define the pain
- 11:55points, define what we hear, right?
- 11:58Define the pain points, create the
- 11:59opportunity map, and then work together,
- 12:02right? This is not me saying I'm going
- 12:04to use one versus the other for that
- 12:06process, that part of the pipeline and
- 12:08workflow. I'm going to use them both. If
- 12:10I'm being honest with you, I'll be using
- 12:11both of them. Um
- 12:13but I I just got to say it.
- 12:16I like how much more professional this
- 12:18one looks. This one looks too much like
- 12:21I'm just This isn't my style,
- 12:24right? So, I don't like the direction it
- 12:26took it in. Had it packaged it
- 12:29Had Had the packaging been a bit
- 12:30different,
- 12:32I might be leaning towards Opus here,
- 12:34but
- 12:35yeah.
- 12:37Yeah.
- 12:39In my book, Codex is the winner here. Um
- 12:42now,
- 12:43I'm very curious. Let's go over to my
- 12:46model meter.
- 12:52And I'm just going to take a screenshot.
- 12:55Please reference this screenshot. They
- 12:57are the
- 12:59audit app Opus session for Opus and then
- 13:03the audit app GPT session for the GPT
- 13:05meter.
- 13:07Are these numbers accurate? I'm just so
- 13:10blown away by the differences. Can you
- 13:12please double-check the numbers,
- 13:15uh cross-reference them to these two
- 13:16links. These are the official uh API
- 13:19pricing docs.
- 13:20And let me know your findings.
- 13:25All right.
- 13:27So, this is just a little bonus part of
- 13:28the episode to make sure we're being
- 13:30thorough and accurate.
- 13:40I'm just blown away at the cost
- 13:42difference. That is kind of insane. And
- 13:44for me to still pick GPT, I thought for
- 13:46sure I was going to go with Opus, but
- 13:47Opus and Anthropic, it defaulted to it's
- 13:50kind of like it feels like a
- 13:55What's the word I want to use to
- 13:56describe this? Uh
- 13:59It's just not my style. It's not my
- 14:00style, right? Uh it doesn't feel that
- 14:03professional, if I'm being brutally
- 14:05honest.
- 14:09Aha! Okay, I was wondering. So, we're
- 14:13going to have the updated pricing here
- 14:14in a moment.
- 14:16It'll be able to fix it in line.
- 14:19Let's just patiently wait.
- 14:21All right. So, as you can see here,
- 14:23applying the Claude fix now, GPT figure
- 14:26will remain unchanged because its
- 14:28published sole rates are already
- 14:29correct. Yeah, okay. So,
- 14:33three times more than what it should be.
- 14:35So, I'm guessing it's going to come down
- 14:36to about 1185
- 14:40or what?
- 14:42Just under $12.
- 14:45$13. Okay, that is more accurate of a
- 14:48demonstration because the numbers
- 14:51weren't adding up for me here. Um
- 14:54>> [laughter]
- 14:55>> Okay. Now, I'll make sure to commit this
- 14:58uh
- 14:58once done.
- 15:00Please commit to get.
- 15:05Merge to main. Thanks.
- 15:08Now, for anybody curious, this is an
- 15:10open-source model meter that I do have
- 15:11live. If you just check out the ClearML
- 15:13GitHub repository or the link in this
- 15:15description, uh you'll be able to access
- 15:17it and use it for yourself. Now,
- 15:21that makes so much more sense.
- 15:26So much more sense. I I I just I'm glad
- 15:29I checked.
- 15:32I'm glad I checked. Wow.
- 15:36Yeah, listen. Like even with the the
- 15:40savings here, like I just like GPT's
- 15:43output more.
- 15:47Listen, Opus
- 15:48and and Anthropic models are always my
- 15:50go-to when I'm in ideation, creation, uh
- 15:54prototyping phases, but once I'm ready
- 15:56to proceed and get ready to production,
- 15:59I trust GPT to be far more efficient um
- 16:03and produce a production-ready codebase.
- 16:06Uh that's where I run all my security
- 16:07audits with. Although some of the tests
- 16:09I've been seeing run with Opus 5 and
- 16:12some of the benchmarks,
- 16:14apparently it does a very good job at
- 16:15security audits. So, I think my new
- 16:17workflow here is going to be running
- 16:18security audits
- 16:20for both, right? With both models
- 16:22against that same repository.
- 16:25So, this is not a I'm going to use one
- 16:27versus the other demonstration. I'm
- 16:29literally going to be using both of
- 16:30these. So,
- 16:31I hope you found some value in today's
- 16:33video. Now, I'm not an AI expert. I'm
- 16:35building in public and sharing what
- 16:36actually works. If you'd like to see me
- 16:38make a particular video or run a
- 16:39specific topic, drop a comment below.
- 16:41Let me know who you are, what you do,
- 16:43who your target audience is, and what
- 16:45your question is, and I will add a video
- 16:46to my queue custom-tailored just for
- 16:48you. Thank you so much for tuning into
- 16:50today's video. My name is Marcelo. This
- 16:52is Clear Mod and Clarity Matters.
- 16:58>> [music]
About this transcript
This page contains the full transcript of Building a $1,500 Lead Artifact: Opus 5 vs GPT-5.6 Sol by Clearmud, generated from the public captions YouTube serves with the video. The transcript has 2,380 words across 415 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.