This Is What Voice Coding Was Supposed to Be (Qwen Audio Agent) — Transcript
Full transcript
- 0:00Voice coding falls apart the moment the
- 0:02agent has to do actual work. Refactor
- 0:05this file, run the build, and then
- 0:08silence. The agent disappears and we're
- 0:10left staring at the mic icon, waiting
- 0:12for it to come back. Quen audio agent
- 0:15works differently. Claude code can be
- 0:17refactoring the background while you
- 0:19keep talking to the voice agent. And
- 0:21when the task finishes, it can interrupt
- 0:23you and tell you it's done. And here's
- 0:24the weird part. Quen audio agent isn't a
- 0:27model. It ships zero model weights. It's
- 0:29a runtime built around one thing most
- 0:31voice coding setups completely miss.
- 0:34Keeping the conversation alive while the
- 0:36actual work happens.
- 0:42Now let me set up why that matters
- 0:45because I think most people have the
- 0:46wrong idea about voice encoding. The
- 0:49usual setup is dictation. Whisper
- 0:52listens, transcribes, pastes text into
- 0:54your editor. That's useful, but look at
- 0:57what actually it replaces. It replaces
- 0:59the whole typing part. You sit there,
- 1:02you still wait. You've swapped the
- 1:04fingers for your mouth and changed
- 1:05really nothing. It does speed things up,
- 1:08though. Full duplex means something
- 1:10different. It means both sides can talk
- 1:12at the same time. You can cut in
- 1:14mid-sentence, and more importantly, the
- 1:16thing on the other end can keep the
- 1:18conversation alive while work happens
- 1:20somewhere else. The second half is the
- 1:22whole reason this exists. If you enjoy
- 1:24coding tools that speed up your
- 1:26workflow, be sure to subscribe. We have
- 1:28videos coming out all the time. All
- 1:29right, now let me show you why this is
- 1:31different. I'm going to make it talk for
- 1:33way too long on purpose. Not going to
- 1:35lie, the setup was actually really
- 1:37frustrating. It was all in Chinese, so I
- 1:40had to go to the Alibaba cloud English
- 1:42site, set the desktop UI for frontend
- 1:45and backend worker with the Alibaba API
- 1:47key, and then add some anthropic API
- 1:50credits for the backend. I did have to
- 1:51translate some of that. Once it's all
- 1:53synced, though, it does work pretty
- 1:55good. I can just speak without pressing
- 1:57anything. It's always listening unless I
- 1:59turn off the mic. So, here we go. I'm
- 2:02going to give you a code repo that I
- 2:04need you to work on. It starts
- 2:06explaining that. Okay. And then halfway
- 2:08through, I'm just going to go here.
- 2:10Stop. Too long. I get it.
- 2:12>> Understood.
- 2:13>> And it cuts off just like that. It it
- 2:15listened to this. Now it's all gone. I
- 2:18can refresh from here. It doesn't finish
- 2:20the sentence because I cut it off.
- 2:21There's no extra second of audio
- 2:23fighting me after I start talking here.
- 2:25The runtime actually kills the old turn
- 2:28and moves on. I'll paste in the path to
- 2:30the repo that we are going to work on
- 2:32here.
- 2:34Now watch what happens with something
- 2:36simple. What's 19 * 24? Don't touch the
- 2:40repo. I got an immediate answer though
- 2:42because clawed code never needed to see
- 2:44that. The voice agent can handle the
- 2:46easy stuff itself, which is great
- 2:48because I don't want every random
- 2:50question turning into a big task or
- 2:52racking up my API credits for anthropic.
- 2:54But now I'm going to give it something
- 2:56that actually takes a bit more time. The
- 2:58code repo that I gave it is just
- 3:00TypeScript practice problems. I could
- 3:02have chose something harder, but let's
- 3:04just keep it straightforward for now.
- 3:05I'm going to say here, refactor record
- 3:08problem TS. Make sure it works
- 3:10correctly. Now watch this.
- 3:13The task leaves the voice agent. Claude
- 3:16code picks it up and this work card
- 3:18appears. It's the proof that Claude
- 3:20actually has the job now. And normally
- 3:22this is where a voice coding setup could
- 3:24fall apart. You've asked it to do
- 3:26something real. So now you have to wait.
- 3:29Except here I don't have to. I can say
- 3:32also can you go make me a normal JS file
- 3:35with an async function to handle HTTP
- 3:37requests?
- 3:40It answers and does both while they're
- 3:43both running side by side. Claude is
- 3:45still working. Now I'm going to fire out
- 3:46a random question. Also, what is the
- 3:49weather like in Dubai?
- 3:52Okay, I can read it from here. Thanks.
- 3:55Another answer. Same conversation. And I
- 3:58can even ask about the task that is
- 4:00being worked on without having to wait
- 4:02for it. Hey, how's the refactor going?
- 4:06Don't wait for it to finish, just the
- 4:07status.
- 4:11Now it can tell me Claude is editing,
- 4:13running tests, or still working. There
- 4:15are basically two things happening at
- 4:16once. I'm talking to one agent, another
- 4:19agent is doing the work, and neither one
- 4:21has to stop because the other one is
- 4:23busy, which means while Claude finishes,
- 4:26I can just keep going. The coding task
- 4:28finished, the result came back into the
- 4:30live conversation, and I didn't have to
- 4:32check anything. I haven't seen an open-
- 4:35source desktop runtime handle that this
- 4:37cleanly. So underneath all this, the
- 4:39setup is actually pretty simple. On one
- 4:41side, we've got the voice agent. Okay,
- 4:44great. That handles conversation,
- 4:46interruptions, status, cancelling, a lot
- 4:48of the fast stuff. On the other side,
- 4:51clogged code is connected over ACP.
- 4:54That's the side touching files, running
- 4:56tools, and doing the slow work. Now,
- 4:58mechanically, it's really two lanes
- 5:00here. And once you see that, this design
- 5:02is a bit more obvious. There's a
- 5:04lightweight front-end voice agent whose
- 5:06only job is to hold the conversation. It
- 5:08answers the easy questions immediately.
- 5:10What's the status? Cancel that. Never
- 5:12mind. Then there is your real coding
- 5:14agent on the back end and they talk over
- 5:16the agent client protocol, the ACP. Real
- 5:20work gets handed over as an async task
- 5:22and runs on its own. When it finishes,
- 5:24the result gets injected back into the
- 5:26live conversation. So the voice layer
- 5:29never blocks on the slow thing. And
- 5:32because it's a protocol rather than an
- 5:34integration, the backend is swappable.
- 5:36Clawed Code, Codeex, Open Code, and a
- 5:39bunch of others. Now, one detail I want
- 5:41to flag because it shows someone was
- 5:43actually paying attention to this
- 5:44interruption isn't handled as an event
- 5:46here. It's a state machine. Detect user
- 5:49speech, mark an interruption, admit an
- 5:51interpreted signal, actively surpass any
- 5:53audio and transcripts still in flight,
- 5:56then start the new term. Their own
- 5:58design note says the engineering quality
- 6:00of interruption determines the
- 6:02reputation. And they're right. The thing
- 6:04that makes a voice assistant feel broken
- 6:06isn't slow responses. It's interrupting
- 6:09it and still hearing another second and
- 6:10a half of the former sentence. That kind
- 6:13of gap is actually rather annoying and a
- 6:15big gap. This has a slight one
- 6:17sometimes. Now, where does all this sit
- 6:20next to what we already know? Well, the
- 6:22OpenAI real-time API is the layer
- 6:24underneath this. It's the thing that
- 6:26does speech in speech out. This runs on
- 6:29top and is provider pluggable. Pycat and
- 6:32LiveKit agents are frameworks. They give
- 6:35you the parts and you assemble the
- 6:37pipeline. This is assembled and pointed
- 6:40at a specific job. And to whisper to LLM
- 6:43to TTS pipeline you wrote yourself is
- 6:45the thing that a lot of us have, a lot
- 6:47of us build out. Now, let's get real
- 6:49here for a second because there are two
- 6:50things here that the readme will not
- 6:52lead with. First, the runtime is open
- 6:54source. The voice is not. The code is
- 6:57Apache 2.0, but the default path, the
- 7:00best sounding one, roots through
- 7:02Alibaba's Dash scope with a paid API
- 7:04key. You get some free credits, but it's
- 7:07still paid. The optimized real-time
- 7:09voice is cloud API only. There is no
- 7:11self-hosted version of the good audio
- 7:13generation. So, open- source voice agent
- 7:16is only partially true here. There is a
- 7:18real escape hatch, and I'll give them
- 7:20credit for documenting it. a fully local
- 7:23mode using a hugging face speechtoext
- 7:25pipeline with an MLX backend on Apple
- 7:27silicone running on MPS. In full local
- 7:31mode, you need no cloud key at all. It's
- 7:34a second Python install and the tune
- 7:36profile in the docs is Chinese only.
- 7:38Okay, so translate it. But the path
- 7:40exists and it's pointed at max
- 7:42specifically. Second thing, and this one
- 7:45matters more, there are no published
- 7:46latency numbers anywhere on any hardware
- 7:50I went looking. Okay, now a final quick
- 7:52points here. It's Chinese first. The
- 7:54wake word is this, which translates to
- 7:57this. The main design article explaining
- 8:00the full duplex architecture is Chinese
- 8:02only. The tuned voice profile is
- 8:04Chinese. The model underneath claims a
- 8:06100 plus languages for recognition and a
- 8:08few dozen for speech. So, I had to
- 8:11translate the page to understand the
- 8:12gist of it when I was signing up for all
- 8:14this. In version 2.0 know is already
- 8:16listed in development with the
- 8:18architecture being rewritten and the
- 8:20stars are running ahead of the usage. 23
- 8:232400 stars here but around 4,000 mpm
- 8:25downloads a month and on the latest
- 8:27release 92 downloaded the Mac version
- 8:30build against 502 on Windows. Now, if
- 8:33you already work inside Cloud Code or
- 8:35Codeex, which I assume most of you all
- 8:37do, and you want it to stop being
- 8:39chained to the keyboard while long tasks
- 8:41run, this is the only thing I found
- 8:43doing the parallel task plus line
- 8:45conversation together, or at least the
- 8:47one that does it really well. Get the
- 8:49DMG. It's a signed universal build. No
- 8:52CUDA, no Python, nothing to compile. It
- 8:54takes about 10, maybe 15 minutes to get
- 8:56everything integrated. If you want a
- 8:58fully open, fully local voice stack,
- 9:00this isn't it. But the MLX path is a
- 9:02real starting point and it's more than
- 9:04most projects offer. And the idea worth
- 9:06taking away whatever happens to this
- 9:07particular repo is the thing that makes
- 9:09voice usable for real work isn't better
- 9:12transcription. Transcription got solved.
- 9:14It's whether the system can keep talking
- 9:15to you while it's busy. That's a
- 9:17concurrency problem, not a speech
- 9:19problem. And almost nobody is treating
- 9:21it like one. I'm Josh from Better Stack.
- 9:23If you enjoy coding tools like this, be
- 9:25sure to subscribe. We'll see you in
- 9:27another video.
About this transcript
This page contains the full transcript of This Is What Voice Coding Was Supposed to Be (Qwen Audio Agent) by Better Stack, generated from the public captions YouTube serves with the video. The transcript has 1,714 words across 254 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.