YouTube2Text

This Is What Voice Coding Was Supposed to Be (Qwen Audio Agent) — Transcript

by Better Stack · 1,714 words · 254 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Voice coding falls apart the moment the
  2. 0:02agent has to do actual work. Refactor
  3. 0:05this file, run the build, and then
  4. 0:08silence. The agent disappears and we're
  5. 0:10left staring at the mic icon, waiting
  6. 0:12for it to come back. Quen audio agent
  7. 0:15works differently. Claude code can be
  8. 0:17refactoring the background while you
  9. 0:19keep talking to the voice agent. And
  10. 0:21when the task finishes, it can interrupt
  11. 0:23you and tell you it's done. And here's
  12. 0:24the weird part. Quen audio agent isn't a
  13. 0:27model. It ships zero model weights. It's
  14. 0:29a runtime built around one thing most
  15. 0:31voice coding setups completely miss.
  16. 0:34Keeping the conversation alive while the
  17. 0:36actual work happens.
  18. 0:42Now let me set up why that matters
  19. 0:45because I think most people have the
  20. 0:46wrong idea about voice encoding. The
  21. 0:49usual setup is dictation. Whisper
  22. 0:52listens, transcribes, pastes text into
  23. 0:54your editor. That's useful, but look at
  24. 0:57what actually it replaces. It replaces
  25. 0:59the whole typing part. You sit there,
  26. 1:02you still wait. You've swapped the
  27. 1:04fingers for your mouth and changed
  28. 1:05really nothing. It does speed things up,
  29. 1:08though. Full duplex means something
  30. 1:10different. It means both sides can talk
  31. 1:12at the same time. You can cut in
  32. 1:14mid-sentence, and more importantly, the
  33. 1:16thing on the other end can keep the
  34. 1:18conversation alive while work happens
  35. 1:20somewhere else. The second half is the
  36. 1:22whole reason this exists. If you enjoy
  37. 1:24coding tools that speed up your
  38. 1:26workflow, be sure to subscribe. We have
  39. 1:28videos coming out all the time. All
  40. 1:29right, now let me show you why this is
  41. 1:31different. I'm going to make it talk for
  42. 1:33way too long on purpose. Not going to
  43. 1:35lie, the setup was actually really
  44. 1:37frustrating. It was all in Chinese, so I
  45. 1:40had to go to the Alibaba cloud English
  46. 1:42site, set the desktop UI for frontend
  47. 1:45and backend worker with the Alibaba API
  48. 1:47key, and then add some anthropic API
  49. 1:50credits for the backend. I did have to
  50. 1:51translate some of that. Once it's all
  51. 1:53synced, though, it does work pretty
  52. 1:55good. I can just speak without pressing
  53. 1:57anything. It's always listening unless I
  54. 1:59turn off the mic. So, here we go. I'm
  55. 2:02going to give you a code repo that I
  56. 2:04need you to work on. It starts
  57. 2:06explaining that. Okay. And then halfway
  58. 2:08through, I'm just going to go here.
  59. 2:10Stop. Too long. I get it.
  60. 2:12>> Understood.
  61. 2:13>> And it cuts off just like that. It it
  62. 2:15listened to this. Now it's all gone. I
  63. 2:18can refresh from here. It doesn't finish
  64. 2:20the sentence because I cut it off.
  65. 2:21There's no extra second of audio
  66. 2:23fighting me after I start talking here.
  67. 2:25The runtime actually kills the old turn
  68. 2:28and moves on. I'll paste in the path to
  69. 2:30the repo that we are going to work on
  70. 2:32here.
  71. 2:34Now watch what happens with something
  72. 2:36simple. What's 19 * 24? Don't touch the
  73. 2:40repo. I got an immediate answer though
  74. 2:42because clawed code never needed to see
  75. 2:44that. The voice agent can handle the
  76. 2:46easy stuff itself, which is great
  77. 2:48because I don't want every random
  78. 2:50question turning into a big task or
  79. 2:52racking up my API credits for anthropic.
  80. 2:54But now I'm going to give it something
  81. 2:56that actually takes a bit more time. The
  82. 2:58code repo that I gave it is just
  83. 3:00TypeScript practice problems. I could
  84. 3:02have chose something harder, but let's
  85. 3:04just keep it straightforward for now.
  86. 3:05I'm going to say here, refactor record
  87. 3:08problem TS. Make sure it works
  88. 3:10correctly. Now watch this.
  89. 3:13The task leaves the voice agent. Claude
  90. 3:16code picks it up and this work card
  91. 3:18appears. It's the proof that Claude
  92. 3:20actually has the job now. And normally
  93. 3:22this is where a voice coding setup could
  94. 3:24fall apart. You've asked it to do
  95. 3:26something real. So now you have to wait.
  96. 3:29Except here I don't have to. I can say
  97. 3:32also can you go make me a normal JS file
  98. 3:35with an async function to handle HTTP
  99. 3:37requests?
  100. 3:40It answers and does both while they're
  101. 3:43both running side by side. Claude is
  102. 3:45still working. Now I'm going to fire out
  103. 3:46a random question. Also, what is the
  104. 3:49weather like in Dubai?
  105. 3:52Okay, I can read it from here. Thanks.
  106. 3:55Another answer. Same conversation. And I
  107. 3:58can even ask about the task that is
  108. 4:00being worked on without having to wait
  109. 4:02for it. Hey, how's the refactor going?
  110. 4:06Don't wait for it to finish, just the
  111. 4:07status.
  112. 4:11Now it can tell me Claude is editing,
  113. 4:13running tests, or still working. There
  114. 4:15are basically two things happening at
  115. 4:16once. I'm talking to one agent, another
  116. 4:19agent is doing the work, and neither one
  117. 4:21has to stop because the other one is
  118. 4:23busy, which means while Claude finishes,
  119. 4:26I can just keep going. The coding task
  120. 4:28finished, the result came back into the
  121. 4:30live conversation, and I didn't have to
  122. 4:32check anything. I haven't seen an open-
  123. 4:35source desktop runtime handle that this
  124. 4:37cleanly. So underneath all this, the
  125. 4:39setup is actually pretty simple. On one
  126. 4:41side, we've got the voice agent. Okay,
  127. 4:44great. That handles conversation,
  128. 4:46interruptions, status, cancelling, a lot
  129. 4:48of the fast stuff. On the other side,
  130. 4:51clogged code is connected over ACP.
  131. 4:54That's the side touching files, running
  132. 4:56tools, and doing the slow work. Now,
  133. 4:58mechanically, it's really two lanes
  134. 5:00here. And once you see that, this design
  135. 5:02is a bit more obvious. There's a
  136. 5:04lightweight front-end voice agent whose
  137. 5:06only job is to hold the conversation. It
  138. 5:08answers the easy questions immediately.
  139. 5:10What's the status? Cancel that. Never
  140. 5:12mind. Then there is your real coding
  141. 5:14agent on the back end and they talk over
  142. 5:16the agent client protocol, the ACP. Real
  143. 5:20work gets handed over as an async task
  144. 5:22and runs on its own. When it finishes,
  145. 5:24the result gets injected back into the
  146. 5:26live conversation. So the voice layer
  147. 5:29never blocks on the slow thing. And
  148. 5:32because it's a protocol rather than an
  149. 5:34integration, the backend is swappable.
  150. 5:36Clawed Code, Codeex, Open Code, and a
  151. 5:39bunch of others. Now, one detail I want
  152. 5:41to flag because it shows someone was
  153. 5:43actually paying attention to this
  154. 5:44interruption isn't handled as an event
  155. 5:46here. It's a state machine. Detect user
  156. 5:49speech, mark an interruption, admit an
  157. 5:51interpreted signal, actively surpass any
  158. 5:53audio and transcripts still in flight,
  159. 5:56then start the new term. Their own
  160. 5:58design note says the engineering quality
  161. 6:00of interruption determines the
  162. 6:02reputation. And they're right. The thing
  163. 6:04that makes a voice assistant feel broken
  164. 6:06isn't slow responses. It's interrupting
  165. 6:09it and still hearing another second and
  166. 6:10a half of the former sentence. That kind
  167. 6:13of gap is actually rather annoying and a
  168. 6:15big gap. This has a slight one
  169. 6:17sometimes. Now, where does all this sit
  170. 6:20next to what we already know? Well, the
  171. 6:22OpenAI real-time API is the layer
  172. 6:24underneath this. It's the thing that
  173. 6:26does speech in speech out. This runs on
  174. 6:29top and is provider pluggable. Pycat and
  175. 6:32LiveKit agents are frameworks. They give
  176. 6:35you the parts and you assemble the
  177. 6:37pipeline. This is assembled and pointed
  178. 6:40at a specific job. And to whisper to LLM
  179. 6:43to TTS pipeline you wrote yourself is
  180. 6:45the thing that a lot of us have, a lot
  181. 6:47of us build out. Now, let's get real
  182. 6:49here for a second because there are two
  183. 6:50things here that the readme will not
  184. 6:52lead with. First, the runtime is open
  185. 6:54source. The voice is not. The code is
  186. 6:57Apache 2.0, but the default path, the
  187. 7:00best sounding one, roots through
  188. 7:02Alibaba's Dash scope with a paid API
  189. 7:04key. You get some free credits, but it's
  190. 7:07still paid. The optimized real-time
  191. 7:09voice is cloud API only. There is no
  192. 7:11self-hosted version of the good audio
  193. 7:13generation. So, open- source voice agent
  194. 7:16is only partially true here. There is a
  195. 7:18real escape hatch, and I'll give them
  196. 7:20credit for documenting it. a fully local
  197. 7:23mode using a hugging face speechtoext
  198. 7:25pipeline with an MLX backend on Apple
  199. 7:27silicone running on MPS. In full local
  200. 7:31mode, you need no cloud key at all. It's
  201. 7:34a second Python install and the tune
  202. 7:36profile in the docs is Chinese only.
  203. 7:38Okay, so translate it. But the path
  204. 7:40exists and it's pointed at max
  205. 7:42specifically. Second thing, and this one
  206. 7:45matters more, there are no published
  207. 7:46latency numbers anywhere on any hardware
  208. 7:50I went looking. Okay, now a final quick
  209. 7:52points here. It's Chinese first. The
  210. 7:54wake word is this, which translates to
  211. 7:57this. The main design article explaining
  212. 8:00the full duplex architecture is Chinese
  213. 8:02only. The tuned voice profile is
  214. 8:04Chinese. The model underneath claims a
  215. 8:06100 plus languages for recognition and a
  216. 8:08few dozen for speech. So, I had to
  217. 8:11translate the page to understand the
  218. 8:12gist of it when I was signing up for all
  219. 8:14this. In version 2.0 know is already
  220. 8:16listed in development with the
  221. 8:18architecture being rewritten and the
  222. 8:20stars are running ahead of the usage. 23
  223. 8:232400 stars here but around 4,000 mpm
  224. 8:25downloads a month and on the latest
  225. 8:27release 92 downloaded the Mac version
  226. 8:30build against 502 on Windows. Now, if
  227. 8:33you already work inside Cloud Code or
  228. 8:35Codeex, which I assume most of you all
  229. 8:37do, and you want it to stop being
  230. 8:39chained to the keyboard while long tasks
  231. 8:41run, this is the only thing I found
  232. 8:43doing the parallel task plus line
  233. 8:45conversation together, or at least the
  234. 8:47one that does it really well. Get the
  235. 8:49DMG. It's a signed universal build. No
  236. 8:52CUDA, no Python, nothing to compile. It
  237. 8:54takes about 10, maybe 15 minutes to get
  238. 8:56everything integrated. If you want a
  239. 8:58fully open, fully local voice stack,
  240. 9:00this isn't it. But the MLX path is a
  241. 9:02real starting point and it's more than
  242. 9:04most projects offer. And the idea worth
  243. 9:06taking away whatever happens to this
  244. 9:07particular repo is the thing that makes
  245. 9:09voice usable for real work isn't better
  246. 9:12transcription. Transcription got solved.
  247. 9:14It's whether the system can keep talking
  248. 9:15to you while it's busy. That's a
  249. 9:17concurrency problem, not a speech
  250. 9:19problem. And almost nobody is treating
  251. 9:21it like one. I'm Josh from Better Stack.
  252. 9:23If you enjoy coding tools like this, be
  253. 9:25sure to subscribe. We'll see you in
  254. 9:27another video.

About this transcript

This page contains the full transcript of This Is What Voice Coding Was Supposed to Be (Qwen Audio Agent) by Better Stack, generated from the public captions YouTube serves with the video. The transcript has 1,714 words across 254 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.