YouTube2Text

Turn Claude into a Faceless Video Engine — Transcript

by The AI Garage · 1,757 words · 279 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Most AI generated faceless videos have
  2. 0:02the same problem. The character keeps
  3. 0:04changing, the setting feels
  4. 0:06inconsistent, and the visuals look like
  5. 0:08random images placed over a voice over.
  6. 0:11With the faceless video engine we're
  7. 0:13going to build today, those problems are
  8. 0:15much easier to avoid. Here is an example
  9. 0:18of what the engine can create. As you
  10. 0:21can see, the character setting, color
  11. 0:23palette, and visual style stay
  12. 0:25consistent. The key idea is simple.
  13. 0:29Build the visual system first, then
  14. 0:31create the scene prompts. Here is the
  15. 0:33process. First, Claude develops the
  16. 0:36story and creates a consistent visual
  17. 0:38system. Then, it turns the script into a
  18. 0:41production board with the image and
  19. 0:43motion prompts for every scene. We use
  20. 0:46that plan to generate the visuals,
  21. 0:48animate the scenes that need movement,
  22. 0:50and assemble everything with voice over
  23. 0:52and music. At the end, I'll show you the
  24. 0:55finished video created with this exact
  25. 0:57workflow. If you are into practical AI
  26. 1:00video workflows like this, subscribe for
  27. 1:02more. I share the complete process and
  28. 1:05every prompt I use. Let's start inside
  29. 1:08Claude. First, I open Claude, create a
  30. 1:11new project, and name it late night
  31. 1:13scroll video engine. Inside project
  32. 1:16instructions, I paste the instructions
  33. 1:18from the prompts file. These stay active
  34. 1:21throughout the project, so I do not need
  35. 1:23to repeat the same rules in every chat.
  36. 1:26They tell Claude to use reliable
  37. 1:28sources, avoid unsupported claims,
  38. 1:30separate evidence from interpretation,
  39. 1:33and keep the narration conversational.
  40. 1:36All the instructions and prompts I use
  41. 1:37are available in the description. Next,
  42. 1:40I start a new chat and paste the web
  43. 1:42research angles and script prompt. At
  44. 1:45the top is a placeholder. For this
  45. 1:47demonstration, I replace it with late
  46. 1:49night scrolling, but you can use any
  47. 1:51subject. Claude searches the web,
  48. 1:54summarizes the strongest findings,
  49. 1:56explains important limitations, and
  50. 1:58provides citations so I can check the
  51. 2:01factual foundation before turning the
  52. 2:03research into a story. Then Claude
  53. 2:05suggests three story angles. The first
  54. 2:08is your brain doesn't know the
  55. 2:10difference between a slot machine and
  56. 2:12your phone. The second is the real
  57. 2:15reason you won't go to sleep, it's not
  58. 2:17the phone. The third is shown on screen.
  59. 2:21Each angle includes the central idea,
  60. 2:23hook, main points, and takeaway. I
  61. 2:27choose the first by typing angle one.
  62. 2:30Claude creates a five-part outline, the
  63. 2:32hook, the machine behind the scroll,
  64. 2:35your brain at night, what the research
  65. 2:38actually shows, and breaking the loop.
  66. 2:41Then it writes the full narration using
  67. 2:43the cited research. I review the
  68. 2:45structure and claims. Type approved and
  69. 2:48Claude saves the final narration as
  70. 2:50final script.txt
  71. 2:52inside the project. Now the story is
  72. 2:55locked before we start designing
  73. 2:57visuals. Next, I return to the Claude
  74. 3:00project, start a new chat, and paste the
  75. 3:02visual system and production board
  76. 3:04prompt. Claude first defines the visual
  77. 3:07system, then uses those rules to build
  78. 3:09the production board. That order matters
  79. 3:12because we decide what the video should
  80. 3:14look like before asking for dozens of
  81. 3:16individual images. When Claude finishes,
  82. 3:20I open the generated HTML file in the
  83. 3:23right panel. At the top is our title.
  84. 3:26Your brain doesn't know the difference
  85. 3:28between a slot machine and your phone.
  86. 3:30Now I expand the visual system summary.
  87. 3:33Our recurring character is Knox, a
  88. 3:36genderneutral flat vector silhouette
  89. 3:38with no face, fixed proportions, and the
  90. 3:41same charcoal navy body in every scene.
  91. 3:44The summary defines the style,
  92. 3:46environment, pallet, camera language,
  93. 3:48and scene mix. Setting the pallet first
  94. 3:51helps avoid the random AI image look and
  95. 3:54makes the video feel intentionally
  96. 3:56designed. Claude plans roughly 35%
  97. 4:00character scenes, 40% diagrams, 15%
  98. 4:04metaphors, and 10% typography.
  99. 4:07Because this story explains a mechanism,
  100. 4:09diagrams get the largest share, while
  101. 4:11the mix keeps the visuals varied. It
  102. 4:14also gives us character and environment
  103. 4:16reference prompts. Their essential
  104. 4:18details are carried into the relevant
  105. 4:20scene prompts, making consistency easier
  106. 4:22later. Each scene card includes the
  107. 4:25voice over, image prompt, motion prompt,
  108. 4:28and a marked produced checkbox. Scene
  109. 4:31one is a character scene, while for
  110. 4:32example, scene 4 is a diagram scene. At
  111. 4:36the top, filters let me isolate
  112. 4:38character, metaphor, diagram, or
  113. 4:41typography scenes. Before I open an
  114. 4:44image generator, the visual rules,
  115. 4:46prompts, scene categories, and
  116. 4:48production tracking are already
  117. 4:50connected. Now I open art to create the
  118. 4:53visuals. One reason I use open art is
  119. 4:56that the latest image and video
  120. 4:57generators are available in one place.
  121. 5:00Under image you can see options such as
  122. 5:03chat GPT nano banana crad
  123. 5:08and gro. For this video I'm using GPT
  124. 5:11image 2. [music] I return to Claude and
  125. 5:14copy Knox's character reference prompt
  126. 5:16from the visual system summary. The
  127. 5:18identity rules are simple. Genderneutral
  128. 5:21flat vector silhouette, no face, fixed
  129. 5:24proportions, and the same charcoal navy
  130. 5:27body. I paste the prompt into open art
  131. 5:30with four generations selected. Even
  132. 5:32when the differences are subtle, several
  133. 5:35options help me choose the best pose,
  134. 5:37proportions, and framing for a reusable
  135. 5:39reference. I generate four clean
  136. 5:42options, compare them, and choose the
  137. 5:44strongest Knox image to use as the
  138. 5:46reference whenever the character
  139. 5:47appears.
  140. 5:49For this demonstration, I'm creating the
  141. 5:51first six scenes. They already include
  142. 5:54character and diagram visuals so we can
  143. 5:56see how the same system handles
  144. 5:58different scene types without producing
  145. 6:00the entire video on screen. For scene
  146. 6:03one, I click the plus icon in Open Art
  147. 6:06and select the Knox image as the
  148. 6:08character reference. [music] Then I copy
  149. 6:10scene 1's exact image prompt from
  150. 6:12Claude, paste it into Open Art and
  151. 6:15generate. While scene one is running, I
  152. 6:18start scene two using the same process.
  153. 6:21When scene one finishes, I compare the
  154. 6:23options and choose the third result
  155. 6:25because the full body is clearly visible
  156. 6:27and the composition works best. For
  157. 6:30scene two, I also choose the third
  158. 6:32option because the clock is easiest to
  159. 6:34read. This is why multiple variations
  160. 6:37help. I am choosing the result that
  161. 6:39communicates the scene most clearly, not
  162. 6:42simply the prettiest image. Then I
  163. 6:44continue through the remaining scenes
  164. 6:46with the same workflow. Copy the prompt
  165. 6:48from Claude. Generate several options
  166. 6:51and select the version that communicates
  167. 6:53the shot best. Scene four is a diagram
  168. 6:56scene. I copy its diagram prompt and
  169. 6:59generate the options. You can see the
  170. 7:02character bodies being used as figures
  171. 7:03inside the diagram while the same pallet
  172. 7:06and line style remain consistent. I
  173. 7:09continue through scene six, checking
  174. 7:11that the character, pallet, and overall
  175. 7:13style still feel connected. These six
  176. 7:16scenes demonstrate the full production
  177. 7:18loop. The remaining cards follow the
  178. 7:20same process. This is why I'm using Open
  179. 7:23Art here. [music] I can move from image
  180. 7:25generation into video generation in the
  181. 7:27same workflow. My affiliate link is
  182. 7:30below if Open Art fits your process.
  183. 7:33Now, let's add motion. I open the
  184. 7:36selected scene 1 image in Open Art and
  185. 7:38choose video reference. Then, I return
  186. 7:41to Claude, copy scene 1's motion prompt,
  187. 7:44paste it into Open Art, and generate.
  188. 7:47While that clip is processing, I repeat
  189. 7:49the same workflow for the other scenes I
  190. 7:51want to animate. The still image already
  191. 7:54establishes the character, composition,
  192. 7:56pallet, and illustration style. The
  193. 7:59motion prompt only describes how that
  194. 8:02existing scene should move, which helps
  195. 8:04reduce unwanted changes.
  196. 8:06After three animated scenes are ready, I
  197. 8:08open the first result. The movement
  198. 8:11makes the shot feel more cinematic
  199. 8:13without losing the original
  200. 8:14illustration. Then I compare the
  201. 8:16animated scenes together. The bigger
  202. 8:19advantage is consistency. Knox still
  203. 8:22reads as the same character and the
  204. 8:24pallet, line style, and overall visual
  205. 8:26language stay connected from shot to
  206. 8:28shot. That consistency matters for
  207. 8:31retention because the viewer can focus
  208. 8:33on the story instead of adjusting to a
  209. 8:35completely different visual style every
  210. 8:37few seconds. The six demonstration
  211. 8:40scenes are ready. So now I put the
  212. 8:42sample together in Cap Cut. You can use
  213. 8:44any editor you prefer. First, I open
  214. 8:48Google AI Studio. I already selected a
  215. 8:51casual middle pitch voice because it
  216. 8:53works well for this storytelling style.
  217. 8:56I copy the narration for the first six
  218. 8:58scenes. Generate the voice over and
  219. 9:00download it. Next, I open YouTube
  220. 9:03Studios audio library for background
  221. 9:06music. I filter by cinematic for genre
  222. 9:09and calm for mood. Listen to a few
  223. 9:11tracks and download the one that best
  224. 9:13matches the late night atmosphere.
  225. 9:16Back in Cap Cut, the voice over goes on
  226. 9:18the timeline first [music] because it
  227. 9:20gives me the structure for the edit.
  228. 9:23Then I import the six animated clips and
  229. 9:25keep Claude's production board open
  230. 9:27beside the editor as the assembly guide
  231. 9:30so I can match each visual to the
  232. 9:31correct narration. I place each scene
  233. 9:34under the matching narration and adjust
  234. 9:36its duration. If a clip is slightly
  235. 9:39short, I slow it down. With these clean
  236. 9:42line-based visuals, subtle, slower
  237. 9:45motion still looks natural and gives the
  238. 9:47viewer time to read the frame. I keep
  239. 9:49the edit simple with clean hard cuts
  240. 9:52instead of unnecessary transitions.
  241. 9:55Then I add the music on a separate track
  242. 9:57and lower the volume so it supports the
  243. 9:59atmosphere without competing with the
  244. 10:01narration. After a final audio and
  245. 10:03visual check, I export the sample. Now,
  246. 10:06let's watch the finished video. [music]
  247. 10:09You told yourself five more minutes.
  248. 10:12That was 40 minutes ago. You're not
  249. 10:15stupid and you're not weak. [music]
  250. 10:17You know you have to be up in 6 hours.
  251. 10:20And yet here you are, thumb moving, eyes
  252. 10:24locked on a screen that's somehow more
  253. 10:26interesting than [music] sleep. If this
  254. 10:28happens to you more nights than you'd
  255. 10:30like to admit, you're in very [music]
  256. 10:32crowded company. Surveys of American
  257. 10:35adults have found that the vast majority
  258. 10:37of people sleep [music] with their phone
  259. 10:38in the bedroom. And a large share say
  260. 10:41they're browsing the internet or
  261. 10:42checking apps within 10 minutes of
  262. 10:44actually falling asleep. [music]
  263. 10:46So before we go further, let's drop the
  264. 10:49idea that this is some [music] rare
  265. 10:51personal failing. It's closer to a
  266. 10:53universal experience. If you want me to
  267. 10:56continue this series with a different
  268. 10:58niche or format, drop the word continue
  269. 11:01plus your niche in the comments. It
  270. 11:03could be fitness, history, business,
  271. 11:06relationships, anything you want to see
  272. 11:08turned into a complete video engine,
  273. 11:10drop it below. And one more thing, this
  274. 11:14workflow isn't limited to this type of
  275. 11:15faceless video. Click the video on your
  276. 11:18screen now to see me apply the same
  277. 11:20video engine idea in a completely
  278. 11:22different way. People really like that
  279. 11:25one. I'll see you there.

About this transcript

This page contains the full transcript of Turn Claude into a Faceless Video Engine by The AI Garage, generated from the public captions YouTube serves with the video. The transcript has 1,757 words across 279 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.