YouTube2Text

How to USE Claude Code for FREE with Ollama ( Local AI FULL Tutorial) — Transcript

by Ksk Royal · 1,354 words · 222 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Whether you have a high-end PC, a
  2. 0:03powerful laptop running Windows 11,
  3. 0:05Linux, or even a MacBook with Apple
  4. 0:08silicon,
  5. 0:10you can start using Claude Code with
  6. 0:12Ollama's local and cloud models. For
  7. 0:15this tutorial, we will be installing
  8. 0:18Claude Code on Linux.
  9. 0:28The first, let's talk about my hardware
  10. 0:30specs. I'm running this on my HP gaming
  11. 0:33laptop. It's decent, but definitely not
  12. 0:36a supercomputer. I'm mentioning this
  13. 0:38because AI performance depends heavily
  14. 0:41on your hardware. You will need at least
  15. 0:4416 GB of RAM and Nvidia GPU with 4 GB of
  16. 0:48VRAM. If your system meets these
  17. 0:51requirements, you can run models like
  18. 0:53Qwen 3.5 and GLM 4.7 Flash locally
  19. 0:58without any major issues.
  20. 1:01If your laptop struggles with local
  21. 1:03models, don't worry. I will share a free
  22. 1:06cloud-based solution later in the video.
  23. 1:13As you can see, I have also installed
  24. 1:15proprietary Nvidia drivers to enable
  25. 1:18partial GPU acceleration.
  26. 1:23Since we are running Claude Code with
  27. 1:25Ollama, the first step is to install
  28. 1:28Ollama on your system.
  29. 1:34Go ahead and open your web browser. Go
  30. 1:37to official website, copy the install
  31. 1:39command, and paste it into your
  32. 1:41terminal. If you're on a Mac, simply
  33. 1:45download the DMG file, drag the icon
  34. 1:47into applications folder, and open
  35. 1:49Ollama. It will start running in the
  36. 1:52background.
  37. 1:54The process is similar on Windows.
  38. 2:00All right, once Ollama is installed, as
  39. 2:02you can see, it has detected the Nvidia
  40. 2:05GPU. This means you can run compatible
  41. 2:08LLMs with full GPU support. Now, let's
  42. 2:12verify that Ollama is running.
  43. 2:16Enter this command to check its status.
  44. 2:19You will notice that the default context
  45. 2:21window is set to 4K.
  46. 2:31By default, Ollama runs on a local port.
  47. 2:34Open this URL in your browser to confirm
  48. 2:37that Ollama server is working properly.
  49. 2:44Next, it's time to install Claude Code.
  50. 2:47This is an AI coding assistant that can
  51. 2:49write, understand, and fix code for you.
  52. 2:52You can install Claude Code based on
  53. 2:55your operating system. On Linux and
  54. 2:57macOS, it's just a simple one-line
  55. 3:00command.
  56. 3:01On Windows, you will need to run the
  57. 3:03relevant command inside PowerShell.
  58. 3:07I will copy this command and paste it
  59. 3:09into the terminal.
  60. 3:13This will start the installation
  61. 3:14process.
  62. 3:19Once it's done, you will need to run
  63. 3:22this command to add Claude Code to your
  64. 3:24system's environment variables.
  65. 3:30Now, you can verify that Claude is
  66. 3:32installed by running this command. As
  67. 3:35you can see, it shows the version, which
  68. 3:38means Claude is successfully installed.
  69. 3:43You can run Claude directly to use
  70. 3:45Anthropic models like Opus if you have
  71. 3:48an account.
  72. 3:49But for now, we will stick with Ollama.
  73. 3:52To exit the TUI, press control plus C
  74. 3:55twice.
  75. 3:58Next, it's time to run Claude Code with
  76. 4:00local Ollama models. This means you will
  77. 4:03be using your own hardware to power the
  78. 4:06AI.
  79. 4:07The first, we need to download a model
  80. 4:09that's both powerful and efficient.
  81. 4:12Let's go to Ollama's model library,
  82. 4:15apply the tool link and thinking
  83. 4:17filters, and choose Qwen 3.5.
  84. 4:21It's a great balance of performance and
  85. 4:23efficiency, especially for low-end
  86. 4:25systems like mine.
  87. 4:28If you have more RAM and VRAM, you can
  88. 4:30try models like Qwen 3 Coder and GLM
  89. 4:34Flash 4.7.
  90. 4:36But for most cases, Qwen 2.5 or 3.5 is
  91. 4:39the great choice.
  92. 4:41Now, go ahead and copy this command and
  93. 4:43run this inside terminal to download the
  94. 4:46model.
  95. 4:50This will install the 6 billion
  96. 4:52parameter version with a large context
  97. 4:54window.
  98. 5:01Once installed, let's test it in REPL.
  99. 5:05As you can see, it runs decently on my
  100. 5:07system using both CPU and partial GPU
  101. 5:11acceleration.
  102. 5:13Now, by default, Ollama intelligently
  103. 5:16splits the workload so performance stays
  104. 5:18stable.
  105. 5:19Now, go ahead and press control plus C
  106. 5:22to stop the model from generating the
  107. 5:23output and exit the REPL environment by
  108. 5:26pressing control plus D.
  109. 5:29Now, by default, the context window is
  110. 5:31limited, which isn't ideal for Claude
  111. 5:34Code. Claude requires minimum of 64K
  112. 5:37context window.
  113. 5:39We will increase it by creating a custom
  114. 5:42model with a larger context size.
  115. 5:45Now, go ahead and create a model file
  116. 5:47using this command.
  117. 5:49Then set the base model and context size
  118. 5:52of 64K, then save it.
  119. 6:02After that, run the command to create a
  120. 6:05brand new custom model with any name of
  121. 6:07your choice.
  122. 6:17You can verify it using Ollama list.
  123. 6:22Now, go ahead and run this brand new
  124. 6:24model. If it loads successfully, your
  125. 6:27system can handle the large context.
  126. 6:35You can also confirm by running this
  127. 6:37command, and you can see it uses the
  128. 6:40larger context window.
  129. 6:42Now, let's connect this model to Claude
  130. 6:44Code. Just go ahead and create a folder
  131. 6:47for your workspace and navigate into it.
  132. 6:54Then launch Claude Code using your
  133. 6:57custom model.
  134. 6:59Now, simply type the brand new Ollama
  135. 7:01launch command with the model name.
  136. 7:09You will see the Claude setup screen.
  137. 7:12Just go ahead and choose your theme and
  138. 7:14trust the workspace. And just like that,
  139. 7:16Qwen 3.5 is now running locally with
  140. 7:20Claude Code.
  141. 7:27Now, let's test it with a simple prompt.
  142. 7:48As expected, responses are slower since
  143. 7:52everything is running locally, but it
  144. 7:54works perfectly.
  145. 8:01Now, for something more practical, let's
  146. 8:04ask it to create a simple C program and
  147. 8:07save it to a file.
  148. 8:18As you can see, it successfully uses
  149. 8:21tool calling to create and manage files.
  150. 8:29We can even ask it to run and test the
  151. 8:32program, and it works perfectly.
  152. 8:53And lastly, let's try generating a shell
  153. 8:56script to create multiple folders.
  154. 9:19And there we go. It completes the task
  155. 9:21without any issues.
  156. 9:23And that's it. Now, you have Claude Code
  157. 9:26running locally with Ollama.
  158. 9:31Next, it's time to show you how to use
  159. 9:33Claude Code with Qwen 3.5 in VS Code to
  160. 9:37build a simple static website. The
  161. 9:39first, create a brand new workspace and
  162. 9:42open it inside VS Code.
  163. 9:50Then open terminal inside VS Code and
  164. 9:53launch Claude using the Qwen 3.5 model.
  165. 10:01I will keep accept edits turned on. This
  166. 10:05way, it can automatically apply changes.
  167. 10:08Now, let's provide a prompt to create a
  168. 10:10simple Hello World website using HTML
  169. 10:14and CSS.
  170. 10:29After a few minutes, it generates a
  171. 10:31complete HTML file with all the content.
  172. 10:38As you can see, this is the code it
  173. 10:41wrote.
  174. 10:45And here's the website it created. It
  175. 10:48actually looks very good. Since I'm
  176. 10:50using a low-end system, the responses
  177. 10:52are a bit slower. But on higher-end
  178. 10:55machines like Apple Silicon with 64 GB
  179. 10:57of unified memory, the generation speed
  180. 11:00will be much faster.
  181. 11:02If your computer doesn't have at least
  182. 11:0516 GB of RAM and 4 GB of VRAM, you may
  183. 11:08not be able to run Ollama models
  184. 11:11locally. In that case, you can use
  185. 11:13Ollama's cloud models. They offer a free
  186. 11:16tier allowing you to run powerful LLMs
  187. 11:20with high-end GPUs. These are the cloud
  188. 11:23models available for free. Now, for this
  189. 11:25video, I will be using Kimik 8 2.5 cloud
  190. 11:29edition, which works great in most
  191. 11:32scenarios.
  192. 11:33Just go ahead and run this simple
  193. 11:35command, and it will prompt you to open
  194. 11:38a web browser and sign in into your
  195. 11:40Ollama account. Log in using your email
  196. 11:43or Google account, then connect your
  197. 11:46device.
  198. 11:51Once it's done, return to the terminal.
  199. 11:55Now, you can use Cloud Code with the
  200. 11:57Kimik 8 2.5 cloud edition. It's very
  201. 12:00good for agent-based task.
  202. 12:06Now, let's test it with the prompt.
  203. 12:11And as you can see, the response is
  204. 12:13super fast.
  205. 12:44You can exit the interface anytime using
  206. 12:47the exit option.
  207. 12:52Now, let's use Kimik 8 2.5 in VS Code to
  208. 12:56create a simple to-do list application.
  209. 13:06It completes the task in 30 to 40
  210. 13:09seconds, which is insanely fast.
  211. 13:15And as you can see, the website looks
  212. 13:17great.
  213. 13:25And that's pretty much it. This is how
  214. 13:27you can run Cloud Code using Ollama
  215. 13:30models.
  216. 13:31Let me know what you think about this in
  217. 13:33the comment section down below.
  218. 13:35If you like this video, hit that like
  219. 13:37button and subscribe to see more
  220. 13:39content. Thank you so much for watching.
  221. 13:42This has been KSK Royal. I will see you
  222. 13:45in the next one.

About this transcript

This page contains the full transcript of How to USE Claude Code for FREE with Ollama ( Local AI FULL Tutorial) by Ksk Royal, generated from the public captions YouTube serves with the video. The transcript has 1,354 words across 222 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.