How to USE Claude Code for FREE with Ollama ( Local AI FULL Tutorial) — Transcript
Full transcript
- 0:00Whether you have a high-end PC, a
- 0:03powerful laptop running Windows 11,
- 0:05Linux, or even a MacBook with Apple
- 0:08silicon,
- 0:10you can start using Claude Code with
- 0:12Ollama's local and cloud models. For
- 0:15this tutorial, we will be installing
- 0:18Claude Code on Linux.
- 0:28The first, let's talk about my hardware
- 0:30specs. I'm running this on my HP gaming
- 0:33laptop. It's decent, but definitely not
- 0:36a supercomputer. I'm mentioning this
- 0:38because AI performance depends heavily
- 0:41on your hardware. You will need at least
- 0:4416 GB of RAM and Nvidia GPU with 4 GB of
- 0:48VRAM. If your system meets these
- 0:51requirements, you can run models like
- 0:53Qwen 3.5 and GLM 4.7 Flash locally
- 0:58without any major issues.
- 1:01If your laptop struggles with local
- 1:03models, don't worry. I will share a free
- 1:06cloud-based solution later in the video.
- 1:13As you can see, I have also installed
- 1:15proprietary Nvidia drivers to enable
- 1:18partial GPU acceleration.
- 1:23Since we are running Claude Code with
- 1:25Ollama, the first step is to install
- 1:28Ollama on your system.
- 1:34Go ahead and open your web browser. Go
- 1:37to official website, copy the install
- 1:39command, and paste it into your
- 1:41terminal. If you're on a Mac, simply
- 1:45download the DMG file, drag the icon
- 1:47into applications folder, and open
- 1:49Ollama. It will start running in the
- 1:52background.
- 1:54The process is similar on Windows.
- 2:00All right, once Ollama is installed, as
- 2:02you can see, it has detected the Nvidia
- 2:05GPU. This means you can run compatible
- 2:08LLMs with full GPU support. Now, let's
- 2:12verify that Ollama is running.
- 2:16Enter this command to check its status.
- 2:19You will notice that the default context
- 2:21window is set to 4K.
- 2:31By default, Ollama runs on a local port.
- 2:34Open this URL in your browser to confirm
- 2:37that Ollama server is working properly.
- 2:44Next, it's time to install Claude Code.
- 2:47This is an AI coding assistant that can
- 2:49write, understand, and fix code for you.
- 2:52You can install Claude Code based on
- 2:55your operating system. On Linux and
- 2:57macOS, it's just a simple one-line
- 3:00command.
- 3:01On Windows, you will need to run the
- 3:03relevant command inside PowerShell.
- 3:07I will copy this command and paste it
- 3:09into the terminal.
- 3:13This will start the installation
- 3:14process.
- 3:19Once it's done, you will need to run
- 3:22this command to add Claude Code to your
- 3:24system's environment variables.
- 3:30Now, you can verify that Claude is
- 3:32installed by running this command. As
- 3:35you can see, it shows the version, which
- 3:38means Claude is successfully installed.
- 3:43You can run Claude directly to use
- 3:45Anthropic models like Opus if you have
- 3:48an account.
- 3:49But for now, we will stick with Ollama.
- 3:52To exit the TUI, press control plus C
- 3:55twice.
- 3:58Next, it's time to run Claude Code with
- 4:00local Ollama models. This means you will
- 4:03be using your own hardware to power the
- 4:06AI.
- 4:07The first, we need to download a model
- 4:09that's both powerful and efficient.
- 4:12Let's go to Ollama's model library,
- 4:15apply the tool link and thinking
- 4:17filters, and choose Qwen 3.5.
- 4:21It's a great balance of performance and
- 4:23efficiency, especially for low-end
- 4:25systems like mine.
- 4:28If you have more RAM and VRAM, you can
- 4:30try models like Qwen 3 Coder and GLM
- 4:34Flash 4.7.
- 4:36But for most cases, Qwen 2.5 or 3.5 is
- 4:39the great choice.
- 4:41Now, go ahead and copy this command and
- 4:43run this inside terminal to download the
- 4:46model.
- 4:50This will install the 6 billion
- 4:52parameter version with a large context
- 4:54window.
- 5:01Once installed, let's test it in REPL.
- 5:05As you can see, it runs decently on my
- 5:07system using both CPU and partial GPU
- 5:11acceleration.
- 5:13Now, by default, Ollama intelligently
- 5:16splits the workload so performance stays
- 5:18stable.
- 5:19Now, go ahead and press control plus C
- 5:22to stop the model from generating the
- 5:23output and exit the REPL environment by
- 5:26pressing control plus D.
- 5:29Now, by default, the context window is
- 5:31limited, which isn't ideal for Claude
- 5:34Code. Claude requires minimum of 64K
- 5:37context window.
- 5:39We will increase it by creating a custom
- 5:42model with a larger context size.
- 5:45Now, go ahead and create a model file
- 5:47using this command.
- 5:49Then set the base model and context size
- 5:52of 64K, then save it.
- 6:02After that, run the command to create a
- 6:05brand new custom model with any name of
- 6:07your choice.
- 6:17You can verify it using Ollama list.
- 6:22Now, go ahead and run this brand new
- 6:24model. If it loads successfully, your
- 6:27system can handle the large context.
- 6:35You can also confirm by running this
- 6:37command, and you can see it uses the
- 6:40larger context window.
- 6:42Now, let's connect this model to Claude
- 6:44Code. Just go ahead and create a folder
- 6:47for your workspace and navigate into it.
- 6:54Then launch Claude Code using your
- 6:57custom model.
- 6:59Now, simply type the brand new Ollama
- 7:01launch command with the model name.
- 7:09You will see the Claude setup screen.
- 7:12Just go ahead and choose your theme and
- 7:14trust the workspace. And just like that,
- 7:16Qwen 3.5 is now running locally with
- 7:20Claude Code.
- 7:27Now, let's test it with a simple prompt.
- 7:48As expected, responses are slower since
- 7:52everything is running locally, but it
- 7:54works perfectly.
- 8:01Now, for something more practical, let's
- 8:04ask it to create a simple C program and
- 8:07save it to a file.
- 8:18As you can see, it successfully uses
- 8:21tool calling to create and manage files.
- 8:29We can even ask it to run and test the
- 8:32program, and it works perfectly.
- 8:53And lastly, let's try generating a shell
- 8:56script to create multiple folders.
- 9:19And there we go. It completes the task
- 9:21without any issues.
- 9:23And that's it. Now, you have Claude Code
- 9:26running locally with Ollama.
- 9:31Next, it's time to show you how to use
- 9:33Claude Code with Qwen 3.5 in VS Code to
- 9:37build a simple static website. The
- 9:39first, create a brand new workspace and
- 9:42open it inside VS Code.
- 9:50Then open terminal inside VS Code and
- 9:53launch Claude using the Qwen 3.5 model.
- 10:01I will keep accept edits turned on. This
- 10:05way, it can automatically apply changes.
- 10:08Now, let's provide a prompt to create a
- 10:10simple Hello World website using HTML
- 10:14and CSS.
- 10:29After a few minutes, it generates a
- 10:31complete HTML file with all the content.
- 10:38As you can see, this is the code it
- 10:41wrote.
- 10:45And here's the website it created. It
- 10:48actually looks very good. Since I'm
- 10:50using a low-end system, the responses
- 10:52are a bit slower. But on higher-end
- 10:55machines like Apple Silicon with 64 GB
- 10:57of unified memory, the generation speed
- 11:00will be much faster.
- 11:02If your computer doesn't have at least
- 11:0516 GB of RAM and 4 GB of VRAM, you may
- 11:08not be able to run Ollama models
- 11:11locally. In that case, you can use
- 11:13Ollama's cloud models. They offer a free
- 11:16tier allowing you to run powerful LLMs
- 11:20with high-end GPUs. These are the cloud
- 11:23models available for free. Now, for this
- 11:25video, I will be using Kimik 8 2.5 cloud
- 11:29edition, which works great in most
- 11:32scenarios.
- 11:33Just go ahead and run this simple
- 11:35command, and it will prompt you to open
- 11:38a web browser and sign in into your
- 11:40Ollama account. Log in using your email
- 11:43or Google account, then connect your
- 11:46device.
- 11:51Once it's done, return to the terminal.
- 11:55Now, you can use Cloud Code with the
- 11:57Kimik 8 2.5 cloud edition. It's very
- 12:00good for agent-based task.
- 12:06Now, let's test it with the prompt.
- 12:11And as you can see, the response is
- 12:13super fast.
- 12:44You can exit the interface anytime using
- 12:47the exit option.
- 12:52Now, let's use Kimik 8 2.5 in VS Code to
- 12:56create a simple to-do list application.
- 13:06It completes the task in 30 to 40
- 13:09seconds, which is insanely fast.
- 13:15And as you can see, the website looks
- 13:17great.
- 13:25And that's pretty much it. This is how
- 13:27you can run Cloud Code using Ollama
- 13:30models.
- 13:31Let me know what you think about this in
- 13:33the comment section down below.
- 13:35If you like this video, hit that like
- 13:37button and subscribe to see more
- 13:39content. Thank you so much for watching.
- 13:42This has been KSK Royal. I will see you
- 13:45in the next one.
About this transcript
This page contains the full transcript of How to USE Claude Code for FREE with Ollama ( Local AI FULL Tutorial) by Ksk Royal, generated from the public captions YouTube serves with the video. The transcript has 1,354 words across 222 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.