The Ultimate Local AI Tier List For 2026 — Transcript
Full transcript
- 0:00Today, I'm ranking the most popular
- 0:01local AI use cases for you based on real
- 0:03engineering experience. For each
- 0:05category, I'll recommend the one model
- 0:07or tool that I would start with, so you
- 0:09can skip hours of research and just get
- 0:11going. Personally, I've spent hundreds
- 0:13of hours testing with my RTX 4090, but a
- 0:16lot of cases work on weaker hardware,
- 0:18too. So, with that being said, let's get
- 0:20right into it. First stop, we're going
- 0:22to start with local code auto complete,
- 0:24and this is straight up S-tier.
- 0:27Code auto complete is one of the first
- 0:28ways that we were using AI for automated
- 0:31coding with something like GitHub
- 0:32Copilot. When you type code, the model
- 0:34finishes your line or even fills in a
- 0:36function body before you can think about
- 0:38it. And while agentic coding seems to be
- 0:40taking over the world, code auto
- 0:41complete is still extremely useful. And
- 0:44the great part is that it absolutely
- 0:45works great with local models. One model
- 0:48you could use is something like Qwen 2.5
- 0:50Coder, because it's a small 7 billion
- 0:53parameter model, which can run at
- 0:55sub-100 millisecond latency, even
- 0:57sometimes on GPUs with just a couple
- 0:59gigabytes of VRAM. Now, you're going to
- 1:02sometimes even get results that are
- 1:03faster than network-based auto
- 1:04completes, because there's no round
- 1:06trip. You're just, you know, executing
- 1:08this with local models. And this gives
- 1:11you a similar experience as the code
- 1:12auto complete models that were
- 1:14state-of-the-art a couple of years ago.
- 1:16You can pair your local model with
- 1:18something like Continue Dev and use
- 1:20Ollama for the inference, and then
- 1:21you've got a self-hosted Copilot
- 1:24replacement, at the very least, you
- 1:25know, the auto complete features of it,
- 1:27that costs nothing after hardware. While
- 1:29this doesn't beat the newest agentic
- 1:31features that you get with things like
- 1:33Cloud Code, Copilot, Codex, it's a great
- 1:35start, and this actually works well
- 1:37locally. Local can beat cloud models
- 1:40here on the metric that matters most
- 1:42when it comes to auto complete, which is
- 1:43response speed, and you can fully
- 1:45customize it yourself, as well. We're
- 1:46going to be talking about other coding
- 1:48use cases later, but first, let's move
- 1:49to something different: photo
- 1:51enhancement. This is something that I
- 1:53would comfortably place in A-tier. Photo
- 1:55enhancement covers things like
- 1:57upscaling, face restoration, background
- 1:59removal, noise reduction, colorization,
- 2:02anything basically where you take an
- 2:03existing image and you make it better
- 2:05without generating new content. I'll
- 2:08focus on upscaling for now as this is an
- 2:09example that is very common, but the
- 2:11whole category is strong locally. A tool
- 2:14that you can get started with easily is
- 2:15something like Upscaley. It's a free,
- 2:18open-source desktop app that uses
- 2:19Real-ESRGAN models. You don't have to
- 2:22screw around too much with Python, you
- 2:24don't need the command line, yet no
- 2:25custom GPU configuration. Just to see
- 2:27what's possible, you drag a photo in and
- 2:30you will get a 4x or even 8x upscaled
- 2:32version out of it. From there, because
- 2:35you can see what's possible, you can go
- 2:36ahead and try and play around with the
- 2:38models yourself and create a custom
- 2:40workflow that works for whatever you
- 2:41need to do. It's great to be able to
- 2:43enhance your photos without having to
- 2:45rely on only a cloud-based version of
- 2:47Photoshop. And any dedicated GPU with
- 2:50something like 4GB of free RAM can
- 2:52handle this use case in just seconds.
- 2:55Another great use case is home
- 2:57automation. Home automation covers many
- 3:00different tasks, but the whole category
- 3:02can definitely be A-tier. With home
- 3:05automation, you can have use cases like
- 3:06having your security cameras detect a
- 3:08person in your driveway and your phone
- 3:10getting a notification with a snapshot,
- 3:12which will all be processed on a box in
- 3:14your closet. And the great part about
- 3:16this is that you will not have a cloud
- 3:18subscription anymore and maybe more
- 3:20importantly, your security footage
- 3:22leaving your network to be stored on
- 3:24maybe some random Chinese cloud that you
- 3:26don't know about. Object detection is
- 3:28the headline feature here, but local
- 3:30home automation can also cover voice
- 3:32control, presence sensing, and even
- 3:34energy monitoring. In any case, the
- 3:36stack that you can start with is Frigate
- 3:38NVR plus Home Assistant. This is
- 3:40probably the most mature local AI
- 3:42ecosystem that exists right now,
- 3:44especially for beginners.
- 3:46With this, you can have use cases like
- 3:48person detection, vehicle detection,
- 3:50pets, packages, license plates, and even
- 3:52some basic facial recognition as part of
- 3:55some of the latest Frigate versions. And
- 3:57all of this will be processed entirely
- 3:59on your hardware. Home Assistant has
- 4:01over 2 million active installations, and
- 4:03Frigate has over 30,000 stars on GitHub,
- 4:05so you're not going to be alone. It's
- 4:06not just a hobby project. It will be a
- 4:08full ecosystem that you can plug into.
- 4:10The only thing keeping us from S tier is
- 4:12the setup complexity. You often will
- 4:14have to configure Docker, camera
- 4:15streams, detection zones, and even if
- 4:17it's running, it might break in a couple
- 4:19of weeks because things update fast in
- 4:21the home automation space. Now, you
- 4:23might be already thinking, "This is a
- 4:25nice tier list so far, but how do I get
- 4:27started?" Well, I put together a
- 4:29collection of my own open-source
- 4:30projects for many of the local AI use
- 4:32cases we're covering today. The link is
- 4:34in the description, and you should
- 4:35definitely check it out. A whole
- 4:36different use case is video generation,
- 4:39and this one is a little bit
- 4:40disappointing. I'm going to put it in
- 4:41the C tier because the idea behind video
- 4:44generation is that you can type a prompt
- 4:45and get a full video clip out. Something
- 4:48you can use for product demos,
- 4:49children's video content, or just
- 4:51content visualization. This is the
- 4:52category everyone wants to work locally
- 4:54because cloud video generation is
- 4:56expensive and rate-limited. Problem is
- 4:58though that it's still very expensive to
- 5:00generate locally in terms of the time
- 5:01that it takes, and the quality is just
- 5:03not great. One model you can try is When
- 5:052.1 from Alibaba, which can beat Sora on
- 5:08several benchmarks, which was, you know,
- 5:10a state-of-the-art model a while back.
- 5:13The issue is though that even on my
- 5:141590, I can't really use a full 14
- 5:17billion model. I already have to kind of
- 5:19go down to the smaller 5 billion
- 5:21parameter one, and the quality is just a
- 5:24lot less. In addition, because it takes
- 5:26a long time to generate video, it is
- 5:28also a lot more time expensive to
- 5:30generate many variants to figure out how
- 5:32we can generate a video clip that
- 5:34actually looks good. Compared to image
- 5:36generation, which we'll get to later,
- 5:37you don't just generate one frame. To
- 5:39get a good video, you need to generate
- 5:41many different frames, and that just
- 5:42takes a long time. Locally, this might
- 5:46be usable for some experimental social
- 5:48media clips and concept work, but real
- 5:50professional production is still cloud
- 5:52territory, and that's why this is going
- 5:54in C tier.
- 5:55Speaking of something that's related,
- 5:57image generation, now this works much
- 5:59better. I would put image generation in
- 6:01the S tier. With image generation, you
- 6:03describe an image and the model can
- 6:05create it. You can use it for
- 6:06thumbnails, marketing assets, product
- 6:08mock-ups, and concept art. This can also
- 6:11include in-painting, out-painting,
- 6:13basically editing existing image, and
- 6:15even generating a new image variant
- 6:17based on an image that you're passing to
- 6:19it. Now, I'll focus on text-to-image
- 6:21generation because that's where most
- 6:23people start. The model you can use here
- 6:25is Flux 2 Dev, which you can run through
- 6:27ComfyUI. On something like my 5090, this
- 6:30is a very comfortable use case, and it's
- 6:32very easy for me to generate an image in
- 6:34just a couple of seconds. Unlike video
- 6:36generation, this means that you can
- 6:37generate hundreds of variants in, you
- 6:39know, a short amount of time, which
- 6:41increases the chances that you actually
- 6:43get an image that you're happy with.
- 6:45And in fact, Flux is a really good model
- 6:47because in some blind tests, it achieved
- 6:49a 71% win rate over some of the older
- 6:52Midjourney versions for editorial
- 6:54photorealism. So, the models are really
- 6:56getting much better.
- 6:59And with image generation, I've also
- 7:00found that there's a really great
- 7:01community where folks have fine-tuned a
- 7:04lot of models for custom characters and
- 7:06styles, and there's even a lot of image
- 7:08generation models that have less content
- 7:10filters than the one in the cloud. It's
- 7:12good that cloud has content filters,
- 7:14especially for copyright, but sometimes
- 7:16they go a little bit too far, and they
- 7:17become a bit unusable, to be honest. And
- 7:19actually training a custom LoRA model,
- 7:22or actually fine-tuning one, only takes
- 7:2450 to 20 images, and you can do that on
- 7:26consumer hardware, as well. Now, one
- 7:28thing that I've learned from using AI
- 7:30image generation for my own content is
- 7:31that iterative editing is where it gets
- 7:33a bit tricky.
- 7:35Image properties like gradients and
- 7:36complex textures make it very difficult
- 7:38for those local models to edit images
- 7:41without losing quality. Personally, this
- 7:43is still where I use the latest native
- 7:45been added model or just use Photoshop
- 7:47to just make sure that I can use some of
- 7:49the cloud features. So, even though it's
- 7:50S tier, there are some use cases where
- 7:53it's going to fall behind a little bit,
- 7:54but it is really getting better every
- 7:56single month. A completely different but
- 7:58fun use case is voice agents. Now, this
- 8:00one is something that I'm going to place
- 8:01in C tier because of the wide variety of
- 8:04use cases where it can work in a variety
- 8:07of ways. Basically, the idea behind
- 8:08voice agents is that you can talk to
- 8:10your computer and it will talk back like
- 8:12a local Alexa or customer support bots
- 8:13that runs on your hardware. But ideally,
- 8:15it should also be able to, you know, do
- 8:17actions for you. For example, you can
- 8:19use it to do hands-free home automation.
- 8:22The best open source option right now
- 8:23would be something like Pipe Cat, which
- 8:25can chain together a speech to text
- 8:27model, an LLM, and then a text to speech
- 8:29model into one single pipeline. And that
- 8:32will allow you to get something like sub
- 8:34800 millisecond voice to voice latency
- 8:36on some standard hardware on something
- 8:39like macOS.
- 8:41The problem is that the model size
- 8:43constraints means that the AI's
- 8:44responses locally are noticeably dumber
- 8:47than you can get from the best local
- 8:48models that you interact with over chat,
- 8:50and of course, the best cloud models. In
- 8:53fact, this is even an issue with cloud
- 8:54voice agents. They can achieve 200 to
- 8:56500 millisecond latency, but their
- 8:58intelligence is more around GPT-4 level
- 9:01and not to the latest models that are
- 9:02out there. And of course, you know, this
- 9:05intelligence gap is much higher for
- 9:07local voice agents. That being said, if
- 9:10you just have a singular command that
- 9:11you want to run, voice agents can work
- 9:13pretty well. The main issue where they
- 9:15lack quality is in longer conversations
- 9:17because with local models, you're just
- 9:19much more constrained on your context
- 9:21window, which means that as you have a
- 9:22long conversation with a voice agent, it
- 9:25will simply go off track. And you don't
- 9:27have that issue if your voice agent is
- 9:29just there to process one command and
- 9:31execute it. So, that would be the
- 9:33recommended use case, I would say. Now,
- 9:35one component of voice agents is a text
- 9:37to speech model and you can use text to
- 9:39speech on its own and on its own this is
- 9:41definitely a tier. With text to speech
- 9:43you feed text in and you get natural
- 9:45sounding audio out. You can use this for
- 9:47audiobook narration, voiceover for
- 9:48video, accessibility features in your
- 9:50app, or even trying to clone your own
- 9:52voice for content which I've tried but
- 9:54that one is a bit ambitious it doesn't
- 9:55work very well. Now this category also
- 9:57covers music generation and sound
- 9:59effects because it's not really my
- 10:01specialty. In general though voice
- 10:03synthesis is where local models have
- 10:04made the biggest gains. Now text to
- 10:06speech has had the most dramatic
- 10:08transformation in my opinion of any
- 10:10local AI category in the past 18 months
- 10:12because I would say that it's almost a
- 10:14solved problem at this point especially
- 10:16for the English language. The model that
- 10:18I would start with is Chatterbox from
- 10:20Resemble AI because it beats 11 labs in
- 10:23a couple of blind tests with over 60%
- 10:26listener preference rates. And the base
- 10:28model might be English only but you can
- 10:30even go to Chatterbox multilingual which
- 10:32covers 23 plus languages. And so the gap
- 10:35versus a cloud offering like 11 labs has
- 10:37closed for a lot of use cases and in
- 10:40some cases the gap has gone entirely.
- 10:43There are some gotchas. A lot of models
- 10:45will hallucinate or really degrade in
- 10:46quality past about 1,000 characters but
- 10:50there's a lot of ways to solve that. You
- 10:51know, you can split up your text in
- 10:53multiple batches and just process them
- 10:55one by one. Now if you're finding this
- 10:56tier list useful make sure to hit
- 10:58subscribe because over 90% of you are
- 11:00missing out on the latest AI engineering
- 11:02news that's based on reality and not
- 11:05hype so make sure to join the club. Next
- 11:07we're going to be talking about speech
- 11:08to text which is another component of
- 11:10voice agents that you can extract on its
- 11:12own and it's another solved problem in
- 11:14my opinion. Speech to text is absolutely
- 11:17S tier.
- 11:18You record audio you pass an audio file
- 11:20you can get a text transcript back. For
- 11:23meeting notes, podcast transcription,
- 11:24subtitle generation, or just turning
- 11:27your voice memos into searchable text. I
- 11:30use this constantly myself for
- 11:31processing my own YouTube video content
- 11:34and for English it works very well. A
- 11:36model you can use is faster Whisper with
- 11:38large V3 Turbo which gives you over four
- 11:41times speed over the original Whisper
- 11:42model. And the real workflow that I use
- 11:45is a two-stage pipeline. I use Whisper
- 11:47to get a fast, accurate raw
- 11:49transcription and then I use local LLM
- 11:52to clean up filler words and extract the
- 11:54core meaning. And this way I can store
- 11:56the notes into something like my
- 11:57Obsidian vault to reference later. The
- 11:59transcription itself is almost instant,
- 12:02but the LM cleanup step does add some
- 12:04time. The thing is though, I can just
- 12:05run it in the background, so that's not
- 12:07an issue at all for me.
- 12:09The gap of cloud models is mainly with
- 12:10speaker diarization, which means that
- 12:13you can split up some audio context into
- 12:16the different speakers. So let's say you
- 12:17have a meeting with 10 people, you want
- 12:19to be able to know who's speaking which
- 12:21sentences, right? Cloud models are a
- 12:23little bit better at doing that, but I
- 12:25fully believe that local models will
- 12:26catch up very quickly. The core use case
- 12:28of transcribing audio just works very
- 12:30well nowadays with local models.
- 12:33Next we're going to be talking about
- 12:34OCR, optical character recognition,
- 12:36which I'm going to comfortably place as
- 12:38our first beat here.
- 12:40OCR covers table extraction, formula
- 12:42recognition, and just converting scanned
- 12:43documents into structured data, which is
- 12:45useful for so many use cases. One tool
- 12:48to start with is Suriya, or you can use
- 12:50a Deep Seek OCR model that has been
- 12:51released recently. Now, a lot of the use
- 12:53cases are a little bit more on the
- 12:54boring side, like being able to transfer
- 12:56and process invoices, but it does work
- 12:59pretty well. Next we're going to be
- 13:00talking about agentic coding, which is
- 13:02much more interesting than just doing
- 13:03code autocomplete and seems to be hyped
- 13:05up more and more for local models. Now,
- 13:08the one disappointing thing about
- 13:09agentic coding is that it doesn't work
- 13:11that well unless you have very good
- 13:13hardware. A lot of YouTube videos they
- 13:15show agentic coding and then they show
- 13:17how it can generate Python code and
- 13:18that's good and all, but with agentic
- 13:20coding you really want a model to be
- 13:22competent enough to read your entire
- 13:24codebase, write code, run tests, and
- 13:26iterate until your feature works, the
- 13:28feature that you asked for with text.
- 13:31The issue is though that most absolutely
- 13:33most hardware is too incompetent for
- 13:35agentic coding.
- 13:37Now, I've created many videos on my
- 13:38channel covering how performant agentic
- 13:41coding is on my 5090, and it's
- 13:43definitely getting closer. But, it is
- 13:45nowhere near as useful as using code
- 13:47autocomplete, nor something like an AI
- 13:50chat, because eventually local AI agents
- 13:53choke on larger projects. That being
- 13:55said, this is something that improves
- 13:56every single month. So, definitely check
- 13:59out the videos that I have in my
- 14:00channel, which I'll put in the link in
- 14:01the description below, to get the most
- 14:03out of your local setup, because it's
- 14:04very difficult to set up properly, and
- 14:06you really need a bit of an expert guide
- 14:08to do this right. With agentic coding,
- 14:10you might be able to match something
- 14:11like GPT-4o locally, but Claude Opus 4.6
- 14:14and similar models have really upped the
- 14:16game here, and there is just a huge gap
- 14:18between what you can do locally versus
- 14:19with the state-of-the-art Claude models.
- 14:22Even if your local model is intelligent
- 14:24enough, the problem is is that it will
- 14:26get slow as the context window fills up.
- 14:28And unlike with these other use cases,
- 14:29where you can clear out the context
- 14:31window after you improve small bits of
- 14:34the task, with agentic coding, you
- 14:36cannot clear the context window with
- 14:37every step. You need the coding model to
- 14:40have a good idea of how your code base
- 14:41works. So, in the end, these models are
- 14:44going to work very slowly on your local
- 14:46hardware, especially compared to cloud
- 14:48models, unless you have a very beefy PC.
- 14:50Most people promoting local models with
- 14:52agentic coding are not using it for
- 14:54serious projects. And if you don't
- 14:55believe me, check out some of our
- 14:57masterclasses on this channel to learn
- 14:59the truth. If you do want to get started
- 15:01with agentic coding, I recommend some of
- 15:02the latest Qwen models, but again, I
- 15:04covered that extensively in the videos
- 15:06on my channel already, so go check those
- 15:08out. As models get stronger, I expect it
- 15:10to move into A or S tier, especially as
- 15:13local models will catch up as well. You
- 15:15know what isn't A tier yet, though? AI
- 15:17chats, a traditional use case. With AI
- 15:19chats, you can ask questions, brainstorm
- 15:21ideas, summarize documents, draft
- 15:23emails, the same stuff you would use
- 15:25something like ChatGPT for, but it can
- 15:27run entirely on your machine with no
- 15:29data leaving your network. I would
- 15:30recommend new models like something like
- 15:32Coin 330B through LM Studio. And this is
- 15:36a pretty optimized model that can run
- 15:37pretty fast, but honestly, there's many
- 15:39choices for AI chat models. You can use
- 15:41a newer Mistral model or even check out
- 15:44the older, but still competent
- 15:46open-source OpenAI model that's out
- 15:47there. In any case, locally, you can
- 15:49really mimic what GPT-4 was able to do a
- 15:52year plus ago and now fully locally. You
- 15:55can use it for so many different
- 15:56purposes that it's just a great use case
- 15:58altogether and a very beginner-friendly
- 16:00one because everyone understands the
- 16:01concept of an AI chat, right? The great
- 16:03part, too, is that you have many choices
- 16:05for the right user interface here. You
- 16:07can use something like LM Studio, which
- 16:09is a desktop app that already has a chat
- 16:10UI, or even create your own web app and
- 16:12customize it to your liking. There are
- 16:14so many open-source repos, and again,
- 16:16check out the link in the description if
- 16:18you want to make a good start on that
- 16:19because I've got many templates for you
- 16:21to customize.
- 16:22A specific subset of AI chats is RAG,
- 16:25retrieval augmented generation, which is
- 16:26a very fancy word that some of you
- 16:28probably already know about, which is
- 16:29the fact that you can point an AI at
- 16:31your own files, company documents,
- 16:33research papers, or notes, and it will
- 16:35answer questions based on that content.
- 16:37It's pretty similar to AI chats, but the
- 16:39thing is is that generally it's a bit
- 16:40more expensive and complex to set up.
- 16:42You might have to put your documents
- 16:44into a vector database, and you have to
- 16:46make sure that those documents are
- 16:48retrieved properly to make sure that the
- 16:50AI chats can answer questions about
- 16:52them. But if you do it right, well, this
- 16:54can work very well for your own use
- 16:56cases and to make sure that these local
- 16:58models are up to date with the latest
- 17:00knowledge in your use case, which is not
- 17:02in the training data of the model,
- 17:04especially because some of these
- 17:05open-source models have been trained
- 17:06half a year or a year ago. So, RAG is a
- 17:09very important paradigm, and I'm putting
- 17:11it in beats here because of the
- 17:12technical complexity to set it up, but
- 17:14it is kind of have
- 17:15for a lot of real AI chat use cases. So,
- 17:18it's definitely something you want to
- 17:19learn to set up yourself. One tool you
- 17:21can start with is something like Open
- 17:23Web UI, which gives you a full rack
- 17:25pipeline out of the box with a clean
- 17:26interface. And the nice part about this
- 17:29is that because it's all running
- 17:30locally, data privacy is the consistent
- 17:32top reason enterprises will use these
- 17:35kinds of self-hosted LLMs. So, if you
- 17:37learn these skills, you can definitely
- 17:38get an AI engineering job with that as
- 17:40well. So, we have talked about some AI
- 17:42use cases now like a chat, a rack chat,
- 17:45we have agent decoding, code auto
- 17:46complete, but what about the ability to
- 17:48create any AI agent of your dream? Where
- 17:51I'm defining an agent as assistant that
- 17:53can autonomously make decisions, execute
- 17:55actual actions for you, and be able to
- 17:58solve a problem in many different ways.
- 18:00Well, the thing is a true AI agent is
- 18:02very difficult to run locally. I'm going
- 18:04to put it in C tier, but I want to make
- 18:06sure I explain this to you because
- 18:08you've probably seen many videos here on
- 18:09YouTube talking about AI agents. The
- 18:11problem is though that most of these
- 18:12videos are not real agents. They're more
- 18:14deterministic workflows than something
- 18:16like Any Ten with a small LLM component
- 18:19in the middle that might make one or two
- 18:20decisions. But a true AI agent that can
- 18:23run like Claude Code simply requires a
- 18:25very good language model, or else it
- 18:28will get confused about all the tools it
- 18:30has access to. It will not be able to
- 18:32run autonomously without you pushing it
- 18:34into the right direction, and you will
- 18:35just find that there are many issues
- 18:37with it in general. That being said, I'm
- 18:39not saying that AI agents are incapable
- 18:41to work locally, it just depends on your
- 18:43use case. And most people that I've
- 18:45talked to want to build the first AI
- 18:47agents are a little bit too ambitious.
- 18:49They might want to build a research
- 18:50agent that works better than something
- 18:52you can use with OpenAI GPT Pro. But to
- 18:55be quite honest, it's very difficult to
- 18:57build a better AI agent than using a
- 18:59platform that you can access over the
- 19:01cloud, again something like Claude Code.
- 19:03Because even with Claude Code, you can
- 19:05point it to a local model, but it just
- 19:07won't perform in the same autonomous
- 19:09way. That being said, AI agents in
- 19:11particular change all the time, and I
- 19:13expect that as models get better, this
- 19:15will move up to the B tier and the A
- 19:17tier eventually. But even then, AI
- 19:20agents simply won't run on weak hardware
- 19:23because of literal mathematical
- 19:25constraints that I've covered in other
- 19:26videos on my channel. So, depending on
- 19:28your use case, this is just not going to
- 19:30work very well. But if you think that
- 19:33that's not true, and you have had good
- 19:34experiences creating local AI agents, I
- 19:37would love to hear it from you in the
- 19:38comments down below. But please tell me
- 19:40what use case you have because most
- 19:42people trying to run local AI agents
- 19:44aren't really running through agents.
- 19:46They're just running regular workflows
- 19:47with LLMs sprinkled somewhere in
- 19:49between. But how about a more simple AI
- 19:51assistant that you can run locally?
- 19:53Like, you know, Open Claw. An always-on
- 19:56personal AI that manages your calendar,
- 19:58triages your email, summarizes your day,
- 20:00and handles whatever you throw at it for
- 20:02personal use cases. The local version of
- 20:04Siri or Google Assistant, but running
- 20:06private and 24/7 on your own hardware.
- 20:09So, tools like Open Claw aim to be
- 20:11exactly this, and sound great on paper.
- 20:13The issue is though that setting them up
- 20:15properly in a secure way that makes sure
- 20:17that you actually keep your account
- 20:18secure is pretty difficult. It does
- 20:20require a little bit of security
- 20:22knowledge. Because of that complexity, I
- 20:24would right now put AI assistants into
- 20:25the B tier because they can work quite
- 20:27well if you are very aware of the
- 20:29security problems with something like
- 20:31Open Claw. Now, personally, I would
- 20:33still use Open Claw with a stated yard
- 20:34cloud model because they're much better
- 20:36protected against things like
- 20:37jailbreaks. But there may be one
- 20:39exception where I would use local
- 20:41models, which is for well-defined cron
- 20:43jobs, which are things that run on a
- 20:45schedule, like every 8 hours. This might
- 20:47be something where you are summarizing
- 20:49your feed into a daily digest or just
- 20:51classifying your incoming emails. For
- 20:53something like this, a local 14 billion
- 20:55model works just fine. And yes, you can
- 20:57use a local model with that. So far, we
- 20:59have a lot of use cases, and they all
- 21:01are pretty competent, right? Even the
- 21:03ones in C tier, they're very usable, and
- 21:05they're getting better every single day.
- 21:07But then, what is the D tier for? Well,
- 21:09to be honest, that's for vibe coding.
- 21:11Vibe coding with local models just
- 21:13doesn't work. The idea with vibe coding
- 21:15is that you just describe an app in
- 21:17plain English and the AI will build the
- 21:18whole thing and you never read the code,
- 21:21you just judge whether the result works.
- 21:23And this is different from agenda coding
- 21:24because you are not explicitly reviewing
- 21:26or steering anything. Now, honestly,
- 21:28this workflow needs a frontier model to
- 21:30cover for the fact that you're not
- 21:31reviewing at all what it writes. A small
- 21:34language model can code quite well with
- 21:36the right guidance in a couple of files.
- 21:38And for models under 14 billion
- 21:40parameters, you cannot even use tool
- 21:42calling properly, which is a requirement
- 21:44for proper agenda coding, let alone vibe
- 21:47coding where you're not guiding the
- 21:48model at all. Vibe coding with cloud
- 21:50models already has serious problems.
- 21:52Researchers found security
- 21:53vulnerabilities in one out of 10 lovable
- 21:56generated apps, for example. But with
- 21:58weaker local models, those problems
- 22:00multiply. So, let's recap this tier list
- 22:03a little bit. The three S tier use cases
- 22:06generally match or sometimes even beat
- 22:08cloud models. Code auto complete, image
- 22:10generation, speech to text. The pattern
- 22:12here is that some of the more boring use
- 22:14cases consistently outperform the hyped
- 22:16ones for local models. However, as
- 22:19models get better, some of the more
- 22:20complex workflows, like AI agents as
- 22:23well as voice agents, will just get
- 22:25better over time and hopefully
- 22:27everything will be in the B tier or
- 22:29above in a couple of years from now. And
- 22:32if you want to get started with local
- 22:33app projects, you should check out the
- 22:35link in the description below and get
- 22:37started today.
About this transcript
This page contains the full transcript of The Ultimate Local AI Tier List For 2026 by Zen van Riel, generated from the public captions YouTube serves with the video. The transcript has 4,689 words across 687 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.