YouTube2Text

The Ultimate Local AI Tier List For 2026 — Transcript

by Zen van Riel · 4,689 words · 687 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Today, I'm ranking the most popular
  2. 0:01local AI use cases for you based on real
  3. 0:03engineering experience. For each
  4. 0:05category, I'll recommend the one model
  5. 0:07or tool that I would start with, so you
  6. 0:09can skip hours of research and just get
  7. 0:11going. Personally, I've spent hundreds
  8. 0:13of hours testing with my RTX 4090, but a
  9. 0:16lot of cases work on weaker hardware,
  10. 0:18too. So, with that being said, let's get
  11. 0:20right into it. First stop, we're going
  12. 0:22to start with local code auto complete,
  13. 0:24and this is straight up S-tier.
  14. 0:27Code auto complete is one of the first
  15. 0:28ways that we were using AI for automated
  16. 0:31coding with something like GitHub
  17. 0:32Copilot. When you type code, the model
  18. 0:34finishes your line or even fills in a
  19. 0:36function body before you can think about
  20. 0:38it. And while agentic coding seems to be
  21. 0:40taking over the world, code auto
  22. 0:41complete is still extremely useful. And
  23. 0:44the great part is that it absolutely
  24. 0:45works great with local models. One model
  25. 0:48you could use is something like Qwen 2.5
  26. 0:50Coder, because it's a small 7 billion
  27. 0:53parameter model, which can run at
  28. 0:55sub-100 millisecond latency, even
  29. 0:57sometimes on GPUs with just a couple
  30. 0:59gigabytes of VRAM. Now, you're going to
  31. 1:02sometimes even get results that are
  32. 1:03faster than network-based auto
  33. 1:04completes, because there's no round
  34. 1:06trip. You're just, you know, executing
  35. 1:08this with local models. And this gives
  36. 1:11you a similar experience as the code
  37. 1:12auto complete models that were
  38. 1:14state-of-the-art a couple of years ago.
  39. 1:16You can pair your local model with
  40. 1:18something like Continue Dev and use
  41. 1:20Ollama for the inference, and then
  42. 1:21you've got a self-hosted Copilot
  43. 1:24replacement, at the very least, you
  44. 1:25know, the auto complete features of it,
  45. 1:27that costs nothing after hardware. While
  46. 1:29this doesn't beat the newest agentic
  47. 1:31features that you get with things like
  48. 1:33Cloud Code, Copilot, Codex, it's a great
  49. 1:35start, and this actually works well
  50. 1:37locally. Local can beat cloud models
  51. 1:40here on the metric that matters most
  52. 1:42when it comes to auto complete, which is
  53. 1:43response speed, and you can fully
  54. 1:45customize it yourself, as well. We're
  55. 1:46going to be talking about other coding
  56. 1:48use cases later, but first, let's move
  57. 1:49to something different: photo
  58. 1:51enhancement. This is something that I
  59. 1:53would comfortably place in A-tier. Photo
  60. 1:55enhancement covers things like
  61. 1:57upscaling, face restoration, background
  62. 1:59removal, noise reduction, colorization,
  63. 2:02anything basically where you take an
  64. 2:03existing image and you make it better
  65. 2:05without generating new content. I'll
  66. 2:08focus on upscaling for now as this is an
  67. 2:09example that is very common, but the
  68. 2:11whole category is strong locally. A tool
  69. 2:14that you can get started with easily is
  70. 2:15something like Upscaley. It's a free,
  71. 2:18open-source desktop app that uses
  72. 2:19Real-ESRGAN models. You don't have to
  73. 2:22screw around too much with Python, you
  74. 2:24don't need the command line, yet no
  75. 2:25custom GPU configuration. Just to see
  76. 2:27what's possible, you drag a photo in and
  77. 2:30you will get a 4x or even 8x upscaled
  78. 2:32version out of it. From there, because
  79. 2:35you can see what's possible, you can go
  80. 2:36ahead and try and play around with the
  81. 2:38models yourself and create a custom
  82. 2:40workflow that works for whatever you
  83. 2:41need to do. It's great to be able to
  84. 2:43enhance your photos without having to
  85. 2:45rely on only a cloud-based version of
  86. 2:47Photoshop. And any dedicated GPU with
  87. 2:50something like 4GB of free RAM can
  88. 2:52handle this use case in just seconds.
  89. 2:55Another great use case is home
  90. 2:57automation. Home automation covers many
  91. 3:00different tasks, but the whole category
  92. 3:02can definitely be A-tier. With home
  93. 3:05automation, you can have use cases like
  94. 3:06having your security cameras detect a
  95. 3:08person in your driveway and your phone
  96. 3:10getting a notification with a snapshot,
  97. 3:12which will all be processed on a box in
  98. 3:14your closet. And the great part about
  99. 3:16this is that you will not have a cloud
  100. 3:18subscription anymore and maybe more
  101. 3:20importantly, your security footage
  102. 3:22leaving your network to be stored on
  103. 3:24maybe some random Chinese cloud that you
  104. 3:26don't know about. Object detection is
  105. 3:28the headline feature here, but local
  106. 3:30home automation can also cover voice
  107. 3:32control, presence sensing, and even
  108. 3:34energy monitoring. In any case, the
  109. 3:36stack that you can start with is Frigate
  110. 3:38NVR plus Home Assistant. This is
  111. 3:40probably the most mature local AI
  112. 3:42ecosystem that exists right now,
  113. 3:44especially for beginners.
  114. 3:46With this, you can have use cases like
  115. 3:48person detection, vehicle detection,
  116. 3:50pets, packages, license plates, and even
  117. 3:52some basic facial recognition as part of
  118. 3:55some of the latest Frigate versions. And
  119. 3:57all of this will be processed entirely
  120. 3:59on your hardware. Home Assistant has
  121. 4:01over 2 million active installations, and
  122. 4:03Frigate has over 30,000 stars on GitHub,
  123. 4:05so you're not going to be alone. It's
  124. 4:06not just a hobby project. It will be a
  125. 4:08full ecosystem that you can plug into.
  126. 4:10The only thing keeping us from S tier is
  127. 4:12the setup complexity. You often will
  128. 4:14have to configure Docker, camera
  129. 4:15streams, detection zones, and even if
  130. 4:17it's running, it might break in a couple
  131. 4:19of weeks because things update fast in
  132. 4:21the home automation space. Now, you
  133. 4:23might be already thinking, "This is a
  134. 4:25nice tier list so far, but how do I get
  135. 4:27started?" Well, I put together a
  136. 4:29collection of my own open-source
  137. 4:30projects for many of the local AI use
  138. 4:32cases we're covering today. The link is
  139. 4:34in the description, and you should
  140. 4:35definitely check it out. A whole
  141. 4:36different use case is video generation,
  142. 4:39and this one is a little bit
  143. 4:40disappointing. I'm going to put it in
  144. 4:41the C tier because the idea behind video
  145. 4:44generation is that you can type a prompt
  146. 4:45and get a full video clip out. Something
  147. 4:48you can use for product demos,
  148. 4:49children's video content, or just
  149. 4:51content visualization. This is the
  150. 4:52category everyone wants to work locally
  151. 4:54because cloud video generation is
  152. 4:56expensive and rate-limited. Problem is
  153. 4:58though that it's still very expensive to
  154. 5:00generate locally in terms of the time
  155. 5:01that it takes, and the quality is just
  156. 5:03not great. One model you can try is When
  157. 5:052.1 from Alibaba, which can beat Sora on
  158. 5:08several benchmarks, which was, you know,
  159. 5:10a state-of-the-art model a while back.
  160. 5:13The issue is though that even on my
  161. 5:141590, I can't really use a full 14
  162. 5:17billion model. I already have to kind of
  163. 5:19go down to the smaller 5 billion
  164. 5:21parameter one, and the quality is just a
  165. 5:24lot less. In addition, because it takes
  166. 5:26a long time to generate video, it is
  167. 5:28also a lot more time expensive to
  168. 5:30generate many variants to figure out how
  169. 5:32we can generate a video clip that
  170. 5:34actually looks good. Compared to image
  171. 5:36generation, which we'll get to later,
  172. 5:37you don't just generate one frame. To
  173. 5:39get a good video, you need to generate
  174. 5:41many different frames, and that just
  175. 5:42takes a long time. Locally, this might
  176. 5:46be usable for some experimental social
  177. 5:48media clips and concept work, but real
  178. 5:50professional production is still cloud
  179. 5:52territory, and that's why this is going
  180. 5:54in C tier.
  181. 5:55Speaking of something that's related,
  182. 5:57image generation, now this works much
  183. 5:59better. I would put image generation in
  184. 6:01the S tier. With image generation, you
  185. 6:03describe an image and the model can
  186. 6:05create it. You can use it for
  187. 6:06thumbnails, marketing assets, product
  188. 6:08mock-ups, and concept art. This can also
  189. 6:11include in-painting, out-painting,
  190. 6:13basically editing existing image, and
  191. 6:15even generating a new image variant
  192. 6:17based on an image that you're passing to
  193. 6:19it. Now, I'll focus on text-to-image
  194. 6:21generation because that's where most
  195. 6:23people start. The model you can use here
  196. 6:25is Flux 2 Dev, which you can run through
  197. 6:27ComfyUI. On something like my 5090, this
  198. 6:30is a very comfortable use case, and it's
  199. 6:32very easy for me to generate an image in
  200. 6:34just a couple of seconds. Unlike video
  201. 6:36generation, this means that you can
  202. 6:37generate hundreds of variants in, you
  203. 6:39know, a short amount of time, which
  204. 6:41increases the chances that you actually
  205. 6:43get an image that you're happy with.
  206. 6:45And in fact, Flux is a really good model
  207. 6:47because in some blind tests, it achieved
  208. 6:49a 71% win rate over some of the older
  209. 6:52Midjourney versions for editorial
  210. 6:54photorealism. So, the models are really
  211. 6:56getting much better.
  212. 6:59And with image generation, I've also
  213. 7:00found that there's a really great
  214. 7:01community where folks have fine-tuned a
  215. 7:04lot of models for custom characters and
  216. 7:06styles, and there's even a lot of image
  217. 7:08generation models that have less content
  218. 7:10filters than the one in the cloud. It's
  219. 7:12good that cloud has content filters,
  220. 7:14especially for copyright, but sometimes
  221. 7:16they go a little bit too far, and they
  222. 7:17become a bit unusable, to be honest. And
  223. 7:19actually training a custom LoRA model,
  224. 7:22or actually fine-tuning one, only takes
  225. 7:2450 to 20 images, and you can do that on
  226. 7:26consumer hardware, as well. Now, one
  227. 7:28thing that I've learned from using AI
  228. 7:30image generation for my own content is
  229. 7:31that iterative editing is where it gets
  230. 7:33a bit tricky.
  231. 7:35Image properties like gradients and
  232. 7:36complex textures make it very difficult
  233. 7:38for those local models to edit images
  234. 7:41without losing quality. Personally, this
  235. 7:43is still where I use the latest native
  236. 7:45been added model or just use Photoshop
  237. 7:47to just make sure that I can use some of
  238. 7:49the cloud features. So, even though it's
  239. 7:50S tier, there are some use cases where
  240. 7:53it's going to fall behind a little bit,
  241. 7:54but it is really getting better every
  242. 7:56single month. A completely different but
  243. 7:58fun use case is voice agents. Now, this
  244. 8:00one is something that I'm going to place
  245. 8:01in C tier because of the wide variety of
  246. 8:04use cases where it can work in a variety
  247. 8:07of ways. Basically, the idea behind
  248. 8:08voice agents is that you can talk to
  249. 8:10your computer and it will talk back like
  250. 8:12a local Alexa or customer support bots
  251. 8:13that runs on your hardware. But ideally,
  252. 8:15it should also be able to, you know, do
  253. 8:17actions for you. For example, you can
  254. 8:19use it to do hands-free home automation.
  255. 8:22The best open source option right now
  256. 8:23would be something like Pipe Cat, which
  257. 8:25can chain together a speech to text
  258. 8:27model, an LLM, and then a text to speech
  259. 8:29model into one single pipeline. And that
  260. 8:32will allow you to get something like sub
  261. 8:34800 millisecond voice to voice latency
  262. 8:36on some standard hardware on something
  263. 8:39like macOS.
  264. 8:41The problem is that the model size
  265. 8:43constraints means that the AI's
  266. 8:44responses locally are noticeably dumber
  267. 8:47than you can get from the best local
  268. 8:48models that you interact with over chat,
  269. 8:50and of course, the best cloud models. In
  270. 8:53fact, this is even an issue with cloud
  271. 8:54voice agents. They can achieve 200 to
  272. 8:56500 millisecond latency, but their
  273. 8:58intelligence is more around GPT-4 level
  274. 9:01and not to the latest models that are
  275. 9:02out there. And of course, you know, this
  276. 9:05intelligence gap is much higher for
  277. 9:07local voice agents. That being said, if
  278. 9:10you just have a singular command that
  279. 9:11you want to run, voice agents can work
  280. 9:13pretty well. The main issue where they
  281. 9:15lack quality is in longer conversations
  282. 9:17because with local models, you're just
  283. 9:19much more constrained on your context
  284. 9:21window, which means that as you have a
  285. 9:22long conversation with a voice agent, it
  286. 9:25will simply go off track. And you don't
  287. 9:27have that issue if your voice agent is
  288. 9:29just there to process one command and
  289. 9:31execute it. So, that would be the
  290. 9:33recommended use case, I would say. Now,
  291. 9:35one component of voice agents is a text
  292. 9:37to speech model and you can use text to
  293. 9:39speech on its own and on its own this is
  294. 9:41definitely a tier. With text to speech
  295. 9:43you feed text in and you get natural
  296. 9:45sounding audio out. You can use this for
  297. 9:47audiobook narration, voiceover for
  298. 9:48video, accessibility features in your
  299. 9:50app, or even trying to clone your own
  300. 9:52voice for content which I've tried but
  301. 9:54that one is a bit ambitious it doesn't
  302. 9:55work very well. Now this category also
  303. 9:57covers music generation and sound
  304. 9:59effects because it's not really my
  305. 10:01specialty. In general though voice
  306. 10:03synthesis is where local models have
  307. 10:04made the biggest gains. Now text to
  308. 10:06speech has had the most dramatic
  309. 10:08transformation in my opinion of any
  310. 10:10local AI category in the past 18 months
  311. 10:12because I would say that it's almost a
  312. 10:14solved problem at this point especially
  313. 10:16for the English language. The model that
  314. 10:18I would start with is Chatterbox from
  315. 10:20Resemble AI because it beats 11 labs in
  316. 10:23a couple of blind tests with over 60%
  317. 10:26listener preference rates. And the base
  318. 10:28model might be English only but you can
  319. 10:30even go to Chatterbox multilingual which
  320. 10:32covers 23 plus languages. And so the gap
  321. 10:35versus a cloud offering like 11 labs has
  322. 10:37closed for a lot of use cases and in
  323. 10:40some cases the gap has gone entirely.
  324. 10:43There are some gotchas. A lot of models
  325. 10:45will hallucinate or really degrade in
  326. 10:46quality past about 1,000 characters but
  327. 10:50there's a lot of ways to solve that. You
  328. 10:51know, you can split up your text in
  329. 10:53multiple batches and just process them
  330. 10:55one by one. Now if you're finding this
  331. 10:56tier list useful make sure to hit
  332. 10:58subscribe because over 90% of you are
  333. 11:00missing out on the latest AI engineering
  334. 11:02news that's based on reality and not
  335. 11:05hype so make sure to join the club. Next
  336. 11:07we're going to be talking about speech
  337. 11:08to text which is another component of
  338. 11:10voice agents that you can extract on its
  339. 11:12own and it's another solved problem in
  340. 11:14my opinion. Speech to text is absolutely
  341. 11:17S tier.
  342. 11:18You record audio you pass an audio file
  343. 11:20you can get a text transcript back. For
  344. 11:23meeting notes, podcast transcription,
  345. 11:24subtitle generation, or just turning
  346. 11:27your voice memos into searchable text. I
  347. 11:30use this constantly myself for
  348. 11:31processing my own YouTube video content
  349. 11:34and for English it works very well. A
  350. 11:36model you can use is faster Whisper with
  351. 11:38large V3 Turbo which gives you over four
  352. 11:41times speed over the original Whisper
  353. 11:42model. And the real workflow that I use
  354. 11:45is a two-stage pipeline. I use Whisper
  355. 11:47to get a fast, accurate raw
  356. 11:49transcription and then I use local LLM
  357. 11:52to clean up filler words and extract the
  358. 11:54core meaning. And this way I can store
  359. 11:56the notes into something like my
  360. 11:57Obsidian vault to reference later. The
  361. 11:59transcription itself is almost instant,
  362. 12:02but the LM cleanup step does add some
  363. 12:04time. The thing is though, I can just
  364. 12:05run it in the background, so that's not
  365. 12:07an issue at all for me.
  366. 12:09The gap of cloud models is mainly with
  367. 12:10speaker diarization, which means that
  368. 12:13you can split up some audio context into
  369. 12:16the different speakers. So let's say you
  370. 12:17have a meeting with 10 people, you want
  371. 12:19to be able to know who's speaking which
  372. 12:21sentences, right? Cloud models are a
  373. 12:23little bit better at doing that, but I
  374. 12:25fully believe that local models will
  375. 12:26catch up very quickly. The core use case
  376. 12:28of transcribing audio just works very
  377. 12:30well nowadays with local models.
  378. 12:33Next we're going to be talking about
  379. 12:34OCR, optical character recognition,
  380. 12:36which I'm going to comfortably place as
  381. 12:38our first beat here.
  382. 12:40OCR covers table extraction, formula
  383. 12:42recognition, and just converting scanned
  384. 12:43documents into structured data, which is
  385. 12:45useful for so many use cases. One tool
  386. 12:48to start with is Suriya, or you can use
  387. 12:50a Deep Seek OCR model that has been
  388. 12:51released recently. Now, a lot of the use
  389. 12:53cases are a little bit more on the
  390. 12:54boring side, like being able to transfer
  391. 12:56and process invoices, but it does work
  392. 12:59pretty well. Next we're going to be
  393. 13:00talking about agentic coding, which is
  394. 13:02much more interesting than just doing
  395. 13:03code autocomplete and seems to be hyped
  396. 13:05up more and more for local models. Now,
  397. 13:08the one disappointing thing about
  398. 13:09agentic coding is that it doesn't work
  399. 13:11that well unless you have very good
  400. 13:13hardware. A lot of YouTube videos they
  401. 13:15show agentic coding and then they show
  402. 13:17how it can generate Python code and
  403. 13:18that's good and all, but with agentic
  404. 13:20coding you really want a model to be
  405. 13:22competent enough to read your entire
  406. 13:24codebase, write code, run tests, and
  407. 13:26iterate until your feature works, the
  408. 13:28feature that you asked for with text.
  409. 13:31The issue is though that most absolutely
  410. 13:33most hardware is too incompetent for
  411. 13:35agentic coding.
  412. 13:37Now, I've created many videos on my
  413. 13:38channel covering how performant agentic
  414. 13:41coding is on my 5090, and it's
  415. 13:43definitely getting closer. But, it is
  416. 13:45nowhere near as useful as using code
  417. 13:47autocomplete, nor something like an AI
  418. 13:50chat, because eventually local AI agents
  419. 13:53choke on larger projects. That being
  420. 13:55said, this is something that improves
  421. 13:56every single month. So, definitely check
  422. 13:59out the videos that I have in my
  423. 14:00channel, which I'll put in the link in
  424. 14:01the description below, to get the most
  425. 14:03out of your local setup, because it's
  426. 14:04very difficult to set up properly, and
  427. 14:06you really need a bit of an expert guide
  428. 14:08to do this right. With agentic coding,
  429. 14:10you might be able to match something
  430. 14:11like GPT-4o locally, but Claude Opus 4.6
  431. 14:14and similar models have really upped the
  432. 14:16game here, and there is just a huge gap
  433. 14:18between what you can do locally versus
  434. 14:19with the state-of-the-art Claude models.
  435. 14:22Even if your local model is intelligent
  436. 14:24enough, the problem is is that it will
  437. 14:26get slow as the context window fills up.
  438. 14:28And unlike with these other use cases,
  439. 14:29where you can clear out the context
  440. 14:31window after you improve small bits of
  441. 14:34the task, with agentic coding, you
  442. 14:36cannot clear the context window with
  443. 14:37every step. You need the coding model to
  444. 14:40have a good idea of how your code base
  445. 14:41works. So, in the end, these models are
  446. 14:44going to work very slowly on your local
  447. 14:46hardware, especially compared to cloud
  448. 14:48models, unless you have a very beefy PC.
  449. 14:50Most people promoting local models with
  450. 14:52agentic coding are not using it for
  451. 14:54serious projects. And if you don't
  452. 14:55believe me, check out some of our
  453. 14:57masterclasses on this channel to learn
  454. 14:59the truth. If you do want to get started
  455. 15:01with agentic coding, I recommend some of
  456. 15:02the latest Qwen models, but again, I
  457. 15:04covered that extensively in the videos
  458. 15:06on my channel already, so go check those
  459. 15:08out. As models get stronger, I expect it
  460. 15:10to move into A or S tier, especially as
  461. 15:13local models will catch up as well. You
  462. 15:15know what isn't A tier yet, though? AI
  463. 15:17chats, a traditional use case. With AI
  464. 15:19chats, you can ask questions, brainstorm
  465. 15:21ideas, summarize documents, draft
  466. 15:23emails, the same stuff you would use
  467. 15:25something like ChatGPT for, but it can
  468. 15:27run entirely on your machine with no
  469. 15:29data leaving your network. I would
  470. 15:30recommend new models like something like
  471. 15:32Coin 330B through LM Studio. And this is
  472. 15:36a pretty optimized model that can run
  473. 15:37pretty fast, but honestly, there's many
  474. 15:39choices for AI chat models. You can use
  475. 15:41a newer Mistral model or even check out
  476. 15:44the older, but still competent
  477. 15:46open-source OpenAI model that's out
  478. 15:47there. In any case, locally, you can
  479. 15:49really mimic what GPT-4 was able to do a
  480. 15:52year plus ago and now fully locally. You
  481. 15:55can use it for so many different
  482. 15:56purposes that it's just a great use case
  483. 15:58altogether and a very beginner-friendly
  484. 16:00one because everyone understands the
  485. 16:01concept of an AI chat, right? The great
  486. 16:03part, too, is that you have many choices
  487. 16:05for the right user interface here. You
  488. 16:07can use something like LM Studio, which
  489. 16:09is a desktop app that already has a chat
  490. 16:10UI, or even create your own web app and
  491. 16:12customize it to your liking. There are
  492. 16:14so many open-source repos, and again,
  493. 16:16check out the link in the description if
  494. 16:18you want to make a good start on that
  495. 16:19because I've got many templates for you
  496. 16:21to customize.
  497. 16:22A specific subset of AI chats is RAG,
  498. 16:25retrieval augmented generation, which is
  499. 16:26a very fancy word that some of you
  500. 16:28probably already know about, which is
  501. 16:29the fact that you can point an AI at
  502. 16:31your own files, company documents,
  503. 16:33research papers, or notes, and it will
  504. 16:35answer questions based on that content.
  505. 16:37It's pretty similar to AI chats, but the
  506. 16:39thing is is that generally it's a bit
  507. 16:40more expensive and complex to set up.
  508. 16:42You might have to put your documents
  509. 16:44into a vector database, and you have to
  510. 16:46make sure that those documents are
  511. 16:48retrieved properly to make sure that the
  512. 16:50AI chats can answer questions about
  513. 16:52them. But if you do it right, well, this
  514. 16:54can work very well for your own use
  515. 16:56cases and to make sure that these local
  516. 16:58models are up to date with the latest
  517. 17:00knowledge in your use case, which is not
  518. 17:02in the training data of the model,
  519. 17:04especially because some of these
  520. 17:05open-source models have been trained
  521. 17:06half a year or a year ago. So, RAG is a
  522. 17:09very important paradigm, and I'm putting
  523. 17:11it in beats here because of the
  524. 17:12technical complexity to set it up, but
  525. 17:14it is kind of have
  526. 17:15for a lot of real AI chat use cases. So,
  527. 17:18it's definitely something you want to
  528. 17:19learn to set up yourself. One tool you
  529. 17:21can start with is something like Open
  530. 17:23Web UI, which gives you a full rack
  531. 17:25pipeline out of the box with a clean
  532. 17:26interface. And the nice part about this
  533. 17:29is that because it's all running
  534. 17:30locally, data privacy is the consistent
  535. 17:32top reason enterprises will use these
  536. 17:35kinds of self-hosted LLMs. So, if you
  537. 17:37learn these skills, you can definitely
  538. 17:38get an AI engineering job with that as
  539. 17:40well. So, we have talked about some AI
  540. 17:42use cases now like a chat, a rack chat,
  541. 17:45we have agent decoding, code auto
  542. 17:46complete, but what about the ability to
  543. 17:48create any AI agent of your dream? Where
  544. 17:51I'm defining an agent as assistant that
  545. 17:53can autonomously make decisions, execute
  546. 17:55actual actions for you, and be able to
  547. 17:58solve a problem in many different ways.
  548. 18:00Well, the thing is a true AI agent is
  549. 18:02very difficult to run locally. I'm going
  550. 18:04to put it in C tier, but I want to make
  551. 18:06sure I explain this to you because
  552. 18:08you've probably seen many videos here on
  553. 18:09YouTube talking about AI agents. The
  554. 18:11problem is though that most of these
  555. 18:12videos are not real agents. They're more
  556. 18:14deterministic workflows than something
  557. 18:16like Any Ten with a small LLM component
  558. 18:19in the middle that might make one or two
  559. 18:20decisions. But a true AI agent that can
  560. 18:23run like Claude Code simply requires a
  561. 18:25very good language model, or else it
  562. 18:28will get confused about all the tools it
  563. 18:30has access to. It will not be able to
  564. 18:32run autonomously without you pushing it
  565. 18:34into the right direction, and you will
  566. 18:35just find that there are many issues
  567. 18:37with it in general. That being said, I'm
  568. 18:39not saying that AI agents are incapable
  569. 18:41to work locally, it just depends on your
  570. 18:43use case. And most people that I've
  571. 18:45talked to want to build the first AI
  572. 18:47agents are a little bit too ambitious.
  573. 18:49They might want to build a research
  574. 18:50agent that works better than something
  575. 18:52you can use with OpenAI GPT Pro. But to
  576. 18:55be quite honest, it's very difficult to
  577. 18:57build a better AI agent than using a
  578. 18:59platform that you can access over the
  579. 19:01cloud, again something like Claude Code.
  580. 19:03Because even with Claude Code, you can
  581. 19:05point it to a local model, but it just
  582. 19:07won't perform in the same autonomous
  583. 19:09way. That being said, AI agents in
  584. 19:11particular change all the time, and I
  585. 19:13expect that as models get better, this
  586. 19:15will move up to the B tier and the A
  587. 19:17tier eventually. But even then, AI
  588. 19:20agents simply won't run on weak hardware
  589. 19:23because of literal mathematical
  590. 19:25constraints that I've covered in other
  591. 19:26videos on my channel. So, depending on
  592. 19:28your use case, this is just not going to
  593. 19:30work very well. But if you think that
  594. 19:33that's not true, and you have had good
  595. 19:34experiences creating local AI agents, I
  596. 19:37would love to hear it from you in the
  597. 19:38comments down below. But please tell me
  598. 19:40what use case you have because most
  599. 19:42people trying to run local AI agents
  600. 19:44aren't really running through agents.
  601. 19:46They're just running regular workflows
  602. 19:47with LLMs sprinkled somewhere in
  603. 19:49between. But how about a more simple AI
  604. 19:51assistant that you can run locally?
  605. 19:53Like, you know, Open Claw. An always-on
  606. 19:56personal AI that manages your calendar,
  607. 19:58triages your email, summarizes your day,
  608. 20:00and handles whatever you throw at it for
  609. 20:02personal use cases. The local version of
  610. 20:04Siri or Google Assistant, but running
  611. 20:06private and 24/7 on your own hardware.
  612. 20:09So, tools like Open Claw aim to be
  613. 20:11exactly this, and sound great on paper.
  614. 20:13The issue is though that setting them up
  615. 20:15properly in a secure way that makes sure
  616. 20:17that you actually keep your account
  617. 20:18secure is pretty difficult. It does
  618. 20:20require a little bit of security
  619. 20:22knowledge. Because of that complexity, I
  620. 20:24would right now put AI assistants into
  621. 20:25the B tier because they can work quite
  622. 20:27well if you are very aware of the
  623. 20:29security problems with something like
  624. 20:31Open Claw. Now, personally, I would
  625. 20:33still use Open Claw with a stated yard
  626. 20:34cloud model because they're much better
  627. 20:36protected against things like
  628. 20:37jailbreaks. But there may be one
  629. 20:39exception where I would use local
  630. 20:41models, which is for well-defined cron
  631. 20:43jobs, which are things that run on a
  632. 20:45schedule, like every 8 hours. This might
  633. 20:47be something where you are summarizing
  634. 20:49your feed into a daily digest or just
  635. 20:51classifying your incoming emails. For
  636. 20:53something like this, a local 14 billion
  637. 20:55model works just fine. And yes, you can
  638. 20:57use a local model with that. So far, we
  639. 20:59have a lot of use cases, and they all
  640. 21:01are pretty competent, right? Even the
  641. 21:03ones in C tier, they're very usable, and
  642. 21:05they're getting better every single day.
  643. 21:07But then, what is the D tier for? Well,
  644. 21:09to be honest, that's for vibe coding.
  645. 21:11Vibe coding with local models just
  646. 21:13doesn't work. The idea with vibe coding
  647. 21:15is that you just describe an app in
  648. 21:17plain English and the AI will build the
  649. 21:18whole thing and you never read the code,
  650. 21:21you just judge whether the result works.
  651. 21:23And this is different from agenda coding
  652. 21:24because you are not explicitly reviewing
  653. 21:26or steering anything. Now, honestly,
  654. 21:28this workflow needs a frontier model to
  655. 21:30cover for the fact that you're not
  656. 21:31reviewing at all what it writes. A small
  657. 21:34language model can code quite well with
  658. 21:36the right guidance in a couple of files.
  659. 21:38And for models under 14 billion
  660. 21:40parameters, you cannot even use tool
  661. 21:42calling properly, which is a requirement
  662. 21:44for proper agenda coding, let alone vibe
  663. 21:47coding where you're not guiding the
  664. 21:48model at all. Vibe coding with cloud
  665. 21:50models already has serious problems.
  666. 21:52Researchers found security
  667. 21:53vulnerabilities in one out of 10 lovable
  668. 21:56generated apps, for example. But with
  669. 21:58weaker local models, those problems
  670. 22:00multiply. So, let's recap this tier list
  671. 22:03a little bit. The three S tier use cases
  672. 22:06generally match or sometimes even beat
  673. 22:08cloud models. Code auto complete, image
  674. 22:10generation, speech to text. The pattern
  675. 22:12here is that some of the more boring use
  676. 22:14cases consistently outperform the hyped
  677. 22:16ones for local models. However, as
  678. 22:19models get better, some of the more
  679. 22:20complex workflows, like AI agents as
  680. 22:23well as voice agents, will just get
  681. 22:25better over time and hopefully
  682. 22:27everything will be in the B tier or
  683. 22:29above in a couple of years from now. And
  684. 22:32if you want to get started with local
  685. 22:33app projects, you should check out the
  686. 22:35link in the description below and get
  687. 22:37started today.

About this transcript

This page contains the full transcript of The Ultimate Local AI Tier List For 2026 by Zen van Riel, generated from the public captions YouTube serves with the video. The transcript has 4,689 words across 687 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.