YouTube2Text

Understanding vLLM with a Hands On Demo — Transcript

by KodeKloud · 2,253 words · 383 segments · language en · Watch on YouTube

Full transcript

  1. 0:00One of the biggest topics around LLM is
  2. 0:02when it comes to inference, which means
  3. 0:04when we actually serve the model to
  4. 0:06start generating output tokens. And
  5. 0:08depending on what system you use to
  6. 0:10serve LLMs, end users will experience
  7. 0:13different results. For example, when we
  8. 0:16go to ChatGPT and interact with the
  9. 0:18chatbot, it might generate its output at
  10. 0:21a certain speed. Well, when we go to
  11. 0:23Gemini, for example, it might generate
  12. 0:25its output at a much different speed. We
  13. 0:27measure this in what's called tokens per
  14. 0:30second. And depending on what system you
  15. 0:32use to actually run your large language
  16. 0:34model, the same model can be served at
  17. 0:36varying speed. And there are tons of
  18. 0:38systems out there like llama.cpp, vLLM,
  19. 0:41SGLang, TensorRT-LLM, Hugging Face TGI,
  20. 0:45LMDeploy. These are all a classification
  21. 0:48of inference engine that allows a model
  22. 0:50to be served. vLLM is a system that
  23. 0:53really stands out as a tool that allows
  24. 0:56you to run your model with high
  25. 0:57throughput. When you start to serve LLM
  26. 1:00beyond a single user or a single
  27. 1:02instance to multiple users across
  28. 1:05multiple instances, what starts to
  29. 1:07happen is that the underlying system
  30. 1:09that runs LLM needs to serve multiple
  31. 1:13requests at a time. And since during
  32. 1:15inference your prompt is stored not only
  33. 1:17in memory in terms of KV cache and auto
  34. 1:20regressively appends during its decoding
  35. 1:23phase, having a system that can actually
  36. 1:26manage this KV cache efficiently is very
  37. 1:28important. vLLM is a system that
  38. 1:31introduced an extremely efficient way to
  39. 1:33manage the KV cache where at any time it
  40. 1:36showed a significant improvement over
  41. 1:38existing systems that often wasted up to
  42. 1:4160 to 80% of its memory in
  43. 1:44fragmentation, over allocation, or both.
  44. 1:47And vLLM introduced what's called the
  45. 1:49paged attention in order to actually
  46. 1:51create virtualized KV cache that was
  47. 1:54inspired by how paging systems work at
  48. 1:57the operating system level. So, instead
  49. 1:59of over allocating your space where you
  50. 2:02prepare for the worst case in terms of
  51. 2:04how much the model will generate in
  52. 2:06output, vLLM, instead of padding the
  53. 2:08memory used in pages that grew in size,
  54. 2:11which frees up the memory to be used
  55. 2:14more efficiently as it processes many
  56. 2:16requests at a time. And not only was
  57. 2:19this more efficient in the, let's call
  58. 2:21it, storage Tetris game, it means that
  59. 2:23the underlying graphics card was busier.
  60. 2:26But, it's important to keep in mind that
  61. 2:28vLLM isn't necessarily an end-all
  62. 2:31solution. It means to solve throughput
  63. 2:33problem on a multi-user and
  64. 2:35multi-requests, but other inference
  65. 2:38engines prioritize other metrics like
  66. 2:41or optimizing to the vendor in the case
  67. 2:43of TensorRT-LLM for Nvidia, as well as
  68. 2:46better optimization support for running
  69. 2:48the model for less RAM requirements.
  70. 2:51But, learning how to use vLLM to serve
  71. 2:53LLM as more and more large language
  72. 2:56models are innovated, but also being
  73. 2:58able to serve it beyond a single user
  74. 3:01instances by sharing them across
  75. 3:03multiple users from server side is going
  76. 3:05to be an extremely important skill to
  77. 3:08monitor and optimize how your hardware
  78. 3:10is being used to provide services.
  79. 3:12Learning how to use vLLM is going beyond
  80. 3:15being impressed by an LLM, but actually
  81. 3:18trying to productize LLM so that you can
  82. 3:20monitor its tokens per speed and its
  83. 3:23outputs and its throughput as the number
  84. 3:25of sessions grow with various prompt
  85. 3:28requests in sizes also changes
  86. 3:30dynamically. And managing this exact
  87. 3:32stack is what vLLM allows you to work to
  88. 3:35lower the latency and maximize
  89. 3:37concurrency using vLLM.
  90. 3:42In this lab, we're going to learn about
  91. 3:44how vLLM turns a basic LLM setup into a
  92. 3:47production-ready inference server.
  93. 3:49[music] We run small LLM 135M with
  94. 3:52Hugging Face, then with vLLM and compare
  95. 3:55the speed. We will visualize the KV
  96. 3:57cache problem, see how paged attention
  97. 3:59fixes it, launch an API server, stress
  98. 4:02test it with concurrent users, tune
  99. 4:04parameters, and build a monitoring
  100. 4:06dashboard. This lab takes about 40 to 50
  101. 4:09minutes to complete. When I open the
  102. 4:11lab, I'm dropped right into the
  103. 4:12scenario. Access the lab using the link
  104. 4:15in the description below and follow
  105. 4:16along with me. The intro page tells us
  106. 4:18that we are an ML engineer at Inference
  107. 4:21IO, a startup building an LLM as a
  108. 4:23service platform. Our CEO just signed a
  109. 4:26deal to serve small LLM to multiple
  110. 4:28concurrent user. Our current Hugging
  111. 4:30Face setup handles one request at a
  112. 4:32time. Our mission is to use vLLM to
  113. 4:36build a production-ready inference
  114. 4:37server. The lab has a setup step, eight
  115. 4:40tasks, and four knowledge checks. The
  116. 4:42tasks cover Hugging Face baseline, vLLM
  117. 4:45inference, KV cache problem, paged
  118. 4:47attention, API server, multi-user load
  119. 4:50testing, parameter tuning, and a
  120. 4:52monitoring dashboard.
  121. 4:54>> [music]
  122. 4:54>> The first step is to verify our
  123. 4:56environment. Run the following command
  124. 4:58and then python/root/code
  125. 4:59verify_environment.py.
  126. 5:02This script checks that all required
  127. 5:03packages like vLLM, Transformers, and
  128. 5:06Gradio are all installed, downloads
  129. 5:08[music] the small LLM 135M model, runs a
  130. 5:12quick test generation, and confirms
  131. 5:14everything is ready. Wait until you see
  132. 5:16the environment verification complete
  133. 5:18message. Now, we move to task one, which
  134. 5:20is about running naive Hugging Face
  135. 5:22inference to establish a baseline. Open
  136. 5:24/root/code/task1_huggingface_baseline.py.
  137. 5:26You have to complete two to-dos. At line
  138. 5:2830, replace the blank with Hugging Face
  139. 5:30TB small LLM 135M to load the model, and
  140. 5:35at line 49, replace the blank with 50 to
  141. 5:37set the maximum number of new tokens to
  142. 5:40generate.
  143. 5:43[music] Run the script with
  144. 5:44python/root/code/task1_huggingface_baseline.py.
  145. 5:49The output shows the model loading,
  146. 5:51generating text, and reporting tokens
  147. 5:53per second. This number is our baseline
  148. 5:55that vLLM will beat. Now, we move to
  149. 5:57task two, which is about running the
  150. 5:59same model with vLLM. Open
  151. 6:01/root/code/task2_vllm_inference.py.
  152. 6:04[music] You need to complete two to-dos.
  153. 6:07At line 36, replace the blank with
  154. 6:09Hugging Face TB small LLM 135M to
  155. 6:12initialize the vLLM engine with the same
  156. 6:14model. At line 40, replace the blank
  157. 6:17with 0.7 for temperature and 50 for max
  158. 6:20tokens to create the sampling
  159. 6:22parameters. Run the script with
  160. 6:23python/root/code/task2_vllm_inference.py.
  161. 6:24[music]
  162. 6:28The output shows a side-by-side
  163. 6:30comparison of Hugging Face versus vLLM
  164. 6:33tokens per second. Notice that vLLM is
  165. 6:35faster, even though it is the exact same
  166. 6:37model and same prompt. This proves that
  167. 6:40the inference engine matters, [music]
  168. 6:42not just the model. We have a knowledge
  169. 6:44check about the inference basics. The
  170. 6:46question asks, "What is a primary metric
  171. 6:48used to measure how fast an LLM
  172. 6:50generates output?" [music] The answer is
  173. 6:52tokens per second because during
  174. 6:54inference the model produces tokens one
  175. 6:56at a time through auto regressive
  176. 6:58decoding. So, tokens per second directly
  177. 7:01tells [music] you how fast user sees
  178. 7:03responses. Now, we move to task three,
  179. 7:05which is about understanding the KV
  180. 7:06cache problem. Open
  181. 7:08/root/code/task3_kvcache_problem.py.
  182. 7:11[music] You need to complete two to-dos.
  183. 7:13At line 26, replace the blank with 512
  184. 7:16to set the maximum sequence length. At
  185. 7:19line 52, replace the blank with
  186. 7:21allocated minus actual divided by
  187. 7:23allocated times 100 to calculate the
  188. 7:25wasted percentage. Run the script with
  189. 7:28python/root/code/task3_kvcache_problem.py.
  190. 7:32The output shows how traditional system
  191. 7:34pre-allocate worst-case memory for every
  192. 7:37request. Short prompts waste [music]
  193. 7:39massive amounts of space because they
  194. 7:41get the same allocation as long as
  195. 7:43prompts. You will see that memory
  196. 7:45utilization is only about 20%, meaning
  197. 7:4780% is wasted. This is a KV cache
  198. 7:50bottleneck. Now, we move to task four,
  199. 7:53which is about paged attention, the
  200. 7:54solution that vLLM uses. Open
  201. 7:56/root/code/task4_paged_attention.py.
  202. 8:00You need to complete three to-dos. At
  203. 8:02line 29, replace the blank with 16 to
  204. 8:05set the page size. At line 51, replace
  205. 8:08the blank with page size to calculate
  206. 8:10how many pages each request [music]
  207. 8:12needs. At line 66, replace the blank
  208. 8:14with total pages allocated to calculate
  209. 8:17the paged memory utilization. Run the
  210. 8:19script with
  211. 8:20python/root/code/task4_paged_attention.py.
  212. 8:24The output compares contiguous
  213. 8:26allocation against paged allocation.
  214. 8:28Instead of pre-allocating worst-case
  215. 8:30memory, paged attention uses small
  216. 8:32fixed-size pages allocated on demand,
  217. 8:35just like how operating system manage
  218. 8:37virtual memory. Memory utilization jumps
  219. 8:40from about 20% to about 95%. [music]
  220. 8:42This means four to five times more
  221. 8:44concurrent users can be served with the
  222. 8:46same hardware. We have a knowledge check
  223. 8:49about paged attention. The question
  224. 8:51asks, "What operating system concept
  225. 8:53[music] inspired vLLM paged attention
  226. 8:55mechanism?" The answer is virtual memory
  227. 8:57paging because just like an OS divides
  228. 8:59RAM into 4KB pages allocated on demand,
  229. 9:02vLLM divides the KV cache into
  230. 9:04token-size pages allocated on demand.
  231. 9:07This eliminates fragmentation and over
  232. 9:10allocation. Now, we move to task five,
  233. 9:12which is about launching vLLM as an
  234. 9:14OpenAI compatible API [music] server.
  235. 9:16Open /root/code/task5_api_server.py.
  236. 9:20You need to complete two to-dos. At line
  237. 9:2399, replace the blank with localhost at
  238. 9:25port 8000 version one for the base URL
  239. 9:28[music] and not needed for the API key
  240. 9:30to configure the OpenAI client. At line
  241. 9:33105, replace the blank with Hugging Face
  242. 9:35TB small LLM 135M to set the model for
  243. 9:39the completion request. Run the script
  244. 9:41with
  245. 9:42python/root/code/task5_api_server.py.
  246. 9:46The script starts with vLLM server in
  247. 9:48the background, waits for it to be
  248. 9:50ready, [music] and then sends a test
  249. 9:52request using the standard OpenAI SDK.
  250. 9:55This means any application that already
  251. 9:57works with the OpenAI API can switch
  252. 9:59[music]
  253. 10:00to your self-hosted vLLM server by just
  254. 10:03changing the base URL. The server stays
  255. 10:06running for the remaining tasks. Now, we
  256. 10:08move to task six, which is about stress
  257. 10:10testing the server with concurrent
  258. 10:12users. Open
  259. 10:13/root/code/task6_multiuser_load.py.
  260. 10:18You need to complete two to-dos. At line
  261. 10:2097, replace the blank with 1 5 10 20 to
  262. 10:24define the list of concurrent user
  263. 10:26counts to test. At line 117, replace the
  264. 10:29blank with total tokens divided by total
  265. 10:31time to calculate the aggregate
  266. 10:33throughput. Run the script with
  267. 10:35python/root/code/task6_multiuser_load.py.
  268. 10:40The output shows a load test table with
  269. 10:43results for each concurrency level. As
  270. 10:46the number of concurrent users
  271. 10:47increases, total throughput goes up
  272. 10:50because the hardware is being used more
  273. 10:52efficiently through continuous batching.
  274. 10:55Per request latency increases slightly,
  275. 10:58but the overall system produces more
  276. 11:00tokens per second. We have a knowledge
  277. 11:02check about throughput. The question
  278. 11:04asked, "What is the main reason vLLM
  279. 11:06achieves high throughput than naive
  280. 11:09inference engines when serving multiple
  281. 11:11users?" The answer is, "vLLM efficiently
  282. 11:13manages the KV cache with paged
  283. 11:15attention,
  284. 11:16>> [music]
  285. 11:16>> reducing memory waste because by using
  286. 11:18paging instead of contiguous
  287. 11:20pre-allocation, more concurrent requests
  288. 11:22fit in memory at once and the hardware
  289. 11:26stays busier processing tokens. Now, we
  290. 11:28move to task seven, which is about
  291. 11:30tuning vLLM parameters for production.
  292. 11:33Open /root/code/task7_tuning.py.
  293. 11:37You need to complete two to-dos. At line
  294. 11:39164,
  295. 11:41replace the blank with 64 to set a
  296. 11:43shorter context length for the
  297. 11:45configuration. At line 173, replace the
  298. 11:48blank with eight to limit the number of
  299. 11:50concurrent sequences. Run the script
  300. 11:52with python/root/code/
  301. 11:55task7_tuning.py.
  302. 11:56[music] The script restarts the vLLM
  303. 11:58server with different configurations and
  304. 12:01benchmarks [music]
  305. 12:01each one. You will see how lowering max
  306. 12:04model length reduces memory per request
  307. 12:07while limiting max number steps controls
  308. 12:09[music]
  309. 12:09how many requests are processed at once.
  310. 12:12The right tuning depends on your
  311. 12:13workload. Short prompts need lower max
  312. 12:16model length and many users need higher
  313. 12:18max number of sequences. Now, we move to
  314. 12:21task eight, which is the capstone.
  315. 12:22[music] Open /root/code/
  316. 12:25task8_dashboard.py.
  317. 12:27You need to complete three to-dos. At
  318. 12:29line 85, replace the blank with
  319. 12:31requests.post to send a test request to
  320. 12:34the vLLM server for live metrics. At
  321. 12:37line 114, replace the blank with
  322. 12:39huggingface_tps vllm_tps to set the
  323. 12:43tokens per second values in the
  324. 12:44comparison chart. [music] At line 119,
  325. 12:47replace the blank with vllm_tps divided
  326. 12:50by huggingface_tps to calculate the
  327. 12:52improvement ratio. Run the script with
  328. 12:55python/root/code/task8_dashboard.py.
  329. 12:56[music]
  330. 12:59Then, click the Gradio UI button in the
  331. 13:01top right to view the dashboard. The
  332. 13:04dashboard shows a comparison chart of
  333. 13:06huggingface versus vLLM performance, the
  334. 13:08improvement ratio, live inference
  335. 13:10metrics, load test results, and tuning
  336. 13:13configuration comparisons. This is a
  337. 13:15simplified version of the monitoring you
  338. 13:18would use in production with tools like
  339. 13:20Prometheus and Grafana. We have a final
  340. 13:23knowledge check about the tradeoffs
  341. 13:25between [music] inference engines. The
  342. 13:26question asked, "When might you choose
  343. 13:29llama.cpp over vLLM?" The answer [music]
  344. 13:31is, "When running inference on CPU/RAM
  345. 13:34without a GPU and [music] needing
  346. 13:36optimized CPU performance because
  347. 13:38llama.cpp is specifically optimized for
  348. 13:41CPU and RAM inference on consumer
  349. 13:43hardware, while vLLM excels at high
  350. 13:46throughput multi-user serving on GPU."
  351. 13:49Before wrapping up, I want to highlight
  352. 13:51a few things. Inference engine matters
  353. 13:53more than you think. The same model runs
  354. 13:56at different speeds depending on the
  355. 13:58engine. KV cache fragmentation wastes 60
  356. 14:00to 80% of memory in traditional systems.
  357. 14:03Paged attention fixes this by borrowing
  358. 14:06the virtual memory paging concept from
  359. 14:08operating systems. vLLM OpenAI
  360. 14:11compatible API means zero code changes
  361. 14:14when migrating from OpenAI. Always tune
  362. 14:16max model length and max number
  363. 14:18sequences for your specific workload.
  364. 14:21Monitor tokens per second and latency in
  365. 14:23production to know when to scale. That
  366. 14:26is it. We went from naive huggingface
  367. 14:28setup that handles one request at a time
  368. 14:31to a production-ready vLLM server with
  369. 14:34live monitoring. We measured baseline
  370. 14:36performance, saw vLLM speed average,
  371. 14:39understood why the KV cache is the
  372. 14:41bottleneck, learned how paged attention
  373. 14:43solves with OS-inspired paging, launched
  374. 14:46an OpenAI compatible API server,
  375. 14:48stress-tested it with up to 20
  376. 14:50concurrent users, tuned the parameters
  377. 14:52for production workloads, and built a
  378. 14:54Gradio monitoring dashboard.
  379. 14:56>> [music]
  380. 14:56>> You now understand why companies use
  381. 14:58inference engines like vLLM to serve
  382. 15:00LLMs efficiently at scale.
  383. 15:03>> [music]

About this transcript

This page contains the full transcript of Understanding vLLM with a Hands On Demo by KodeKloud, generated from the public captions YouTube serves with the video. The transcript has 2,253 words across 383 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.