Understanding vLLM with a Hands On Demo — Transcript
Full transcript
- 0:00One of the biggest topics around LLM is
- 0:02when it comes to inference, which means
- 0:04when we actually serve the model to
- 0:06start generating output tokens. And
- 0:08depending on what system you use to
- 0:10serve LLMs, end users will experience
- 0:13different results. For example, when we
- 0:16go to ChatGPT and interact with the
- 0:18chatbot, it might generate its output at
- 0:21a certain speed. Well, when we go to
- 0:23Gemini, for example, it might generate
- 0:25its output at a much different speed. We
- 0:27measure this in what's called tokens per
- 0:30second. And depending on what system you
- 0:32use to actually run your large language
- 0:34model, the same model can be served at
- 0:36varying speed. And there are tons of
- 0:38systems out there like llama.cpp, vLLM,
- 0:41SGLang, TensorRT-LLM, Hugging Face TGI,
- 0:45LMDeploy. These are all a classification
- 0:48of inference engine that allows a model
- 0:50to be served. vLLM is a system that
- 0:53really stands out as a tool that allows
- 0:56you to run your model with high
- 0:57throughput. When you start to serve LLM
- 1:00beyond a single user or a single
- 1:02instance to multiple users across
- 1:05multiple instances, what starts to
- 1:07happen is that the underlying system
- 1:09that runs LLM needs to serve multiple
- 1:13requests at a time. And since during
- 1:15inference your prompt is stored not only
- 1:17in memory in terms of KV cache and auto
- 1:20regressively appends during its decoding
- 1:23phase, having a system that can actually
- 1:26manage this KV cache efficiently is very
- 1:28important. vLLM is a system that
- 1:31introduced an extremely efficient way to
- 1:33manage the KV cache where at any time it
- 1:36showed a significant improvement over
- 1:38existing systems that often wasted up to
- 1:4160 to 80% of its memory in
- 1:44fragmentation, over allocation, or both.
- 1:47And vLLM introduced what's called the
- 1:49paged attention in order to actually
- 1:51create virtualized KV cache that was
- 1:54inspired by how paging systems work at
- 1:57the operating system level. So, instead
- 1:59of over allocating your space where you
- 2:02prepare for the worst case in terms of
- 2:04how much the model will generate in
- 2:06output, vLLM, instead of padding the
- 2:08memory used in pages that grew in size,
- 2:11which frees up the memory to be used
- 2:14more efficiently as it processes many
- 2:16requests at a time. And not only was
- 2:19this more efficient in the, let's call
- 2:21it, storage Tetris game, it means that
- 2:23the underlying graphics card was busier.
- 2:26But, it's important to keep in mind that
- 2:28vLLM isn't necessarily an end-all
- 2:31solution. It means to solve throughput
- 2:33problem on a multi-user and
- 2:35multi-requests, but other inference
- 2:38engines prioritize other metrics like
- 2:41or optimizing to the vendor in the case
- 2:43of TensorRT-LLM for Nvidia, as well as
- 2:46better optimization support for running
- 2:48the model for less RAM requirements.
- 2:51But, learning how to use vLLM to serve
- 2:53LLM as more and more large language
- 2:56models are innovated, but also being
- 2:58able to serve it beyond a single user
- 3:01instances by sharing them across
- 3:03multiple users from server side is going
- 3:05to be an extremely important skill to
- 3:08monitor and optimize how your hardware
- 3:10is being used to provide services.
- 3:12Learning how to use vLLM is going beyond
- 3:15being impressed by an LLM, but actually
- 3:18trying to productize LLM so that you can
- 3:20monitor its tokens per speed and its
- 3:23outputs and its throughput as the number
- 3:25of sessions grow with various prompt
- 3:28requests in sizes also changes
- 3:30dynamically. And managing this exact
- 3:32stack is what vLLM allows you to work to
- 3:35lower the latency and maximize
- 3:37concurrency using vLLM.
- 3:42In this lab, we're going to learn about
- 3:44how vLLM turns a basic LLM setup into a
- 3:47production-ready inference server.
- 3:49[music] We run small LLM 135M with
- 3:52Hugging Face, then with vLLM and compare
- 3:55the speed. We will visualize the KV
- 3:57cache problem, see how paged attention
- 3:59fixes it, launch an API server, stress
- 4:02test it with concurrent users, tune
- 4:04parameters, and build a monitoring
- 4:06dashboard. This lab takes about 40 to 50
- 4:09minutes to complete. When I open the
- 4:11lab, I'm dropped right into the
- 4:12scenario. Access the lab using the link
- 4:15in the description below and follow
- 4:16along with me. The intro page tells us
- 4:18that we are an ML engineer at Inference
- 4:21IO, a startup building an LLM as a
- 4:23service platform. Our CEO just signed a
- 4:26deal to serve small LLM to multiple
- 4:28concurrent user. Our current Hugging
- 4:30Face setup handles one request at a
- 4:32time. Our mission is to use vLLM to
- 4:36build a production-ready inference
- 4:37server. The lab has a setup step, eight
- 4:40tasks, and four knowledge checks. The
- 4:42tasks cover Hugging Face baseline, vLLM
- 4:45inference, KV cache problem, paged
- 4:47attention, API server, multi-user load
- 4:50testing, parameter tuning, and a
- 4:52monitoring dashboard.
- 4:54>> [music]
- 4:54>> The first step is to verify our
- 4:56environment. Run the following command
- 4:58and then python/root/code
- 4:59verify_environment.py.
- 5:02This script checks that all required
- 5:03packages like vLLM, Transformers, and
- 5:06Gradio are all installed, downloads
- 5:08[music] the small LLM 135M model, runs a
- 5:12quick test generation, and confirms
- 5:14everything is ready. Wait until you see
- 5:16the environment verification complete
- 5:18message. Now, we move to task one, which
- 5:20is about running naive Hugging Face
- 5:22inference to establish a baseline. Open
- 5:24/root/code/task1_huggingface_baseline.py.
- 5:26You have to complete two to-dos. At line
- 5:2830, replace the blank with Hugging Face
- 5:30TB small LLM 135M to load the model, and
- 5:35at line 49, replace the blank with 50 to
- 5:37set the maximum number of new tokens to
- 5:40generate.
- 5:43[music] Run the script with
- 5:44python/root/code/task1_huggingface_baseline.py.
- 5:49The output shows the model loading,
- 5:51generating text, and reporting tokens
- 5:53per second. This number is our baseline
- 5:55that vLLM will beat. Now, we move to
- 5:57task two, which is about running the
- 5:59same model with vLLM. Open
- 6:01/root/code/task2_vllm_inference.py.
- 6:04[music] You need to complete two to-dos.
- 6:07At line 36, replace the blank with
- 6:09Hugging Face TB small LLM 135M to
- 6:12initialize the vLLM engine with the same
- 6:14model. At line 40, replace the blank
- 6:17with 0.7 for temperature and 50 for max
- 6:20tokens to create the sampling
- 6:22parameters. Run the script with
- 6:23python/root/code/task2_vllm_inference.py.
- 6:24[music]
- 6:28The output shows a side-by-side
- 6:30comparison of Hugging Face versus vLLM
- 6:33tokens per second. Notice that vLLM is
- 6:35faster, even though it is the exact same
- 6:37model and same prompt. This proves that
- 6:40the inference engine matters, [music]
- 6:42not just the model. We have a knowledge
- 6:44check about the inference basics. The
- 6:46question asks, "What is a primary metric
- 6:48used to measure how fast an LLM
- 6:50generates output?" [music] The answer is
- 6:52tokens per second because during
- 6:54inference the model produces tokens one
- 6:56at a time through auto regressive
- 6:58decoding. So, tokens per second directly
- 7:01tells [music] you how fast user sees
- 7:03responses. Now, we move to task three,
- 7:05which is about understanding the KV
- 7:06cache problem. Open
- 7:08/root/code/task3_kvcache_problem.py.
- 7:11[music] You need to complete two to-dos.
- 7:13At line 26, replace the blank with 512
- 7:16to set the maximum sequence length. At
- 7:19line 52, replace the blank with
- 7:21allocated minus actual divided by
- 7:23allocated times 100 to calculate the
- 7:25wasted percentage. Run the script with
- 7:28python/root/code/task3_kvcache_problem.py.
- 7:32The output shows how traditional system
- 7:34pre-allocate worst-case memory for every
- 7:37request. Short prompts waste [music]
- 7:39massive amounts of space because they
- 7:41get the same allocation as long as
- 7:43prompts. You will see that memory
- 7:45utilization is only about 20%, meaning
- 7:4780% is wasted. This is a KV cache
- 7:50bottleneck. Now, we move to task four,
- 7:53which is about paged attention, the
- 7:54solution that vLLM uses. Open
- 7:56/root/code/task4_paged_attention.py.
- 8:00You need to complete three to-dos. At
- 8:02line 29, replace the blank with 16 to
- 8:05set the page size. At line 51, replace
- 8:08the blank with page size to calculate
- 8:10how many pages each request [music]
- 8:12needs. At line 66, replace the blank
- 8:14with total pages allocated to calculate
- 8:17the paged memory utilization. Run the
- 8:19script with
- 8:20python/root/code/task4_paged_attention.py.
- 8:24The output compares contiguous
- 8:26allocation against paged allocation.
- 8:28Instead of pre-allocating worst-case
- 8:30memory, paged attention uses small
- 8:32fixed-size pages allocated on demand,
- 8:35just like how operating system manage
- 8:37virtual memory. Memory utilization jumps
- 8:40from about 20% to about 95%. [music]
- 8:42This means four to five times more
- 8:44concurrent users can be served with the
- 8:46same hardware. We have a knowledge check
- 8:49about paged attention. The question
- 8:51asks, "What operating system concept
- 8:53[music] inspired vLLM paged attention
- 8:55mechanism?" The answer is virtual memory
- 8:57paging because just like an OS divides
- 8:59RAM into 4KB pages allocated on demand,
- 9:02vLLM divides the KV cache into
- 9:04token-size pages allocated on demand.
- 9:07This eliminates fragmentation and over
- 9:10allocation. Now, we move to task five,
- 9:12which is about launching vLLM as an
- 9:14OpenAI compatible API [music] server.
- 9:16Open /root/code/task5_api_server.py.
- 9:20You need to complete two to-dos. At line
- 9:2399, replace the blank with localhost at
- 9:25port 8000 version one for the base URL
- 9:28[music] and not needed for the API key
- 9:30to configure the OpenAI client. At line
- 9:33105, replace the blank with Hugging Face
- 9:35TB small LLM 135M to set the model for
- 9:39the completion request. Run the script
- 9:41with
- 9:42python/root/code/task5_api_server.py.
- 9:46The script starts with vLLM server in
- 9:48the background, waits for it to be
- 9:50ready, [music] and then sends a test
- 9:52request using the standard OpenAI SDK.
- 9:55This means any application that already
- 9:57works with the OpenAI API can switch
- 9:59[music]
- 10:00to your self-hosted vLLM server by just
- 10:03changing the base URL. The server stays
- 10:06running for the remaining tasks. Now, we
- 10:08move to task six, which is about stress
- 10:10testing the server with concurrent
- 10:12users. Open
- 10:13/root/code/task6_multiuser_load.py.
- 10:18You need to complete two to-dos. At line
- 10:2097, replace the blank with 1 5 10 20 to
- 10:24define the list of concurrent user
- 10:26counts to test. At line 117, replace the
- 10:29blank with total tokens divided by total
- 10:31time to calculate the aggregate
- 10:33throughput. Run the script with
- 10:35python/root/code/task6_multiuser_load.py.
- 10:40The output shows a load test table with
- 10:43results for each concurrency level. As
- 10:46the number of concurrent users
- 10:47increases, total throughput goes up
- 10:50because the hardware is being used more
- 10:52efficiently through continuous batching.
- 10:55Per request latency increases slightly,
- 10:58but the overall system produces more
- 11:00tokens per second. We have a knowledge
- 11:02check about throughput. The question
- 11:04asked, "What is the main reason vLLM
- 11:06achieves high throughput than naive
- 11:09inference engines when serving multiple
- 11:11users?" The answer is, "vLLM efficiently
- 11:13manages the KV cache with paged
- 11:15attention,
- 11:16>> [music]
- 11:16>> reducing memory waste because by using
- 11:18paging instead of contiguous
- 11:20pre-allocation, more concurrent requests
- 11:22fit in memory at once and the hardware
- 11:26stays busier processing tokens. Now, we
- 11:28move to task seven, which is about
- 11:30tuning vLLM parameters for production.
- 11:33Open /root/code/task7_tuning.py.
- 11:37You need to complete two to-dos. At line
- 11:39164,
- 11:41replace the blank with 64 to set a
- 11:43shorter context length for the
- 11:45configuration. At line 173, replace the
- 11:48blank with eight to limit the number of
- 11:50concurrent sequences. Run the script
- 11:52with python/root/code/
- 11:55task7_tuning.py.
- 11:56[music] The script restarts the vLLM
- 11:58server with different configurations and
- 12:01benchmarks [music]
- 12:01each one. You will see how lowering max
- 12:04model length reduces memory per request
- 12:07while limiting max number steps controls
- 12:09[music]
- 12:09how many requests are processed at once.
- 12:12The right tuning depends on your
- 12:13workload. Short prompts need lower max
- 12:16model length and many users need higher
- 12:18max number of sequences. Now, we move to
- 12:21task eight, which is the capstone.
- 12:22[music] Open /root/code/
- 12:25task8_dashboard.py.
- 12:27You need to complete three to-dos. At
- 12:29line 85, replace the blank with
- 12:31requests.post to send a test request to
- 12:34the vLLM server for live metrics. At
- 12:37line 114, replace the blank with
- 12:39huggingface_tps vllm_tps to set the
- 12:43tokens per second values in the
- 12:44comparison chart. [music] At line 119,
- 12:47replace the blank with vllm_tps divided
- 12:50by huggingface_tps to calculate the
- 12:52improvement ratio. Run the script with
- 12:55python/root/code/task8_dashboard.py.
- 12:56[music]
- 12:59Then, click the Gradio UI button in the
- 13:01top right to view the dashboard. The
- 13:04dashboard shows a comparison chart of
- 13:06huggingface versus vLLM performance, the
- 13:08improvement ratio, live inference
- 13:10metrics, load test results, and tuning
- 13:13configuration comparisons. This is a
- 13:15simplified version of the monitoring you
- 13:18would use in production with tools like
- 13:20Prometheus and Grafana. We have a final
- 13:23knowledge check about the tradeoffs
- 13:25between [music] inference engines. The
- 13:26question asked, "When might you choose
- 13:29llama.cpp over vLLM?" The answer [music]
- 13:31is, "When running inference on CPU/RAM
- 13:34without a GPU and [music] needing
- 13:36optimized CPU performance because
- 13:38llama.cpp is specifically optimized for
- 13:41CPU and RAM inference on consumer
- 13:43hardware, while vLLM excels at high
- 13:46throughput multi-user serving on GPU."
- 13:49Before wrapping up, I want to highlight
- 13:51a few things. Inference engine matters
- 13:53more than you think. The same model runs
- 13:56at different speeds depending on the
- 13:58engine. KV cache fragmentation wastes 60
- 14:00to 80% of memory in traditional systems.
- 14:03Paged attention fixes this by borrowing
- 14:06the virtual memory paging concept from
- 14:08operating systems. vLLM OpenAI
- 14:11compatible API means zero code changes
- 14:14when migrating from OpenAI. Always tune
- 14:16max model length and max number
- 14:18sequences for your specific workload.
- 14:21Monitor tokens per second and latency in
- 14:23production to know when to scale. That
- 14:26is it. We went from naive huggingface
- 14:28setup that handles one request at a time
- 14:31to a production-ready vLLM server with
- 14:34live monitoring. We measured baseline
- 14:36performance, saw vLLM speed average,
- 14:39understood why the KV cache is the
- 14:41bottleneck, learned how paged attention
- 14:43solves with OS-inspired paging, launched
- 14:46an OpenAI compatible API server,
- 14:48stress-tested it with up to 20
- 14:50concurrent users, tuned the parameters
- 14:52for production workloads, and built a
- 14:54Gradio monitoring dashboard.
- 14:56>> [music]
- 14:56>> You now understand why companies use
- 14:58inference engines like vLLM to serve
- 15:00LLMs efficiently at scale.
- 15:03>> [music]
About this transcript
This page contains the full transcript of Understanding vLLM with a Hands On Demo by KodeKloud, generated from the public captions YouTube serves with the video. The transcript has 2,253 words across 383 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.