vLLM and Ray cluster to start LLM on multiple servers with multiple GPUs — Transcript
Full transcript
- 0:00There are large language models. There
- 0:02are small language models. Uh you can
- 0:05train language models and process of
- 0:08using language models asking questions
- 0:11called inferencing. I will try to
- 0:13inference large language model that does
- 0:15not fit on one computer even with
- 0:17multiple GPUs because it needs a lot of
- 0:20video memory. I will try to use DeepSeek
- 0:23R1. It has 671 billion parameters and it
- 0:29is estimated that it will require 16
- 0:31GPUs with 80 GBTE memory each. There are
- 0:35different techniques to optimize large
- 0:36language model distillation and
- 0:39quantization. Distillation is a process
- 0:41that creates a new model when smaller
- 0:43models like llama and quen are learning
- 0:46from the teacher dips R1 and creating
- 0:49smaller distilled models. Quantization
- 0:51is decreasing model numerical precision
- 0:54from for example flow 32 to integer 8 to
- 0:57create quantized model. Different
- 0:59versions can be downloaded from hugging
- 1:00face. I will download this one deepsec
- 1:03car 1. It is possible to run optimized
- 1:05version on a single computer for example
- 1:07to use very popular software or llama.
- 1:09It can run on Linux, Mac and windows and
- 1:12it also support multiple GPUs. It's even
- 1:15possible to run on a smartphone if you
- 1:18download very optimized version. I think
- 1:20on this stage it's called small language
- 1:22model. I tried this pocket pal AI
- 1:24application. It works on iPhone and
- 1:26Android and it can run small language
- 1:29models. But in case of multiple
- 1:31computers software like VLM with ray
- 1:34cluster can inference LLM on multiple
- 1:37computers with multiple GPUs. This is
- 1:40will be set up hardware for servers with
- 1:4216 GPUs. Software Linux 96 with CUDA
- 1:4712.9. I will use default pre-installed
- 1:51Python 39. These IP addresses I will use
- 1:54on these four servers and all servers
- 1:57have mounted shared file system. So all
- 1:59machines can see the same files.
- 2:01Important note if servers have multiple
- 2:03IP addresses. It's very important to set
- 2:06VLM host IP to avoid network related
- 2:09errors. Now brace yourself. After this
- 2:12point there will be many commands. I
- 2:14will add a link to the text version of
- 2:16this video in the video description.
- 2:18Let's change colors for each server. Uh
- 2:22I have Python 39. All of them running on
- 2:25Rocky Linux 96. Each server has four
- 2:28GPUs. In addition, I need to install
- 2:32Python 3 Develop package. Now I will
- 2:35create directory on a shared file system
- 2:38cluster lm. And here I will store all
- 2:43files. Now I'm creating Python
- 2:45environment variable. Activate.
- 2:48Installing array cluster. Done.
- 2:50Installing VLM. Done. Now I'm setting
- 2:54VLM host IP because I have multiple IP
- 2:56addresses on each server. So each server
- 2:59will get its own VLM IP address.
- 3:03Starting ray head node with IP address
- 3:06and port number. I can run ray status.
- 3:09And I can see one node is active and I
- 3:12see four GPUs. Now I'm adding additional
- 3:15nodes to the ray cluster. Ray status
- 3:18again. Now I can see four nodes in the
- 3:20cluster and I can see 16 GPUs on the
- 3:24head node. You can also run array list
- 3:27nodes. It will list IP addresses that
- 3:31are configured to be used in the
- 3:32cluster. Now it is time to download DCR
- 3:351. I will use g clone. Full size after
- 3:38download will be 1.3 terabyte but it is
- 3:41possible to save space just by deleting
- 3:44Git subdirectory which is taking half of
- 3:46the space what is not needed for
- 3:49inferencing this model now I can use VLM
- 3:52to inference model because I use SH file
- 3:55system and it's very large model
- 3:58it takes a lot of time to start in my
- 4:03case it was close to 1 hour at the end
- 4:06you will get application startup
- 4:08complete. Now it's ready to be used. I
- 4:12can access model from a command line. I
- 4:14can use VLM chat with URL to the node
- 4:18where I started it. I started on a head
- 4:21node with IP address 61. Let's draw a
- 4:24cat with ASKI code. I like this cat.
- 4:28This cat looks like ice cream. This look
- 4:30like a sleeping cat, but it writes it's
- 4:33sitting cat. Now I can connect VLM to
- 4:37open web UI. I will use open AI API.
- 4:41What I can find in admin panel settings
- 4:45uh connections
- 4:49add open AI API connection
- 4:52using URL verify. And now I can see
- 4:56deepsear one here. And I can make it
- 5:00public.
- 5:02Save and update. Now I can create new
- 5:06chat. Let's ask how to make a battery in
- 5:08the wild.
- 5:11And while it's answering, we can check
- 5:14on one of the servers. Uh what is the
- 5:17GPU doing?
- 5:21And Smi watch. And all GPUs are very
- 5:26busy. And this is how you run large
- 5:28language model on multiple servers with
- 5:31multiple GPUs.
About this transcript
This page contains the full transcript of vLLM and Ray cluster to start LLM on multiple servers with multiple GPUs by Pavlo Khmel HPC, generated from the public captions YouTube serves with the video. The transcript has 792 words across 115 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.