YouTube2Text

vLLM and Ray cluster to start LLM on multiple servers with multiple GPUs — Transcript

by Pavlo Khmel HPC · 792 words · 115 segments · language en · Watch on YouTube

Full transcript

  1. 0:00There are large language models. There
  2. 0:02are small language models. Uh you can
  3. 0:05train language models and process of
  4. 0:08using language models asking questions
  5. 0:11called inferencing. I will try to
  6. 0:13inference large language model that does
  7. 0:15not fit on one computer even with
  8. 0:17multiple GPUs because it needs a lot of
  9. 0:20video memory. I will try to use DeepSeek
  10. 0:23R1. It has 671 billion parameters and it
  11. 0:29is estimated that it will require 16
  12. 0:31GPUs with 80 GBTE memory each. There are
  13. 0:35different techniques to optimize large
  14. 0:36language model distillation and
  15. 0:39quantization. Distillation is a process
  16. 0:41that creates a new model when smaller
  17. 0:43models like llama and quen are learning
  18. 0:46from the teacher dips R1 and creating
  19. 0:49smaller distilled models. Quantization
  20. 0:51is decreasing model numerical precision
  21. 0:54from for example flow 32 to integer 8 to
  22. 0:57create quantized model. Different
  23. 0:59versions can be downloaded from hugging
  24. 1:00face. I will download this one deepsec
  25. 1:03car 1. It is possible to run optimized
  26. 1:05version on a single computer for example
  27. 1:07to use very popular software or llama.
  28. 1:09It can run on Linux, Mac and windows and
  29. 1:12it also support multiple GPUs. It's even
  30. 1:15possible to run on a smartphone if you
  31. 1:18download very optimized version. I think
  32. 1:20on this stage it's called small language
  33. 1:22model. I tried this pocket pal AI
  34. 1:24application. It works on iPhone and
  35. 1:26Android and it can run small language
  36. 1:29models. But in case of multiple
  37. 1:31computers software like VLM with ray
  38. 1:34cluster can inference LLM on multiple
  39. 1:37computers with multiple GPUs. This is
  40. 1:40will be set up hardware for servers with
  41. 1:4216 GPUs. Software Linux 96 with CUDA
  42. 1:4712.9. I will use default pre-installed
  43. 1:51Python 39. These IP addresses I will use
  44. 1:54on these four servers and all servers
  45. 1:57have mounted shared file system. So all
  46. 1:59machines can see the same files.
  47. 2:01Important note if servers have multiple
  48. 2:03IP addresses. It's very important to set
  49. 2:06VLM host IP to avoid network related
  50. 2:09errors. Now brace yourself. After this
  51. 2:12point there will be many commands. I
  52. 2:14will add a link to the text version of
  53. 2:16this video in the video description.
  54. 2:18Let's change colors for each server. Uh
  55. 2:22I have Python 39. All of them running on
  56. 2:25Rocky Linux 96. Each server has four
  57. 2:28GPUs. In addition, I need to install
  58. 2:32Python 3 Develop package. Now I will
  59. 2:35create directory on a shared file system
  60. 2:38cluster lm. And here I will store all
  61. 2:43files. Now I'm creating Python
  62. 2:45environment variable. Activate.
  63. 2:48Installing array cluster. Done.
  64. 2:50Installing VLM. Done. Now I'm setting
  65. 2:54VLM host IP because I have multiple IP
  66. 2:56addresses on each server. So each server
  67. 2:59will get its own VLM IP address.
  68. 3:03Starting ray head node with IP address
  69. 3:06and port number. I can run ray status.
  70. 3:09And I can see one node is active and I
  71. 3:12see four GPUs. Now I'm adding additional
  72. 3:15nodes to the ray cluster. Ray status
  73. 3:18again. Now I can see four nodes in the
  74. 3:20cluster and I can see 16 GPUs on the
  75. 3:24head node. You can also run array list
  76. 3:27nodes. It will list IP addresses that
  77. 3:31are configured to be used in the
  78. 3:32cluster. Now it is time to download DCR
  79. 3:351. I will use g clone. Full size after
  80. 3:38download will be 1.3 terabyte but it is
  81. 3:41possible to save space just by deleting
  82. 3:44Git subdirectory which is taking half of
  83. 3:46the space what is not needed for
  84. 3:49inferencing this model now I can use VLM
  85. 3:52to inference model because I use SH file
  86. 3:55system and it's very large model
  87. 3:58it takes a lot of time to start in my
  88. 4:03case it was close to 1 hour at the end
  89. 4:06you will get application startup
  90. 4:08complete. Now it's ready to be used. I
  91. 4:12can access model from a command line. I
  92. 4:14can use VLM chat with URL to the node
  93. 4:18where I started it. I started on a head
  94. 4:21node with IP address 61. Let's draw a
  95. 4:24cat with ASKI code. I like this cat.
  96. 4:28This cat looks like ice cream. This look
  97. 4:30like a sleeping cat, but it writes it's
  98. 4:33sitting cat. Now I can connect VLM to
  99. 4:37open web UI. I will use open AI API.
  100. 4:41What I can find in admin panel settings
  101. 4:45uh connections
  102. 4:49add open AI API connection
  103. 4:52using URL verify. And now I can see
  104. 4:56deepsear one here. And I can make it
  105. 5:00public.
  106. 5:02Save and update. Now I can create new
  107. 5:06chat. Let's ask how to make a battery in
  108. 5:08the wild.
  109. 5:11And while it's answering, we can check
  110. 5:14on one of the servers. Uh what is the
  111. 5:17GPU doing?
  112. 5:21And Smi watch. And all GPUs are very
  113. 5:26busy. And this is how you run large
  114. 5:28language model on multiple servers with
  115. 5:31multiple GPUs.

About this transcript

This page contains the full transcript of vLLM and Ray cluster to start LLM on multiple servers with multiple GPUs by Pavlo Khmel HPC, generated from the public captions YouTube serves with the video. The transcript has 792 words across 115 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.