How to Build a Text Summarizer using Huggingface Transformers | Turingtalks — Transcript
Full transcript
- 0:00hello guys welcome to this channel in
- 0:02today's video we're going to see how to
- 0:04build a tech summarizer using hugging
- 0:06phas
- 0:07Transformers so in this video I'll walk
- 0:09you through what a summarizer is its use
- 0:11cases what huging face Transformers are
- 0:14and how you can build your own text
- 0:15summarizer using huging face
- 0:16Transformers Library so let's get
- 0:19started a summarizer does exactly what
- 0:21the name suggests it takes a large block
- 0:24of text and condenses it into a shorter
- 0:26version so this shorter version keeps
- 0:28only the key points think think of it as
- 0:30the difference between reading a whole
- 0:31novel and glancing at its back cover so
- 0:34the aim is to save time while still
- 0:36getting the essence of the content now
- 0:38where do we use the summarizer for
- 0:40example journalists use them to quickly
- 0:42sift through reports and studies
- 0:44students can use them to summarize
- 0:46lengthy readings businesses can use them
- 0:48to condense market analysis or lengthy
- 0:50reports in essence anyone who needs to
- 0:53process large amounts of text can
- 0:55quickly benefit from a
- 0:57summarizer now what are hugin phace
- 0:58Transformers hugging face is a company
- 1:01that has created a state-of-the-art
- 1:03platform for natural language processing
- 1:05they have a library called the
- 1:07Transformers Library it is like a
- 1:09treasure toe for NLP tasks it includes
- 1:11pre-train models that can do everything
- 1:13from translation sentiment analysis and
- 1:15yes summarization so these models have
- 1:18learned from vast amounts of text and
- 1:20can understand and generate language in
- 1:21a surprisingly human-like way so we
- 1:23don't have to do a lot of the training
- 1:24in these models these are pre-trained
- 1:26models that we can take them and even
- 1:28fine-tune them to use it for our needs
- 1:31now let's see how we can build a
- 1:33summarizer using the hugging face
- 1:34Transformers model for this example we
- 1:37will be using Facebook's Bart model so
- 1:39Bart model is a pre-trained model on the
- 1:42English language so it is a sequence to
- 1:44sequence model and it is great for uh
- 1:45text generation for example
- 1:47summarization translation but it also
- 1:49works well for comprehension tasks like
- 1:51text classification question answering
- 1:53and so many others so a sequence to
- 1:54sequence model is where you input a
- 1:56sequence and it outputs a sequence for
- 1:58example Chach PT Hing Fai also has a
- 2:01concept called pipelines so pipelines
- 2:03offer a simpler approach to implementing
- 2:05various tasks so instead of preparing a
- 2:07data set training it with the model and
- 2:08then using it pipeline simplifies the
- 2:10code because it HIDs away a lot of the
- 2:13manual work uh like tokenization and
- 2:15model customization and you can just get
- 2:17started working with the model so we
- 2:18will be using a Google collab notebook
- 2:20for this project you can also find a
- 2:22link for the finished project in the
- 2:23description below so let's open the
- 2:25notebook and let's first install
- 2:27Transformers so installing Transformers
- 2:29will be a sh command whereas colloud
- 2:31notebooks read python so we will add an
- 2:33exclamation mark at the beginning so pip
- 2:35install Transformers let's wait for the
- 2:38Transformers library to be installed
- 2:40great Transformers library is now
- 2:42installed so now let's initialize a text
- 2:45summarization by bline using the huging
- 2:47face Transformers Library here I'll be
- 2:49typing summarizer equal to by blind
- 2:51summarizer model Facebook but L CNN so
- 2:54let's see what each part does the
- 2:57pipeline is a function provided by the
- 2:58huging face Transformers SL to make it
- 3:00easy to apply different types of natural
- 3:02language processing tasks so the
- 3:04function returns a ready to use pipeline
- 3:06object for the specified tasks
- 3:08summarization is the first argument we
- 3:09are telling the pipeline that task that
- 3:11we are going to perform is summarization
- 3:13and the model here is going to be
- 3:15Facebook's bot so we'll be using
- 3:17facebook/ bot CNN so this argument
- 3:20specifies the pre-train model that will
- 3:23be used for this project for this
- 3:24summarization task is Facebook's B Cloud
- 3:28CNN so this mod is provided by Facebook
- 3:30and it is placed on the bot model bot
- 3:32architecture which is effective for uh
- 3:34General NLP tasks so let's execute the
- 3:37code how creates a summarizer object so
- 3:39this object can be used to perform text
- 3:41summarization by passing Text data to it
- 3:44the model will generate a shorter
- 3:45version of the input text and it will
- 3:47capture the most important or relevant
- 3:49information according to its training on
- 3:50the summarization tasks so yeah we are
- 3:52now ready to use the model uh that's all
- 3:55it takes to build a summarization model
- 3:57using hugging the face pipeline so you
- 3:58can see how easy it is to build build
- 4:00NLP models using hugging face
- 4:02Transformers so we will fetch an anal
- 4:04report from FDA here it
- 4:06is so let's copy the text and create a
- 4:10variable called text and put in our text
- 4:13now let's use this text as input and
- 4:14call our summarizer here I'll be writing
- 4:16summarizer equal to summarizer text Max
- 4:18length Min length and do
- 4:22sample this line of code is using the
- 4:24summarizer object created from a hugging
- 4:26face pipeline to generate a summary of
- 4:28the input text so summarizer is the
- 4:30object initialized previously with the
- 4:32bip planine function the text we have
- 4:34already declared we setting the Max and
- 4:36Min length so the max length specifies
- 4:39the maximum length uh obviously and uh
- 4:42it'll be in terms of the number of
- 4:43tokens it's not the number of characters
- 4:45it's the number of tokens and the Min
- 4:46length is also the minimum length of the
- 4:48summary so we don't want the summary to
- 4:50be too short so now we can go ahead and
- 4:53print the summary and here's the
- 4:55response so in short this code loads a
- 4:57summarization pipeline that is
- 4:58preconfigured to use the Facebook B
- 5:00model feeds the text to A summarizer
- 5:03outputs a summary with a specified
- 5:04minimum and maximum length so you can
- 5:06see how easy it is to work with the
- 5:08hugging face Transformers Library there
- 5:10are many other features or NLP functions
- 5:13that are offered by hugging face
- 5:14Transformers Library so I'll be posting
- 5:16more videos on those topics if you want
- 5:18to get daily AI tutorials we have a
- 5:21newsletter called touring talks you can
- 5:23find the newsletter at during talks. you
- 5:26can also subscribe to this channel to
- 5:28get notified when we post new videos
- 5:30thanks for watching if you have any
- 5:32questions please let us know in the
- 5:33comments see you soon with a new topic
- 5:35cheers
About this transcript
This page contains the full transcript of How to Build a Text Summarizer using Huggingface Transformers | Turingtalks by TuringTalks, generated from the public captions YouTube serves with the video. The transcript has 1,088 words across 164 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.