AI-Assisted Data Extraction | HubMeta Tutorial #15 — Transcript
Full transcript
- 0:15Welcome back. Now, we have set up our
- 0:18codebook for our project and we are all
- 0:22ready to start capturing data. Now, if
- 0:26you want to capture the data manually,
- 0:30you can of course go ahead and do that.
- 0:32All you got to do is to come inside a
- 0:35paper
- 0:36to present this. I will go ahead and
- 0:38delete the existing data that I have for
- 0:41my project. And then, you will see that
- 0:44on the right I have my PDF and here I
- 0:47can say, "Okay, let's start by adding a
- 0:51first sample." It's done. Now, what do I
- 0:55want to extract? I want to, let's say, a
- 0:59study source information, which is first
- 1:02author,
- 1:04for example, last name. In this case, I
- 1:07will just start manually doing that.
- 1:09What is the year? It is 2019.
- 1:13And then, the title, I will
- 1:18copy
- 1:20and paste here. And you get the idea.
- 1:23So, you will
- 1:25manually enter them one by one.
- 1:28You can even add a correlation matrix.
- 1:32Here, you know, you will add row by row.
- 1:35You define a variable. You add the
- 1:38correlational values, you know, all the
- 1:41other things that you want. And with
- 1:44that, you will manually enter your
- 1:47stuff.
- 1:48But, we have a much better way and
- 1:50that's the exciting part. So, in this
- 1:53case, I have the PDF, I have the
- 1:56templates that I want to extract. And
- 1:58now, I'm going to use the AI extract
- 2:02function that we have here.
- 2:05This will open up a model for me, which
- 2:09allows me to select, all right, which of
- 2:12this data you want to extract. So, let's
- 2:15select all of them.
- 2:18Um I will just uncheck this last one.
- 2:20This If you remember, this is just a
- 2:22sample one that we created
- 2:25uh just to test this. And then, of
- 2:28course, check both correlation and
- 2:32regression table.
- 2:34And here, down here, we can select the
- 2:38large language model that will be used
- 2:41in the process of extraction.
- 2:44One thing to note is how many AI credits
- 2:47will be used in this. This is the most
- 2:50expensive part of AI using across all of
- 2:54HubMed data. Because, let me explain
- 2:57what's happening after I uh hit
- 2:59extraction because it takes like two or
- 3:02three minutes to extract one paper.
- 3:05So, what is happening now is that HubMed
- 3:08goes through multiple steps to capture
- 3:11the data. First, it will convert your
- 3:14PDF and all of its data into a readable
- 3:18format. Into a format that is not
- 3:21picture, that is not, you know, um some
- 3:25weird numbers, into clear
- 3:28uh extractable text.
- 3:31Uh
- 3:32then, it will send it through a
- 3:34multi-agent
- 3:36extraction process through multiple
- 3:39rounds of calls to large language
- 3:41models. And this is something that we
- 3:44have been working on for a few months.
- 3:48And I I think this extraction
- 3:52uh is going to be a standardized
- 3:56in the future. We will definitely keep
- 3:59improving our process, but by our count
- 4:02and with our test uh
- 4:05the the version that we have on Hub Meta
- 4:07working right now as of May 2026
- 4:11is pretty accurate. It's It's really
- 4:13good in capturing the full correlation
- 4:16table, the actual data from the
- 4:19templates, uh defining all the variables
- 4:23and measurements. So, it's doing a very
- 4:25good job uh overall.
- 4:29So, how long does it take? It usually
- 4:32takes around anywhere between 2 to 5
- 4:36minutes per paper to to capture all of
- 4:39that. And that is just how long it takes
- 4:43the AI models to come back with the
- 4:46response after all of these things that
- 4:48we are asking them. So, we have to be
- 4:51patient here, but let's review the
- 4:55results after one study is captured. If
- 4:58you were happy with the extraction, we
- 5:00will move on to the final stage in prep,
- 5:03which is selecting multiple papers and
- 5:07then extracting their data with AI
- 5:11uh all at the same time.
- 5:14Perfect. Now AI has finished extracting
- 5:18the data that we want, and this is how
- 5:21we will see it. This is for us to review
- 5:24the extraction and fix any issues that
- 5:27we find
- 5:29um like uh you know, my
- 5:32template by template it shows uh like
- 5:35what has been extracted like the last
- 5:38name, the year, um
- 5:40the type of publication, you know,
- 5:43number of firms, number of observations,
- 5:46like all the things that we wanted to
- 5:49include here. So, that is as much as AI
- 5:53could actually extract. Down here, we
- 5:57also see a correlation table. You will
- 6:00see that mean and a standard deviation
- 6:03for eight variables were extracted.
- 6:07And there are a lot of other options
- 6:09which we need not have data in this
- 6:12project, but here is the actual
- 6:15correlation table and the actual values.
- 6:17We have implemented a lot of uh checks
- 6:21here to make sure that these extracted
- 6:24values are correct, like, you know, no
- 6:27value uh
- 6:28higher than one or lower than minus one.
- 6:31Uh we hope that this are these are good.
- 6:34Like, we have tested it in a few
- 6:35projects and so far it's been very good.
- 6:38Down here, the and this part is very
- 6:42very important, is
- 6:44the definition of these variables, of
- 6:48the measurements, according to the
- 6:50paper. So, you can see, for example,
- 6:53share of foreign outsourcing, how is
- 6:57that defined in the paper? So, it says
- 7:01uh ratio of total cost of foreign
- 7:04outsourcing. So, this is the exact
- 7:06definition as was used in the paper. So,
- 7:08basically, AI reads the text of the
- 7:11paper and says how this specific
- 7:15variable was defined. And then it picks
- 7:18that. And then, when we save this data,
- 7:22it will save all of these variables into
- 7:25our taxonomy, in the taxonomy of
- 7:28variables, which is something we will
- 7:30get to in the next video.
- 7:33And then, after all of these are done,
- 7:36we will see our regression table. Many
- 7:40papers also report a regression
- 7:42analysis. So, multiple uh versions. And
- 7:45then down here you can give us some
- 7:48feedback like how accurate was this
- 7:51compared to your actual paper.
- 7:53You can click view PDF and then see all
- 7:57of these data side by side to make a
- 7:59comparison and see if everything is
- 8:04accurate. And when you were happy, you
- 8:07will hit save results. And then these
- 8:11data will appear inside this page
- 8:15like it was before, right? So, our data
- 8:20is extracted with AI. This usually took
- 8:23anywhere between two to three hours up
- 8:26to a half day from from a trained
- 8:29research assistant to extract all of
- 8:31this data. Now it all happened within a
- 8:34few minutes. Now, if we want to do this
- 8:37not on just one paper, all we got to do
- 8:41is go to prep AI extract and then we get
- 8:46started. Here again we select the
- 8:49templates we want to extract data from.
- 8:52These are the the ones that we set in
- 8:54our code book. Then we continue. We pick
- 8:57our language model. By default Haiku
- 9:004.5. If you don't have a reason to
- 9:02change this, please don't. This is the
- 9:04one that gives us the most accurate
- 9:06ones. This project has already had all
- 9:10of its data extracted. So,
- 9:13you know,
- 9:14you know, if this hide extracted one
- 9:17is selected, it will not show them.
- 9:20Uncheck them if you want. You can select
- 9:23one, two, three more or, you know, use
- 9:26this and,
- 9:27you know, select more papers to extract
- 9:30all at one or just select all and do all
- 9:32of that. You can also
- 9:34do some different types of filters like
- 9:38let's just look at the ones that are
- 9:40missing a correlation table and missing
- 9:43a regression, so no data has been
- 9:45extracted. Let's just, for example,
- 9:47select all of these. And when when you
- 9:50select these papers, you continue, it
- 9:53again says how much AI credits this is
- 9:56going to cost you. And then you can
- 9:58start AI batch screening, and then it
- 10:01will go paper by paper and extract all
- 10:03of the data. Uh of course,
- 10:06um you might not have enough credits. If
- 10:08you don't, um this will warn you there
- 10:12and will stop after your credits run out
- 10:14and save the data up to that place. To
- 10:17add credits
- 10:19uh to your account, you go to this add
- 10:22credit button. This will take you to a
- 10:24payment step, and then you can add AI
- 10:27credits. As mentioned before, you only
- 10:29pay for what you use, and and only for
- 10:33the functions inside have meta that are
- 10:36using AI heavily, and we have to pay
- 10:39external parties uh for that. In the
- 10:42next video, we will talk about what we
- 10:45should do with data that is now
- 10:47extracted.
- 10:58>> [music]
About this transcript
This page contains the full transcript of AI-Assisted Data Extraction | HubMeta Tutorial #15 by HubMeta, generated from the public captions YouTube serves with the video. The transcript has 1,457 words across 216 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.