YouTube2Text

AI-Assisted Data Extraction | HubMeta Tutorial #15 — Transcript

by HubMeta · 1,457 words · 216 segments · language en · Watch on YouTube

Full transcript

  1. 0:15Welcome back. Now, we have set up our
  2. 0:18codebook for our project and we are all
  3. 0:22ready to start capturing data. Now, if
  4. 0:26you want to capture the data manually,
  5. 0:30you can of course go ahead and do that.
  6. 0:32All you got to do is to come inside a
  7. 0:35paper
  8. 0:36to present this. I will go ahead and
  9. 0:38delete the existing data that I have for
  10. 0:41my project. And then, you will see that
  11. 0:44on the right I have my PDF and here I
  12. 0:47can say, "Okay, let's start by adding a
  13. 0:51first sample." It's done. Now, what do I
  14. 0:55want to extract? I want to, let's say, a
  15. 0:59study source information, which is first
  16. 1:02author,
  17. 1:04for example, last name. In this case, I
  18. 1:07will just start manually doing that.
  19. 1:09What is the year? It is 2019.
  20. 1:13And then, the title, I will
  21. 1:18copy
  22. 1:20and paste here. And you get the idea.
  23. 1:23So, you will
  24. 1:25manually enter them one by one.
  25. 1:28You can even add a correlation matrix.
  26. 1:32Here, you know, you will add row by row.
  27. 1:35You define a variable. You add the
  28. 1:38correlational values, you know, all the
  29. 1:41other things that you want. And with
  30. 1:44that, you will manually enter your
  31. 1:47stuff.
  32. 1:48But, we have a much better way and
  33. 1:50that's the exciting part. So, in this
  34. 1:53case, I have the PDF, I have the
  35. 1:56templates that I want to extract. And
  36. 1:58now, I'm going to use the AI extract
  37. 2:02function that we have here.
  38. 2:05This will open up a model for me, which
  39. 2:09allows me to select, all right, which of
  40. 2:12this data you want to extract. So, let's
  41. 2:15select all of them.
  42. 2:18Um I will just uncheck this last one.
  43. 2:20This If you remember, this is just a
  44. 2:22sample one that we created
  45. 2:25uh just to test this. And then, of
  46. 2:28course, check both correlation and
  47. 2:32regression table.
  48. 2:34And here, down here, we can select the
  49. 2:38large language model that will be used
  50. 2:41in the process of extraction.
  51. 2:44One thing to note is how many AI credits
  52. 2:47will be used in this. This is the most
  53. 2:50expensive part of AI using across all of
  54. 2:54HubMed data. Because, let me explain
  55. 2:57what's happening after I uh hit
  56. 2:59extraction because it takes like two or
  57. 3:02three minutes to extract one paper.
  58. 3:05So, what is happening now is that HubMed
  59. 3:08goes through multiple steps to capture
  60. 3:11the data. First, it will convert your
  61. 3:14PDF and all of its data into a readable
  62. 3:18format. Into a format that is not
  63. 3:21picture, that is not, you know, um some
  64. 3:25weird numbers, into clear
  65. 3:28uh extractable text.
  66. 3:31Uh
  67. 3:32then, it will send it through a
  68. 3:34multi-agent
  69. 3:36extraction process through multiple
  70. 3:39rounds of calls to large language
  71. 3:41models. And this is something that we
  72. 3:44have been working on for a few months.
  73. 3:48And I I think this extraction
  74. 3:52uh is going to be a standardized
  75. 3:56in the future. We will definitely keep
  76. 3:59improving our process, but by our count
  77. 4:02and with our test uh
  78. 4:05the the version that we have on Hub Meta
  79. 4:07working right now as of May 2026
  80. 4:11is pretty accurate. It's It's really
  81. 4:13good in capturing the full correlation
  82. 4:16table, the actual data from the
  83. 4:19templates, uh defining all the variables
  84. 4:23and measurements. So, it's doing a very
  85. 4:25good job uh overall.
  86. 4:29So, how long does it take? It usually
  87. 4:32takes around anywhere between 2 to 5
  88. 4:36minutes per paper to to capture all of
  89. 4:39that. And that is just how long it takes
  90. 4:43the AI models to come back with the
  91. 4:46response after all of these things that
  92. 4:48we are asking them. So, we have to be
  93. 4:51patient here, but let's review the
  94. 4:55results after one study is captured. If
  95. 4:58you were happy with the extraction, we
  96. 5:00will move on to the final stage in prep,
  97. 5:03which is selecting multiple papers and
  98. 5:07then extracting their data with AI
  99. 5:11uh all at the same time.
  100. 5:14Perfect. Now AI has finished extracting
  101. 5:18the data that we want, and this is how
  102. 5:21we will see it. This is for us to review
  103. 5:24the extraction and fix any issues that
  104. 5:27we find
  105. 5:29um like uh you know, my
  106. 5:32template by template it shows uh like
  107. 5:35what has been extracted like the last
  108. 5:38name, the year, um
  109. 5:40the type of publication, you know,
  110. 5:43number of firms, number of observations,
  111. 5:46like all the things that we wanted to
  112. 5:49include here. So, that is as much as AI
  113. 5:53could actually extract. Down here, we
  114. 5:57also see a correlation table. You will
  115. 6:00see that mean and a standard deviation
  116. 6:03for eight variables were extracted.
  117. 6:07And there are a lot of other options
  118. 6:09which we need not have data in this
  119. 6:12project, but here is the actual
  120. 6:15correlation table and the actual values.
  121. 6:17We have implemented a lot of uh checks
  122. 6:21here to make sure that these extracted
  123. 6:24values are correct, like, you know, no
  124. 6:27value uh
  125. 6:28higher than one or lower than minus one.
  126. 6:31Uh we hope that this are these are good.
  127. 6:34Like, we have tested it in a few
  128. 6:35projects and so far it's been very good.
  129. 6:38Down here, the and this part is very
  130. 6:42very important, is
  131. 6:44the definition of these variables, of
  132. 6:48the measurements, according to the
  133. 6:50paper. So, you can see, for example,
  134. 6:53share of foreign outsourcing, how is
  135. 6:57that defined in the paper? So, it says
  136. 7:01uh ratio of total cost of foreign
  137. 7:04outsourcing. So, this is the exact
  138. 7:06definition as was used in the paper. So,
  139. 7:08basically, AI reads the text of the
  140. 7:11paper and says how this specific
  141. 7:15variable was defined. And then it picks
  142. 7:18that. And then, when we save this data,
  143. 7:22it will save all of these variables into
  144. 7:25our taxonomy, in the taxonomy of
  145. 7:28variables, which is something we will
  146. 7:30get to in the next video.
  147. 7:33And then, after all of these are done,
  148. 7:36we will see our regression table. Many
  149. 7:40papers also report a regression
  150. 7:42analysis. So, multiple uh versions. And
  151. 7:45then down here you can give us some
  152. 7:48feedback like how accurate was this
  153. 7:51compared to your actual paper.
  154. 7:53You can click view PDF and then see all
  155. 7:57of these data side by side to make a
  156. 7:59comparison and see if everything is
  157. 8:04accurate. And when you were happy, you
  158. 8:07will hit save results. And then these
  159. 8:11data will appear inside this page
  160. 8:15like it was before, right? So, our data
  161. 8:20is extracted with AI. This usually took
  162. 8:23anywhere between two to three hours up
  163. 8:26to a half day from from a trained
  164. 8:29research assistant to extract all of
  165. 8:31this data. Now it all happened within a
  166. 8:34few minutes. Now, if we want to do this
  167. 8:37not on just one paper, all we got to do
  168. 8:41is go to prep AI extract and then we get
  169. 8:46started. Here again we select the
  170. 8:49templates we want to extract data from.
  171. 8:52These are the the ones that we set in
  172. 8:54our code book. Then we continue. We pick
  173. 8:57our language model. By default Haiku
  174. 9:004.5. If you don't have a reason to
  175. 9:02change this, please don't. This is the
  176. 9:04one that gives us the most accurate
  177. 9:06ones. This project has already had all
  178. 9:10of its data extracted. So,
  179. 9:13you know,
  180. 9:14you know, if this hide extracted one
  181. 9:17is selected, it will not show them.
  182. 9:20Uncheck them if you want. You can select
  183. 9:23one, two, three more or, you know, use
  184. 9:26this and,
  185. 9:27you know, select more papers to extract
  186. 9:30all at one or just select all and do all
  187. 9:32of that. You can also
  188. 9:34do some different types of filters like
  189. 9:38let's just look at the ones that are
  190. 9:40missing a correlation table and missing
  191. 9:43a regression, so no data has been
  192. 9:45extracted. Let's just, for example,
  193. 9:47select all of these. And when when you
  194. 9:50select these papers, you continue, it
  195. 9:53again says how much AI credits this is
  196. 9:56going to cost you. And then you can
  197. 9:58start AI batch screening, and then it
  198. 10:01will go paper by paper and extract all
  199. 10:03of the data. Uh of course,
  200. 10:06um you might not have enough credits. If
  201. 10:08you don't, um this will warn you there
  202. 10:12and will stop after your credits run out
  203. 10:14and save the data up to that place. To
  204. 10:17add credits
  205. 10:19uh to your account, you go to this add
  206. 10:22credit button. This will take you to a
  207. 10:24payment step, and then you can add AI
  208. 10:27credits. As mentioned before, you only
  209. 10:29pay for what you use, and and only for
  210. 10:33the functions inside have meta that are
  211. 10:36using AI heavily, and we have to pay
  212. 10:39external parties uh for that. In the
  213. 10:42next video, we will talk about what we
  214. 10:45should do with data that is now
  215. 10:47extracted.
  216. 10:58>> [music]

About this transcript

This page contains the full transcript of AI-Assisted Data Extraction | HubMeta Tutorial #15 by HubMeta, generated from the public captions YouTube serves with the video. The transcript has 1,457 words across 216 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.