YouTube2Text

The Taxonomy Wizard: Mapping Measurements to Constructs | HubMeta Tutorial #17 — Transcript

by HubMeta · 2,131 words · 315 segments · language en · Watch on YouTube

Full transcript

  1. 0:15All right, let's talk about our taxonomy
  2. 0:18wizard. So, what happens here is that
  3. 0:22the first stage in creating an
  4. 0:26analyzable set of constructs
  5. 0:29in your project is that you want to
  6. 0:33define some seed constructs.
  7. 0:36What seed constructs are the variables,
  8. 0:40the constructs that you know you
  9. 0:42actually have to analyze in your
  10. 0:45project. Usually, this comes from
  11. 0:48theory. This comes from you. You are the
  12. 0:50expert in your field. You're the one
  13. 0:52that have defined this project. So, you
  14. 0:54know better
  15. 0:55what are the constructs that you have
  16. 1:00absolutely want to measure here.
  17. 1:03Usually, these are the things that would
  18. 1:04go into your final model to to do the
  19. 1:08analysis. Of course, your main
  20. 1:10independent variable or variables,
  21. 1:12dependent variable and variables, all
  22. 1:15the control variables are the things
  23. 1:17that you define here. One of the ways to
  24. 1:20do this is again to look at some anchor
  25. 1:23studies and see how they have defined
  26. 1:26their variables. For example, in this
  27. 1:29case, my anchor study has things such as
  28. 1:34multinationality, financial performance,
  29. 1:37and they clearly define what
  30. 1:39measurements were used in all of that.
  31. 1:42Like in my case, my anchor study has
  32. 1:45things such as financial performance,
  33. 1:48multinationality,
  34. 1:50R&D intensity, and firm size, firm age,
  35. 1:53and things such as that. And I will make
  36. 1:57sure that I will at the very least have
  37. 1:59these as my seed constructs. But in
  38. 2:03addition to that, maybe there are other
  39. 2:06things that I can measure in this study.
  40. 2:10Whatever the case may be, I can add seed
  41. 2:13constructs one by one. I can just say
  42. 2:16firm multi- multinationality
  43. 2:20and then add a definition here. I can
  44. 2:24then create the seed. That's one way,
  45. 2:26manually add them one by one. The other
  46. 2:30way is we have defined a batch AIC
  47. 2:34option.
  48. 2:36What this one does is it will ask you to
  49. 2:39give it a seed paper. Usually a one or
  50. 2:43two core defining papers in your field.
  51. 2:47You put it here. You give it some
  52. 2:48context of what your construct is and
  53. 2:51then this can use that to generate a set
  54. 2:56of seed articles
  55. 2:59nationality and performance. Like I
  56. 3:02added my core paper here. I would say
  57. 3:05what is your my target size? Let's just
  58. 3:07say in this case I don't want
  59. 3:11a lot. I will just do 10 of them.
  60. 3:15Control variables and moderators too and
  61. 3:18then generate. This takes around 20 to
  62. 3:2160 seconds. It will read the PDF. It
  63. 3:25will process that. It has some prompts
  64. 3:28to suggest the constructs that are good
  65. 3:33and can be analyzed in this project. And
  66. 3:36this is what it came up with. So, if you
  67. 3:39remember, these were actually the ones
  68. 3:42that were measured in in that paper and
  69. 3:46we did a good job in suggesting what
  70. 3:48needs to be captured. It says is this an
  71. 3:51outcome variable a control or
  72. 3:54independent I can delete edit or add to
  73. 3:58this and then or save these 14 to my
  74. 4:02seat or I can regenerate with more
  75. 4:05things. So that our first step. We are
  76. 4:08starting with some seeds. I have already
  77. 4:11defined my seeds. I'm happy with their
  78. 4:13definition. So my next step is that
  79. 4:16okay, I know the parents. I know the
  80. 4:19parent constructs. Now my job is to
  81. 4:22assign the measurements like the
  82. 4:24different ways that all of these
  83. 4:26constructs have been measured and put
  84. 4:29them under the relevant parent. So if
  85. 4:31you if you think about it
  86. 4:34this is actually a an an assignment of
  87. 4:39function. It is something about checking
  88. 4:42the similarity
  89. 4:43of your measurement definition with your
  90. 4:47constructs. We have different steps of
  91. 4:51doing that.
  92. 4:53And we have which include suggestions
  93. 4:58based on cosine similarity like
  94. 5:01similarity between the definition of the
  95. 5:05measurement and the construct similarity
  96. 5:08with according to AI embeddings like a
  97. 5:11little bit more involved machine
  98. 5:13learning method of doing that. Then for
  99. 5:16all the ones that were not assigned
  100. 5:18using this process we have this
  101. 5:20clustering model. Here we explain
  102. 5:23exactly how this process works and how
  103. 5:26all of these functions work. You can
  104. 5:28spend some time working with that. But
  105. 5:32they what I recommend running is this
  106. 5:35thing that will do all of these steps
  107. 5:38that I mentioned the automatically. So
  108. 5:42after you have defined your seeds, you
  109. 5:45You go to our auto tab and assign all of
  110. 5:48these. You don't have to do anything
  111. 5:49else and you just run the pipeline and
  112. 5:51it just automatically assigns everything
  113. 5:54for you. But let me just explain what is
  114. 5:56happening here. The in the first step,
  115. 5:59which is our tier one assignment, what
  116. 6:02the
  117. 6:03what have meta does is it just checks
  118. 6:06pure
  119. 6:08word-by-word similarity. If in our
  120. 6:10construct we have defined firm size as
  121. 6:13the number of employees or how big the
  122. 6:16firm is, it is looking for those words
  123. 6:20in the definition of the measurement. So
  124. 6:23when if it finds them, if the definition
  125. 6:25of the measurement is close enough in
  126. 6:28terms of cosine similarity, that will be
  127. 6:31included.
  128. 6:32And that's the first one. And here I'll
  129. 6:36set it to usually a high number like
  130. 6:39it's anything above a 75% is means very
  131. 6:43similar. So I try to be conservative
  132. 6:47here like even set it at 0.8 because
  133. 6:49these will be automatically assigned. If
  134. 6:52you saw that level of similarity between
  135. 6:54that measurement and our construct, just
  136. 6:57go for it. Just pick that and assign
  137. 6:59them.
  138. 7:00Then the next one that you set is what
  139. 7:04happens is that
  140. 7:06the ones that are below 0.8 similarity
  141. 7:10but it's still above 0.5
  142. 7:13means it's some sort of a a gray area.
  143. 7:17We don't know. So what we do for that is
  144. 7:20to include a large language model. So we
  145. 7:23will send the definition of the
  146. 7:25construct and the measurements to a
  147. 7:28large language model and say,
  148. 7:30"Do you think that this measurement
  149. 7:34belongs to this construct? Is this
  150. 7:36measuring the meaning of this construct?
  151. 7:39Is Is a good way of measuring that?" And
  152. 7:42it can say yes or no. And honestly, the
  153. 7:45way we have designed it is that LLM is
  154. 7:47not just saying yes or no. It's more
  155. 7:49than that. It says, "I'm almost around
  156. 7:5270% sure or 80% sure that these are the
  157. 7:56same." And that's where the accept score
  158. 7:59comes in. If So, if I set it to 80%,
  159. 8:03that means the LLM's response has to be
  160. 8:05that these two, the measurement and the
  161. 8:08construct, have to be more than 80%
  162. 8:12according to the LLM's decision close to
  163. 8:15each other for it to receive an accept
  164. 8:17decision. So, that's how it goes. And in
  165. 8:20the third round, if it's below this, it
  166. 8:24will just do
  167. 8:26some final round of measuring and just
  168. 8:30make sure if it is relevant to any of
  169. 8:33our remaining. So, basically, a final
  170. 8:35check if this is relevant or not. But,
  171. 8:38the main decision is made in tier one
  172. 8:40and tier two. So, if you go ahead and
  173. 8:43run the pipeline, it will
  174. 8:46it's not going to cost you any AI
  175. 8:48credits. It just takes some time for it
  176. 8:51to run because it's a long process. So,
  177. 8:54usually,
  178. 8:55for something that is per about 1,000
  179. 8:59article measurements, it takes about 10
  180. 9:03minutes or so. So, you can figure out
  181. 9:05how long it will take total.
  182. 9:08So, you can see if you read this, it's a
  183. 9:11it auto assigned a few of them. And
  184. 9:14then, it it auto assigned only 566
  185. 9:18measurements because I set it the auto
  186. 9:21assign limit to very high. And then, the
  187. 9:24rest of it that was not auto assigned,
  188. 9:27but was still above .5, it is being sent
  189. 9:31to an LLM. So, the LLM is making a
  190. 9:34decision on 5,000 papers. So, these are
  191. 9:38borderline items. We want to see where
  192. 9:40they belong and we will pick a parent
  193. 9:43for them.
  194. 9:44So, it this will continue going on and
  195. 9:47on. Because we don't have time for that,
  196. 9:49I will just cancel this pipeline and
  197. 9:52then look at a generation that I've done
  198. 9:55in the past. So, let's look at this one
  199. 9:58for example. You can see your past runs
  200. 10:01here. So, in this one for example, what
  201. 10:04the LLM did was it assigned 567
  202. 10:07of my 14,000 measurements. And the rest
  203. 10:11of them, which were in the gray area,
  204. 10:13were sent to an LLM.
  205. 10:1618,000 received a accept vote. 1,300
  206. 10:22received a reject vote. Some of them
  207. 10:25were accepted after a second pass. And
  208. 10:28then you can see all the rest of the all
  209. 10:30the other ones. So, if I accept save
  210. 10:34these 2,500 measurements, my taxonomy
  211. 10:38becomes ready. So, that means my
  212. 10:40measurements will be assigned to a
  213. 10:43parent construct. So, I am done with my
  214. 10:46taxonomy. Of course, even after this
  215. 10:49saving, I still have a residual of
  216. 10:5310,000
  217. 10:55measurements. These 10,000 measurements
  218. 10:57are not relevant to the seed constructs
  219. 11:01that I had defined in my project. So, I
  220. 11:04have not found a parent for them within
  221. 11:07those. But that doesn't mean they are
  222. 11:09useless. I can still find some clusters
  223. 11:12for them. And for that, we have this
  224. 11:15clusters tab. What this one does is that
  225. 11:18it only looks at the ones that do not
  226. 11:20have a parent. And here we have multiple
  227. 11:24different ways of generating a cluster.
  228. 11:27And you can read all about it in this
  229. 11:29note. When you run this, I just want to
  230. 11:31show you the final result, so you have
  231. 11:34an understanding of what we mean by
  232. 11:36that. In a nutshell, what this process
  233. 11:39is doing is it looks at the semantic
  234. 11:42proximity of this remaining residual
  235. 11:46measurements and will suggest groupings
  236. 11:50for them. So, basically it will say,
  237. 11:52"Okay, this pack of 100 looks similar
  238. 11:55enough to each other. Perhaps they are
  239. 11:57measuring a construct like this." Let's
  240. 12:00see the results and I will explain more.
  241. 12:02So, as you can see, it is running right
  242. 12:04now. It's rather long process because
  243. 12:07it's basically running across
  244. 12:10my project 10,000 measurements and find
  245. 12:14those clusters or groups of similarity
  246. 12:17in meanings and putting them all
  247. 12:19together. We'll give it a second. We
  248. 12:22will give it a second and come back to
  249. 12:24it when it is finished. All right. Now
  250. 12:27that it is finished, you will see that
  251. 12:29it has found different clusters like 466
  252. 12:35clusters to be exact and each of them is
  253. 12:39its own group. And it also creates this
  254. 12:41visualization. In this case, because my
  255. 12:44data is very big, it might not be as
  256. 12:47easy and meaningful to look at this
  257. 12:50visualization.
  258. 12:51But if it is lower, you will see like
  259. 12:53the ones that are closer in meaning to
  260. 12:55each other are closer in this map and a
  261. 12:58similar color. But let's take a look at
  262. 13:01one of these clusters of meaning. So, if
  263. 13:05I click on it, it will say, "Okay, all
  264. 13:06the measurements that are
  265. 13:10have been detected to be close enough to
  266. 13:12each other." And when I look at this, I
  267. 13:15see a lot of things about foreign direct
  268. 13:19investment
  269. 13:20or trade
  270. 13:23maybe trade barriers. They are not
  271. 13:26completely relevant to each other, but
  272. 13:29they are not that far-fetched either.
  273. 13:32So, another one is Tobin's Q. It's a
  274. 13:37measurement for performance. You can see
  275. 13:39that all of these are actually different
  276. 13:41ways of Tobin's Q as a performance
  277. 13:44measurement. So, maybe there is
  278. 13:47something here. So, 121 measurements in
  279. 13:50this cluster. So, the next step we do is
  280. 13:53to when we click on generate construct,
  281. 13:56it will feed all of these measurements
  282. 13:59in this definition to an AI and suggest
  283. 14:04a name and a definition
  284. 14:07for that cluster. So, here it says, you
  285. 14:10know, an obvious name which is Tobin's Q
  286. 14:13or firm market valuation. Here is the
  287. 14:15definition.
  288. 14:17These measurements predominantly capture
  289. 14:19a firm's market valuation relative to
  290. 14:22its book value or asset replacement
  291. 14:25cost. This is a very good definition, I
  292. 14:27think. It can even be called firm
  293. 14:29performance, the Tobin's Q. And then, if
  294. 14:33I happy with this, I will define this as
  295. 14:37the parent construct and assign all of
  296. 14:40these 121 measurements under that. So,
  297. 14:45if I click save node,
  298. 14:47I have one new construct which is parent
  299. 14:52to all of this. And to double-check, you
  300. 14:54can always come here
  301. 14:57and filter your constructs. It will see
  302. 15:00It will show all the ones that you have
  303. 15:03already defined. Like this was the one
  304. 15:05that we just created, the firm
  305. 15:08performance Tobin's Q.
  306. 15:10And that's That's what we have. It's
  307. 15:12still It's more refreshed to reflect the
  308. 15:15correct number.
  309. 15:16But that's it about our taxonomy. It's a
  310. 15:20different platform of its own. so this
  311. 15:22one it was a longer video compared to
  312. 15:25the rest of them. See you in the next
  313. 15:27one to talk about exporting and
  314. 15:29analyzing our data.
  315. 15:41>> [music]

About this transcript

This page contains the full transcript of The Taxonomy Wizard: Mapping Measurements to Constructs | HubMeta Tutorial #17 by HubMeta, generated from the public captions YouTube serves with the video. The transcript has 2,131 words across 315 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.