The Taxonomy Wizard: Mapping Measurements to Constructs | HubMeta Tutorial #17 — Transcript
Full transcript
- 0:15All right, let's talk about our taxonomy
- 0:18wizard. So, what happens here is that
- 0:22the first stage in creating an
- 0:26analyzable set of constructs
- 0:29in your project is that you want to
- 0:33define some seed constructs.
- 0:36What seed constructs are the variables,
- 0:40the constructs that you know you
- 0:42actually have to analyze in your
- 0:45project. Usually, this comes from
- 0:48theory. This comes from you. You are the
- 0:50expert in your field. You're the one
- 0:52that have defined this project. So, you
- 0:54know better
- 0:55what are the constructs that you have
- 1:00absolutely want to measure here.
- 1:03Usually, these are the things that would
- 1:04go into your final model to to do the
- 1:08analysis. Of course, your main
- 1:10independent variable or variables,
- 1:12dependent variable and variables, all
- 1:15the control variables are the things
- 1:17that you define here. One of the ways to
- 1:20do this is again to look at some anchor
- 1:23studies and see how they have defined
- 1:26their variables. For example, in this
- 1:29case, my anchor study has things such as
- 1:34multinationality, financial performance,
- 1:37and they clearly define what
- 1:39measurements were used in all of that.
- 1:42Like in my case, my anchor study has
- 1:45things such as financial performance,
- 1:48multinationality,
- 1:50R&D intensity, and firm size, firm age,
- 1:53and things such as that. And I will make
- 1:57sure that I will at the very least have
- 1:59these as my seed constructs. But in
- 2:03addition to that, maybe there are other
- 2:06things that I can measure in this study.
- 2:10Whatever the case may be, I can add seed
- 2:13constructs one by one. I can just say
- 2:16firm multi- multinationality
- 2:20and then add a definition here. I can
- 2:24then create the seed. That's one way,
- 2:26manually add them one by one. The other
- 2:30way is we have defined a batch AIC
- 2:34option.
- 2:36What this one does is it will ask you to
- 2:39give it a seed paper. Usually a one or
- 2:43two core defining papers in your field.
- 2:47You put it here. You give it some
- 2:48context of what your construct is and
- 2:51then this can use that to generate a set
- 2:56of seed articles
- 2:59nationality and performance. Like I
- 3:02added my core paper here. I would say
- 3:05what is your my target size? Let's just
- 3:07say in this case I don't want
- 3:11a lot. I will just do 10 of them.
- 3:15Control variables and moderators too and
- 3:18then generate. This takes around 20 to
- 3:2160 seconds. It will read the PDF. It
- 3:25will process that. It has some prompts
- 3:28to suggest the constructs that are good
- 3:33and can be analyzed in this project. And
- 3:36this is what it came up with. So, if you
- 3:39remember, these were actually the ones
- 3:42that were measured in in that paper and
- 3:46we did a good job in suggesting what
- 3:48needs to be captured. It says is this an
- 3:51outcome variable a control or
- 3:54independent I can delete edit or add to
- 3:58this and then or save these 14 to my
- 4:02seat or I can regenerate with more
- 4:05things. So that our first step. We are
- 4:08starting with some seeds. I have already
- 4:11defined my seeds. I'm happy with their
- 4:13definition. So my next step is that
- 4:16okay, I know the parents. I know the
- 4:19parent constructs. Now my job is to
- 4:22assign the measurements like the
- 4:24different ways that all of these
- 4:26constructs have been measured and put
- 4:29them under the relevant parent. So if
- 4:31you if you think about it
- 4:34this is actually a an an assignment of
- 4:39function. It is something about checking
- 4:42the similarity
- 4:43of your measurement definition with your
- 4:47constructs. We have different steps of
- 4:51doing that.
- 4:53And we have which include suggestions
- 4:58based on cosine similarity like
- 5:01similarity between the definition of the
- 5:05measurement and the construct similarity
- 5:08with according to AI embeddings like a
- 5:11little bit more involved machine
- 5:13learning method of doing that. Then for
- 5:16all the ones that were not assigned
- 5:18using this process we have this
- 5:20clustering model. Here we explain
- 5:23exactly how this process works and how
- 5:26all of these functions work. You can
- 5:28spend some time working with that. But
- 5:32they what I recommend running is this
- 5:35thing that will do all of these steps
- 5:38that I mentioned the automatically. So
- 5:42after you have defined your seeds, you
- 5:45You go to our auto tab and assign all of
- 5:48these. You don't have to do anything
- 5:49else and you just run the pipeline and
- 5:51it just automatically assigns everything
- 5:54for you. But let me just explain what is
- 5:56happening here. The in the first step,
- 5:59which is our tier one assignment, what
- 6:02the
- 6:03what have meta does is it just checks
- 6:06pure
- 6:08word-by-word similarity. If in our
- 6:10construct we have defined firm size as
- 6:13the number of employees or how big the
- 6:16firm is, it is looking for those words
- 6:20in the definition of the measurement. So
- 6:23when if it finds them, if the definition
- 6:25of the measurement is close enough in
- 6:28terms of cosine similarity, that will be
- 6:31included.
- 6:32And that's the first one. And here I'll
- 6:36set it to usually a high number like
- 6:39it's anything above a 75% is means very
- 6:43similar. So I try to be conservative
- 6:47here like even set it at 0.8 because
- 6:49these will be automatically assigned. If
- 6:52you saw that level of similarity between
- 6:54that measurement and our construct, just
- 6:57go for it. Just pick that and assign
- 6:59them.
- 7:00Then the next one that you set is what
- 7:04happens is that
- 7:06the ones that are below 0.8 similarity
- 7:10but it's still above 0.5
- 7:13means it's some sort of a a gray area.
- 7:17We don't know. So what we do for that is
- 7:20to include a large language model. So we
- 7:23will send the definition of the
- 7:25construct and the measurements to a
- 7:28large language model and say,
- 7:30"Do you think that this measurement
- 7:34belongs to this construct? Is this
- 7:36measuring the meaning of this construct?
- 7:39Is Is a good way of measuring that?" And
- 7:42it can say yes or no. And honestly, the
- 7:45way we have designed it is that LLM is
- 7:47not just saying yes or no. It's more
- 7:49than that. It says, "I'm almost around
- 7:5270% sure or 80% sure that these are the
- 7:56same." And that's where the accept score
- 7:59comes in. If So, if I set it to 80%,
- 8:03that means the LLM's response has to be
- 8:05that these two, the measurement and the
- 8:08construct, have to be more than 80%
- 8:12according to the LLM's decision close to
- 8:15each other for it to receive an accept
- 8:17decision. So, that's how it goes. And in
- 8:20the third round, if it's below this, it
- 8:24will just do
- 8:26some final round of measuring and just
- 8:30make sure if it is relevant to any of
- 8:33our remaining. So, basically, a final
- 8:35check if this is relevant or not. But,
- 8:38the main decision is made in tier one
- 8:40and tier two. So, if you go ahead and
- 8:43run the pipeline, it will
- 8:46it's not going to cost you any AI
- 8:48credits. It just takes some time for it
- 8:51to run because it's a long process. So,
- 8:54usually,
- 8:55for something that is per about 1,000
- 8:59article measurements, it takes about 10
- 9:03minutes or so. So, you can figure out
- 9:05how long it will take total.
- 9:08So, you can see if you read this, it's a
- 9:11it auto assigned a few of them. And
- 9:14then, it it auto assigned only 566
- 9:18measurements because I set it the auto
- 9:21assign limit to very high. And then, the
- 9:24rest of it that was not auto assigned,
- 9:27but was still above .5, it is being sent
- 9:31to an LLM. So, the LLM is making a
- 9:34decision on 5,000 papers. So, these are
- 9:38borderline items. We want to see where
- 9:40they belong and we will pick a parent
- 9:43for them.
- 9:44So, it this will continue going on and
- 9:47on. Because we don't have time for that,
- 9:49I will just cancel this pipeline and
- 9:52then look at a generation that I've done
- 9:55in the past. So, let's look at this one
- 9:58for example. You can see your past runs
- 10:01here. So, in this one for example, what
- 10:04the LLM did was it assigned 567
- 10:07of my 14,000 measurements. And the rest
- 10:11of them, which were in the gray area,
- 10:13were sent to an LLM.
- 10:1618,000 received a accept vote. 1,300
- 10:22received a reject vote. Some of them
- 10:25were accepted after a second pass. And
- 10:28then you can see all the rest of the all
- 10:30the other ones. So, if I accept save
- 10:34these 2,500 measurements, my taxonomy
- 10:38becomes ready. So, that means my
- 10:40measurements will be assigned to a
- 10:43parent construct. So, I am done with my
- 10:46taxonomy. Of course, even after this
- 10:49saving, I still have a residual of
- 10:5310,000
- 10:55measurements. These 10,000 measurements
- 10:57are not relevant to the seed constructs
- 11:01that I had defined in my project. So, I
- 11:04have not found a parent for them within
- 11:07those. But that doesn't mean they are
- 11:09useless. I can still find some clusters
- 11:12for them. And for that, we have this
- 11:15clusters tab. What this one does is that
- 11:18it only looks at the ones that do not
- 11:20have a parent. And here we have multiple
- 11:24different ways of generating a cluster.
- 11:27And you can read all about it in this
- 11:29note. When you run this, I just want to
- 11:31show you the final result, so you have
- 11:34an understanding of what we mean by
- 11:36that. In a nutshell, what this process
- 11:39is doing is it looks at the semantic
- 11:42proximity of this remaining residual
- 11:46measurements and will suggest groupings
- 11:50for them. So, basically it will say,
- 11:52"Okay, this pack of 100 looks similar
- 11:55enough to each other. Perhaps they are
- 11:57measuring a construct like this." Let's
- 12:00see the results and I will explain more.
- 12:02So, as you can see, it is running right
- 12:04now. It's rather long process because
- 12:07it's basically running across
- 12:10my project 10,000 measurements and find
- 12:14those clusters or groups of similarity
- 12:17in meanings and putting them all
- 12:19together. We'll give it a second. We
- 12:22will give it a second and come back to
- 12:24it when it is finished. All right. Now
- 12:27that it is finished, you will see that
- 12:29it has found different clusters like 466
- 12:35clusters to be exact and each of them is
- 12:39its own group. And it also creates this
- 12:41visualization. In this case, because my
- 12:44data is very big, it might not be as
- 12:47easy and meaningful to look at this
- 12:50visualization.
- 12:51But if it is lower, you will see like
- 12:53the ones that are closer in meaning to
- 12:55each other are closer in this map and a
- 12:58similar color. But let's take a look at
- 13:01one of these clusters of meaning. So, if
- 13:05I click on it, it will say, "Okay, all
- 13:06the measurements that are
- 13:10have been detected to be close enough to
- 13:12each other." And when I look at this, I
- 13:15see a lot of things about foreign direct
- 13:19investment
- 13:20or trade
- 13:23maybe trade barriers. They are not
- 13:26completely relevant to each other, but
- 13:29they are not that far-fetched either.
- 13:32So, another one is Tobin's Q. It's a
- 13:37measurement for performance. You can see
- 13:39that all of these are actually different
- 13:41ways of Tobin's Q as a performance
- 13:44measurement. So, maybe there is
- 13:47something here. So, 121 measurements in
- 13:50this cluster. So, the next step we do is
- 13:53to when we click on generate construct,
- 13:56it will feed all of these measurements
- 13:59in this definition to an AI and suggest
- 14:04a name and a definition
- 14:07for that cluster. So, here it says, you
- 14:10know, an obvious name which is Tobin's Q
- 14:13or firm market valuation. Here is the
- 14:15definition.
- 14:17These measurements predominantly capture
- 14:19a firm's market valuation relative to
- 14:22its book value or asset replacement
- 14:25cost. This is a very good definition, I
- 14:27think. It can even be called firm
- 14:29performance, the Tobin's Q. And then, if
- 14:33I happy with this, I will define this as
- 14:37the parent construct and assign all of
- 14:40these 121 measurements under that. So,
- 14:45if I click save node,
- 14:47I have one new construct which is parent
- 14:52to all of this. And to double-check, you
- 14:54can always come here
- 14:57and filter your constructs. It will see
- 15:00It will show all the ones that you have
- 15:03already defined. Like this was the one
- 15:05that we just created, the firm
- 15:08performance Tobin's Q.
- 15:10And that's That's what we have. It's
- 15:12still It's more refreshed to reflect the
- 15:15correct number.
- 15:16But that's it about our taxonomy. It's a
- 15:20different platform of its own. so this
- 15:22one it was a longer video compared to
- 15:25the rest of them. See you in the next
- 15:27one to talk about exporting and
- 15:29analyzing our data.
- 15:41>> [music]
About this transcript
This page contains the full transcript of The Taxonomy Wizard: Mapping Measurements to Constructs | HubMeta Tutorial #17 by HubMeta, generated from the public captions YouTube serves with the video. The transcript has 2,131 words across 315 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.