YouTube2Text

The Zipf Mystery — Transcript

by Vsauce · 2,813 words · 222 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Hey, Vsauce. Michael here. About 6 percent of everything you say and read and write is
  2. 0:08the
  3. 0:12"the" - is the most used word in the English language. About one out of every
  4. 0:1816 words we encounter on a daily basis is "the." The top 20 most common English
  5. 0:25words in order are "the," "of," "and," "to," "a," "in," "is," "I," "that," "it," "for," "you,"
  6. 0:32"was," "with," "on," "as," "have," "but," "be," "they." That's a fun fact. A piece of trivia but it's
  7. 0:39also more. You see, whether the most commonly used words are ranked across an
  8. 0:44entire language, or in just one book or article, almost every time a bizarre
  9. 0:51pattern emerges. The second most used word will appear about half as often as
  10. 0:57the most used. The third one third as often. The fourth one fourth as often. The
  11. 1:04fifth one fifth as often. The sixth one sixth as often, and so on all the way down.
  12. 1:10Seriously. For some reason, the amount of times a word is used is just
  13. 1:16proportional to one over its rank. Word frequency and ranking on a log log graph
  14. 1:23follow a nice straight line. A power-law. This phenomenon is called Zipf's Law and
  15. 1:30it doesn't only apply to English. It also applies to other languages, like, well,
  16. 1:38all of them.
  17. 1:39Even ancient languages we haven't been able to translate yet.
  18. 1:43And here's the thing. We have no idea why. It's surprising that something as
  19. 1:50complex as reality should be conveyed by something as creative as language in
  20. 1:56such a predictable way. How predictable? Well, watch this. According to WordCount.org,
  21. 2:03which ranks words as found in the British National Corpus, "sauce" is the
  22. 2:085,555th most common English word. Now, here is a list of how many times
  23. 2:15every word on Wikipedia and in the entire Gutenberg Corpus of tens of
  24. 2:21thousands of public domain books shows up. The most used word, 'the,' shows up about
  25. 2:27181 million times. Knowing these two things, we can estimate that the word
  26. 2:34"sauce" should appear about thirty thousand times on Wikipedia and
  27. 2:39Gutenberg combined. And it pretty much does.
  28. 2:45What gives? The world is chaotic. Things are distributed in myriad of ways, not just
  29. 2:51power laws. And language is personal,
  30. 2:54intentional, idiosyncratic. What about the world and ourselves could cause such
  31. 3:00complex activities and behaviors to follow such a basic rule? We literally
  32. 3:08don't know. More than a century of research has yet to close the case.
  33. 3:13Moreover, Zipf's law doesn't just mysteriously describe word use. It's
  34. 3:20also found in city populations, solar flare intensities, protein sequences and
  35. 3:25immune receptors, the amount of traffic websites get, earthquake magnitudes, the
  36. 3:30number of times academic papers are cited, last names, the firing patterns of
  37. 3:34neural networks, ingredients used in cookbooks, the number of phone calls
  38. 3:38people received, the diameter of Moon craters, the number of people that die
  39. 3:42in wars, the popularity of opening chess moves, even the rate at which we forget.
  40. 3:47There are plenty of theories about why language is 'zipf-y,' but no firm conclusions
  41. 3:54and this video doesn't contain a definite explanation either. Sorry, I know
  42. 4:00that's a bummer, since we appear to like knowing more than mystery. But that said,
  43. 4:06we also ask more than we answer. So let's dive into Zipf's ramifications, some
  44. 4:13related patterns, some possible explanations and the depth of the
  45. 4:18mystery itself. Zipf's law was popularized by George Zipf,
  46. 4:23a linguist at Harvard University. It is a discrete form of the continuous Pareto
  47. 4:29distribution from which we get the Pareto Principle. Because so many
  48. 4:35real-world processes behave this way, the Pareto Principle tells us that, as a rule
  49. 4:40of thumb, it's worth assuming that 20% of the causes are responsible for 80% of
  50. 4:47the outcome,
  51. 4:49like in language, where the most frequently used 18 percent of words
  52. 4:54account for over 80% of word occurrences. In 1896, Vilfredo Pareto showed that
  53. 5:02approximately 80% of the land in Italy was owned by just twenty percent of the
  54. 5:08population. It is said that he later noticed in his garden 20 percent of his
  55. 5:13pea pods contained eighty percent of the peas. He and other researchers looked at
  56. 5:20other datasets and found that this 80-20 imbalance comes up a lot in the world.
  57. 5:26The richest 20% of humans have 82.7% of the world's income. In the US, 20% of
  58. 5:35patients use eighty percent of health care resources. In 2002, Microsoft
  59. 5:41reported that 80% of the errors and crashes in Windows and Office are caused
  60. 5:46by 20% of the bugs detected. A common rule of thumb in the business world
  61. 5:51states that 20% of your customers are responsible for 80% of your profits and
  62. 5:57eighty percent of the complaints you receive will come from 20% of your
  63. 6:02customers. A book titled "The 80/20 Principle" even says that in a home or
  64. 6:08office,
  65. 6:0920% of the carpet receives 80 percent of the wear. Oh, and as Woody Allen famously
  66. 6:16said, "eighty percent of success is just showing up." The Pareto Principle is
  67. 6:24everywhere, which is good.
  68. 6:27By focusing on just 20 percent of what's wrong, you can often expect to solve
  69. 6:32eighty percent of the problems. A variety of different unrelated factors cause
  70. 6:38this to be true from case to case, but if we can get to the bottom of what causes
  71. 6:43some of them,
  72. 6:44maybe we'll find that one or more of those mechanisms is responsible for
  73. 6:49Zipf's law in language. George Zipf himself thought languages' interesting rank
  74. 6:55frequency distribution was a consequence of the Principle of Least Effort. The
  75. 7:00tendency for life and things to follow the path of least resistance. Zipf believed
  76. 7:07it drove much of human behavior and hypothesized that as language developed
  77. 7:12in our species, speakers naturally preferred drawing from as few words as
  78. 7:18possible to get their thoughts out there. It was easier. But in order to understand
  79. 7:23what was being said,
  80. 7:25listeners preferred larger vocabularies that gave more specificity, so that they
  81. 7:30had to do less work. The compromise between listening and speaking, Zipf felt,
  82. 7:36led to the current state of language. A few words are used often and many many
  83. 7:43many words are used rarely.
  84. 7:47Recent papers have suggested that having a few short, often used, predictable words
  85. 7:52helps dissipate information load density on listeners, spacing out important vocab
  86. 7:57so that the information rate is more constant. This makes sense and much has
  87. 8:03been learned by applying the least effort principle to other behaviors, but
  88. 8:07later researchers argued that for language, the explanation was even more
  89. 8:13simple. Just a few years after Zipf's seminal paper, Benoit Mandelbrot showed
  90. 8:19that there may be nothing mysterious about Zipf's law at all, because even if you
  91. 8:24just randomly type on a keyboard you will produce words distributed according
  92. 8:29to Zipf's law. It's a pretty cool point and this is why it happens. There are
  93. 8:35exponentially more different long words than short words. For instance, the English
  94. 8:41alphabet can be used to make 26 one letter words, but 26 squared 2 letter
  95. 8:48words. Also, in random typing, whenever the space bar is pressed a word terminates.
  96. 8:55Since there's always a certain chance that the space bar will be pressed, longer
  97. 9:00stretches of time before it happens
  98. 9:03are exponentially less likely than shorter ones. The combination of these
  99. 9:08exponentials is pretty 'Zipf-y.' For example, if all 26 letters and the
  100. 9:14spacebar are equally likely to be typed, after a letter is typed and a word has
  101. 9:20begun, the probability that the next input will be a space, thus creating a
  102. 9:25one letter word, is just one in 27. And sure enough, if you randomly generate
  103. 9:31characters or hire a proverbial typing monkey, about one out of every 27 or 3.7
  104. 9:39percent of the stuff between spaces, will be single letters. Two letter words
  105. 9:44appear when after beginning a word any character but the space bar is hit - a 26
  106. 9:50in 27 chance and then the space bar.
  107. 9:54A three-letter word is the probability of a letter, another letter and then a
  108. 9:59space. If we divide by the number of unique words of each length there can be,
  109. 10:04we get the frequency of occurrence expected for any particular word given
  110. 10:08its length. For example, the letter V will make up about 0.142 percent of
  111. 10:14random typing. The word "Vsauce" 0.0000000993 percent. Longer words are
  112. 10:24less likely, but watch this. Let's spread these frequencies out according to the
  113. 10:29ranks they'd take up on a most often used list. There are 26 possible one
  114. 10:35letter words, so each of the top 26 ranked words are expected to occur
  115. 10:40about this often. The next 676 ranks will be taken up
  116. 10:45by two letter words that show up about this often. If we extend each frequency
  117. 10:50according to how many members it has, we get Zipf. Subsequent researchers have
  118. 10:56detailed how changing up the initial conditions can smooth the steps out. Our
  119. 11:02mysterious distribution has been created out of nothing but the inevitabilities
  120. 11:08of math.
  121. 11:09So maybe there is no mystery. Maybe words are just the result of humans randomly
  122. 11:16segmenting the observable world and the mental world into labels and Zipf's law
  123. 11:21describes what naturally happens when you do that. Case closed. and as always
  124. 11:27And as always,
  125. 11:28thanks for... wait a minute!
  126. 11:31Actual language is very different from random typing. Communication is
  127. 11:36deterministic to a certain extent. Utterances and topics arrive based on
  128. 11:41what was said before. And the vocabulary we have to work with certainly isn't the
  129. 11:46result of purely random naming. For example, the monkey typing model can't
  130. 11:51explain why even the names of the elements, the planets and the days of the
  131. 11:56week are used in language according to Zipf's law. Sets like these are constrained
  132. 12:02by the natural world and they're not the result of us randomly segmenting the
  133. 12:06world into labels. Furthermore, when given a list of novel words, words they've
  134. 12:12never heard or used before, like when prompted to write a story about alien
  135. 12:16creatures with strange names, people will naturally tend to use the name of one
  136. 12:21alien twice as often as another, three times as often as another... Zipf's law appears to
  137. 12:29be built into our brains. Perhaps there is something about the way thoughts and
  138. 12:35topics of discussion ebb and flow that contributes to Zipf's law.
  139. 12:40Another way 'Zipf-ian' distributions occur is via processes that change
  140. 12:44according to how they've previously operated. These are called preferential
  141. 12:49attachment processes. They occur when something - money, views,
  142. 12:55attention, variation, friends, jobs, anything really is given out according
  143. 12:59to how much is already possessed. To go back to the carpet example, if most
  144. 13:05people walk from the living room to the kitchen across a certain path, furniture
  145. 13:11will be placed elsewhere, making that path even more popular. The more views
  146. 13:17a video or image or post has, the more likely it is to get recommended
  147. 13:22automatically or make the news for having so many views, both of which give
  148. 13:28it more views.
  149. 13:29It's like a snowball rolling down a snowy hill. The more snow it accumulates, the
  150. 13:34bigger its surface area becomes for collecting more and the faster it grows.
  151. 13:38There doesn't have to be a deliberate choice driving a preferential attachment
  152. 13:43process. It can happen naturally. Try this. Take a bunch of paper clips and grab any
  153. 13:50two at random.
  154. 13:52Link them together and then throw them back in the pile. Now, repeat over and
  155. 13:56over again. If you grab paper clips that are already part of a chain, link 'em anyway.
  156. 14:02More often than not after a while you will have a distribution that looks
  157. 14:06'Zipf-ian.' A small number of chains contain a disproportionate amount of the
  158. 14:11total paperclip count. This is simply because the longer a chain gets, the
  159. 14:16greater proportion of the whole it contains, which gives it a better chance
  160. 14:20of being picked up in the future and consequently made even longer. The rich
  161. 14:26get richer, the big get bigger, the popular get popular-er. It's just math.
  162. 14:33Perhaps languages' Zipf mystery is, if not caused by it, at least strengthened by
  163. 14:39preferential attachment. Once a word is used, it's more likely to be used again soon.
  164. 14:45Critical points may play a role as well.
  165. 14:49Writing and conversation often stick to a topic until a critical point is reached
  166. 14:54and the subject is changed and the vocabulary shifts. Processes like these
  167. 14:59are known to result in power laws. So, in the end, it seems tenable that all these
  168. 15:05mechanisms might collude to make Zipf's law the most natural way for language to
  169. 15:11be. Perhaps some of our vocabulary and grammar was developed randomly, according
  170. 15:16to Mandelbrot's theory. And the natural way conversation and discussion follow
  171. 15:22preferential attachment and criticality, coupled with the principle of least
  172. 15:26effort when speaking and listening are all responsible for the relationship
  173. 15:31between word rank and frequency.
  174. 15:35It's a shame that the answer isn't simpler, but it's fascinating because of
  175. 15:40the consequences it has on what communication is made of. Roughly
  176. 15:45speaking, and this is mind blowing, nearly half of any book, conversation or article
  177. 15:52will be nothing but the same 50 to 100 words. And nearly the other half will be
  178. 15:58words that appear in that selection only once. That's not so surprising when you
  179. 16:04consider the fact that one word accounts for 6 percent of what we say. The top 25
  180. 16:11most used words make up about a third of everything we say and the top 100 about
  181. 16:18half. Seriously. I mean, whether it's all the words in "Wet Hot American Summer," or all
  182. 16:25the words in Plato's "Complete Works" or in the complete works of Edgar Allan Poe
  183. 16:30or the Bible itself, only about 100 words are used for nearly half of everything
  184. 16:37written or said. In Alice's Adventures in Wonderland 44% and in Tom Sawyer 49.8%
  185. 16:48of the unique words used appear only once in the book. A word that is used
  186. 16:55only once in a given selection of words is called a 'hapax legomenon.'
  187. 17:01Hapax legomena are vitally important to understanding languages. If a word has
  188. 17:06only been found once in the entire known collection of an ancient language, it can
  189. 17:11be very difficult to figure out what it means. Now, there is no corpus of
  190. 17:17everything ever said or written in English, but there are very very large
  191. 17:23collections and it's fun to find hapax legomena in them. For instance, and this
  192. 17:29probably won't be the case after I mention it, but the word "quizzaciously"
  193. 17:34is in the Oxford English Dictionary, but appears nowhere on Wikipedia or in the
  194. 17:41Gutenberg corpus or in the British National Corpus or the American National
  195. 17:46Corpus, but it does appear when searched in just one result on Google. Fittingly, in a
  196. 17:53book titled "ElderSpeak" that lists it as a 'rare word.' Quizzaciously, by the way,
  197. 18:01means "in a mocking manner," as in "The paradist rattled off quizzaciously,
  198. 18:07'Hey, Vsauce. Michael here. But who is Michael and how much does here
  199. 18:12weigh?'" It's a little sad that quizzaciously has been used so infrequently. It's a
  200. 18:20fun word, but that's the way things go in a 'Zipf-ian' system. Some things get all the
  201. 18:26love, some get little. Most of what you experience on a day-to-day basis is
  202. 18:33forgotten, forgettable. The Dictionary of Obscure Sorrows, as it often does, has a
  203. 18:39word for this - Olēka - the awareness of how few days are memorable.
  204. 18:45I've been alive for almost 11,000 days but I couldn't tell you something about
  205. 18:52each one of them. I mean, not even close.
  206. 18:54Most of what we do and see and think and say and hear and feel is forgotten
  207. 19:00at a rate quite similar to Zipf's law, which makes sense. If a number of factors
  208. 19:07naturally selected for thinking and talking about the world with tools in
  209. 19:12a 'Zipf-ian' way, it makes sense we'd remember it that way too. Some things
  210. 19:17really well, most things hardly at all. But it bums me out sometimes because it
  211. 19:24means that so much is forgotten, even things that at the time you thought you
  212. 19:29could never forget. My locker number -
  213. 19:33senior year - its combination, the jokes I liked when I saw a comedian on stage,
  214. 19:38the names of people I saw every day 10 years ago. So many memories are gone. When
  215. 19:45I look at all the books I've read and realize that I can't remember every
  216. 19:48detail from them, it's a little disappointing. I mean, why even bother if
  217. 19:53the Pareto Principle dictates that my 'Zipf-ian' mind will consciously remember
  218. 19:58pretty much only the titles and a few basic reactions years later
  219. 20:04Ralph Waldo Emerson makes me feel better. He once said, "I cannot remember the books
  220. 20:09I've read any more than the meals I have eaten. Even so, they have made me."
  221. 20:17And as always,
  222. 20:19thanks for watching.

About this transcript

This page contains the full transcript of The Zipf Mystery by Vsauce, generated from the public captions YouTube serves with the video. The transcript has 2,813 words across 222 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.