YouTube2Text

[CS61C FA20] Lecture 06.6 - Floating Point: Other Floating Point Representations — Transcript

by CS 61C Departmental · 1,853 words · 266 segments · language en · Watch on YouTube

Full transcript

  1. 0:00PROFESSOR: And welcome back.
  2. 0:01In our final video in this module,
  3. 0:04let's actually talk about other floating point representations.
  4. 0:07We like the single precision 32-bit representation,
  5. 0:09but are there others out there?
  6. 0:11Yes.
  7. 0:12What if you need more?
  8. 0:13What if I need more precision to get more accuracy, hopefully?
  9. 0:19Well, let's just double it.
  10. 0:21I've got 64 bits.
  11. 0:22Many machines now are 64-bit wide architectures.
  12. 0:26Let's actually think about having a 64-bit floating point
  13. 0:28number.
  14. 0:29By the way, this works on 32-bit machines.
  15. 0:32It just one by one sends the top 32 and the second 32.
  16. 0:35So it'll work in smaller machines.
  17. 0:37This can work on a 16-bit machine.
  18. 0:38It can work on 1-bit machine.
  19. 0:40Just send all 64 bits one by one.
  20. 0:41So it can work.
  21. 0:43We call this "binary 64."
  22. 0:44The other one we called "FP32" or also "binary 32."
  23. 0:49This we call "binary 64."
  24. 0:51This is the next multiple of the word size.
  25. 0:53And so now look at the size of this, 11 bits of exponent,
  26. 0:5752 bits of significand called a "double."
  27. 1:00You've seen floats and doubles.
  28. 1:02By the way, doubles are great.
  29. 1:03And so I encourage you, if you ever
  30. 1:05use any computational science, anything
  31. 1:07where you need a floating point value, use doubles.
  32. 1:09If you're just doing small ones, you're
  33. 1:10doing some change in fraction stuff, floats are fine.
  34. 1:13But if you're actually doing some calculation where
  35. 1:15you need big numbers or small numbers and much more accuracy,
  36. 1:18you're going to use a double because that precision can
  37. 1:22buy you that accuracy.
  38. 1:24How big and small can you get to?
  39. 1:25Well, 2 times 10 to the minus 308 is pretty small,
  40. 1:30times 2 to-- and all the way up to 2 to the 10 to the--
  41. 1:322 times 10 to the 308.
  42. 1:34So that's pretty big.
  43. 1:36Those are big numbers.
  44. 1:37And again, that significand is really the key here.
  45. 1:41Are there bigger than that?
  46. 1:42Yes, there are, what is called "quad precision," 128 bits,
  47. 1:47otherwise known as "binary128."
  48. 1:49Again, unbelievable range and precision, 15 exponent bits
  49. 1:54and 120--
  50. 1:56112 significand bits.
  51. 1:58Do we have oct-precision for people who need even more?
  52. 2:01Yes.
  53. 2:01That's called "binary256."
  54. 2:03What if you actually need less?
  55. 2:05Does binary16 exist?
  56. 2:07Yes, it does.
  57. 2:08That's half precision.
  58. 2:09That's using what as known are short or 2 bytes to store that,
  59. 2:12and that's 1, 5 exponent, and 10 significand.
  60. 2:17There's another comparison called "bfloat16."
  61. 2:22And B stands for "brain."
  62. 2:23And the idea is, what if you kept the exponent width
  63. 2:27the same?
  64. 2:29Okay, let's take a look here.
  65. 2:31So this is my 32 bits.
  66. 2:33There's my 1.
  67. 2:34There's my 8.
  68. 2:35There's my 23 for FP32.
  69. 2:38If I said, you know what?
  70. 2:39I only have 2 bytes not 4 bytes to store a float, then, well,
  71. 2:42let's kind of keep the same ratio.
  72. 2:44I've got a 1 of 5 and a 10 in here.
  73. 2:47But what if you wanted to keep the same big-to-small range?
  74. 2:52Well, wouldn't it be nice if I kept the same 8?
  75. 2:55So actually, notice this width is exactly the same.
  76. 2:59"Bfloat16" stands for "brain floating point" format.
  77. 3:04And this is the same 8.
  78. 3:05Notice it's the same 8.
  79. 3:06I just sacrifice precision.
  80. 3:09So this is a lot smaller.
  81. 3:10This is 1 out of 16.
  82. 3:11It's 1, 8, and 7.
  83. 3:12Has to be, obviously, right?
  84. 3:14So that is a little smaller.
  85. 3:15And that's used for faster machine learning.
  86. 3:18And in fact, there's now an entire soup
  87. 3:21of people interested in doing machine learning,
  88. 3:23deep neural nets.
  89. 3:25There's now a TF32, which also has
  90. 3:288 but has 10 significand bits, and it's 18 bits.
  91. 3:31So it's not a multiple of 8.
  92. 3:34That's a little weird.
  93. 3:35And in fact, domain accelerators have even more support,
  94. 3:40or I would say a spotty support of [? all them ?] all.
  95. 3:43And I'm looking at this table.
  96. 3:44And the key thing I'm seeing is this Nvidia
  97. 3:46is the only one who supports all of them.
  98. 3:49Notice that most of them support different sizes like, well,
  99. 3:55to optimize, I'll only work with the bfloats,
  100. 3:58or I'll only work with int8s.
  101. 4:00And by the way, there's a int4 as well.
  102. 4:02This means an integer.
  103. 4:03By the way, this means a nibble.
  104. 4:04This means a byte.
  105. 4:06This means a 2 bytes of that.
  106. 4:08And then the floating point space,
  107. 4:10these are not all supported by all
  108. 4:12of these domain accelerators.
  109. 4:13So think about that when you choose a particular domain
  110. 4:17accelerator that you might pay for on the cloud
  111. 4:18or try to buy yourself or build yourself.
  112. 4:21So those are not full support for all the different kinds
  113. 4:24of numbers that we're going to use.
  114. 4:26Here's a really interesting idea.
  115. 4:29This is Dr. John Gustafson.
  116. 4:31He said, remember what we didn't like about fixed point?
  117. 4:35You kind of locked in where the binary point was.
  118. 4:38Let's be clever.
  119. 4:39Let's have two numbers that in which
  120. 4:42you saw that all those period formats locked
  121. 4:45in what the sine bit was, the exponent bit was,
  122. 4:50and the significand bit was.
  123. 4:51You saw them, right, all the different things.
  124. 4:53So he stepped back and said, well,
  125. 4:54why don't we abstract it even more?
  126. 4:56Why don't we allow you to be able to vary
  127. 5:01the bit width of the exponent and significand.
  128. 5:06You see that connection?
  129. 5:07And so that was his idea.
  130. 5:08So what if you made that variable so it's almost
  131. 5:10like a third variable?
  132. 5:12This tells you how many bits for an each or maybe just
  133. 5:15where that line is.
  134. 5:16And then these two is what the exponent significand are.
  135. 5:20He's also going to add this very clever idea called a "u-bit"
  136. 5:24that says, you know, remember before when we were on this
  137. 5:26floating point, and you did some calculation,
  138. 5:27and you had some computation that ended up being not exactly
  139. 5:30representable?
  140. 5:31And we had to say, we're talking about rounding.
  141. 5:32Does it round down?
  142. 5:33Does it round up?
  143. 5:34What happens?
  144. 5:34Well, when it would have been in the middle,
  145. 5:37he stores the bit to say, you know what?
  146. 5:39We did some rounding.
  147. 5:40It wasn't exact.
  148. 5:42And this u-bit tells you whether the number in my calculation
  149. 5:46was exact or not.
  150. 5:47So if it actually can end up 2 to the something,
  151. 5:49that's certainly representable perfectly because 2
  152. 5:52to the something we can do really well
  153. 5:53if it's in our range.
  154. 5:54And what if I double that?
  155. 5:55Well, that's also in the range because we can
  156. 5:57do powers of 2 really easily.
  157. 5:58Well, then that would say the u-bit
  158. 6:00is off because I'm exactly snapped to a value.
  159. 6:03But if ever the calculation was in the middle, then
  160. 6:05I'm going to eventually have to snap
  161. 6:07to a value of a u-bit, whatever the unum is, the closest unum.
  162. 6:09They call them "unums"--
  163. 6:10the closest unum is, I would snap to it.
  164. 6:12And I'll set the u-bit to 1 saying, you know what?
  165. 6:14I had to do some rounding.
  166. 6:15So it remembers whether you did rounding
  167. 6:17or not and whether you had to.
  168. 6:18If I didn't have to, your bit is off.
  169. 6:20If I had to, your bit is on.
  170. 6:21That's very clever.
  171. 6:22So this, his quote, his boastful quote, I'll say--
  172. 6:26I'll even go on record saying that-- said, "It promises
  173. 6:28to be to floating point what floating point was
  174. 6:31to fixed point," which is kind of true in a way.
  175. 6:33You think about we locked in where the binary point was.
  176. 6:36Let's make it more general with this knob over here
  177. 6:39telling you where that point should float.
  178. 6:41And he's adding another knob that
  179. 6:43says how many bits in each of the exponent and significand
  180. 6:46I'm going to give.
  181. 6:47And so that is actually a bit more interesting.
  182. 6:49If I needed a lot of precision, then give a lot
  183. 6:51more bits to the significand.
  184. 6:53A lot more range, I'll give a lot more bits to the exponent.
  185. 6:56That makes sense?
  186. 6:56So it's very clever to think about how you do this to able
  187. 6:59to have even more range.
  188. 7:00So unum hasn't been formally adopted
  189. 7:03by every piece of hardware, but it's getting more adoption.
  190. 7:06And I think it's effectually a very clever solution, so claims
  191. 7:08to save power as well.
  192. 7:09I don't know whether I can support that claim,
  193. 7:12but that's another claim they have.
  194. 7:14In summary, last slide, floating point
  195. 7:19lets us represent numbers containing integer
  196. 7:21and floating point parts.
  197. 7:22We can do big numbers.
  198. 7:23We can do small numbers.
  199. 7:25Amazing.
  200. 7:26IEEE 754 floating point standard earned one of my colleagues
  201. 7:31a Turing Award, the Nobel Prize in computing for that work.
  202. 7:34Every computer since 1997-- before then
  203. 7:37was having a different way that did it.
  204. 7:40HP was different from Apple which was different from Intel.
  205. 7:42It was crazy.
  206. 7:43It was chaos.
  207. 7:45This floating point standardized everybody.
  208. 7:47And now we're all using the floating point standard.
  209. 7:49By the way, that has revisions.
  210. 7:50There were revisions to this where
  211. 7:51they tried to correct some things and add things.
  212. 7:53There's always a revision to that.
  213. 7:54You saw the revisions to C. There's
  214. 7:55revisions to the IEEE floating point standard as well.
  215. 7:58And you should say IEEE 74 underscore
  216. 8:00the year, which is kind of nice, in four digits,
  217. 8:02not 79 with two digits.
  218. 8:04This had four digits for that.
  219. 8:05So it'll say underscore 2000 and the year or something
  220. 8:07like that.
  221. 8:07And I think there's a 2019 out there recently as well.
  222. 8:09So summary, this is it.
  223. 8:12You've seen examples.
  224. 8:13You've seen pictures.
  225. 8:16You've seen ways you can play with an interactive,
  226. 8:20touch the bits and play with it and explore that.
  227. 8:22I encourage you strongly to do that.
  228. 8:24S, exponent, significand, everything,
  229. 8:26it follows this format.
  230. 8:27Every other format basically is just
  231. 8:30changing the widths of the exponent
  232. 8:31and significand to do different things.
  233. 8:34But FP32 is the default. When you say "float,"
  234. 8:36that's what you get.
  235. 8:38Remember, exponent is telling the significand how much,
  236. 8:41what power to to count by.
  237. 8:43Are you counting by quarters or halves or ones or twos?
  238. 8:46So you're always counting by power or minus 149's.
  239. 8:49That's what the smallest one.
  240. 8:50So the exponent tells how much to count by.
  241. 8:52Then the significand is just counting by that much.
  242. 8:54That says, how wide are the ruler values?
  243. 8:57So the exponent says, tell me how wide my ruler lines are?
  244. 9:01And then the significand counts by those ruler lines.
  245. 9:04It's always an equal.
  246. 9:05So it's [ROLLS TONGUE],, and then bup, bup, bup, bup,
  247. 9:07and then bup, bup, bup, like that,
  248. 9:09if you [INAUDIBLE] this with the sound.
  249. 9:11And this benefit that you can store NaNs.
  250. 9:14You can store plus or minus infinity.
  251. 9:15You can start debugging information in your NaNs.
  252. 9:17We love all this.
  253. 9:19I'm going to take you out with one of the greatest
  254. 9:23'70s soul R&B songs.
  255. 9:27It's called "Float On" by The Floaters.
  256. 9:28Just boogie along to it.
  257. 9:30[MUSIC - THE FLOATERS, "FLOAT ON"]
  258. 9:31(SINGING) And float on.
  259. 9:32Yeah.
  260. 9:35Yeah.
  261. 9:36PROFESSOR: I'm gonna dance.
  262. 9:38(SINGING) You better float on.
  263. 9:40PROFESSOR: And we'll see you the next module.
  264. 9:42(SINGING) Float on.
  265. 9:45Ah.
  266. 9:49PROFESSOR: Take care.

About this transcript

This page contains the full transcript of [CS61C FA20] Lecture 06.6 - Floating Point: Other Floating Point Representations by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 1,853 words across 266 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.