[CS61C FA20] Lecture 06.5 - Floating Point: Floating Point Discussion — Transcript
Full transcript
- 0:00and welcome back now let's have a
- 0:02discussion about some of the issues of
- 0:04floating point
- 0:06so first there's some fallacies we want
- 0:08to
- 0:09dispel um is add associative you know ad
- 0:14addition should be associative right
- 0:15that's you know one plus two plus three
- 0:17whether you add the one and two first or
- 0:19the two and three
- 0:20first it should be the same number it
- 0:22should be six
- 0:23well is that true for floating point add
- 0:26here's an example
- 0:26so let's make x a really small
- 0:30uh a really big negative number so minus
- 0:321.5 times 10 to the 38.
- 0:34let's make y exactly the opposite but a
- 0:37positive number
- 0:38and let z just be one and we even saw
- 0:40how to do one in a previous video
- 0:43let's first add y and z together when
- 0:46you add y and z together get this really
- 0:48big number
- 0:49and this really small number and the
- 0:51number of bits that you can play with is
- 0:53really 23
- 0:54bits total if you include the
- 0:56normalization there's like 24 total bits
- 0:58from like left most bit to the rightmost
- 1:00bit that's a one
- 1:01and the problem is you don't have enough
- 1:03bits at least in the floating point and
- 1:05a float
- 1:06to store both the one and this really
- 1:08big number over here
- 1:10so you can't do that um so when that
- 1:13then when that when this right most let
- 1:14me get my thing when this when this adds
- 1:16together that one goes away
- 1:18i can't store this one and that one the
- 1:20bigger one wins
- 1:21we lose the one well that means it's as
- 1:24if the z wasn't added at all
- 1:26and i get x plus y which is zero
- 1:30well now let's try it again seeing if
- 1:32it's associative or not let's try
- 1:35the leftmost one
- 1:38so the leftmost one says
- 1:42hear the leftmost guy
- 1:45and x and y and they cancel and then you
- 1:49have one and so
- 1:50these two are not equal so therefore in
- 1:53summary floating point add
- 1:54is not associative really important that
- 1:57you saw that
- 1:59and the reason for that is we you just
- 2:01have a fixed number of bits you can't do
- 2:03more than that if you have a really big
- 2:04number and really small stuff you can't
- 2:06keep both those guys alive you could
- 2:08imagine a third
- 2:09maybe you wanted to invent a third way
- 2:11to have big numbers and small numbers
- 2:13and have two pieces of them and
- 2:14i mean that's kind of like having two
- 2:16separate numbers anyway but you can't a
- 2:17single number can't have b both really
- 2:19big and really small
- 2:21so that's i mean this bullet says
- 2:24exactly what i said
- 2:25the number is so much bigger than one
- 2:26that it can't the representation cannot
- 2:28store both of them at once and so it
- 2:29loses it in one
- 2:30keeps it in the other because they
- 2:31cancel each other out and one is not
- 2:33equal to zero
- 2:36i wanted to give you a definition just
- 2:38so that you're using the right words as
- 2:39you're talking about floating point
- 2:40numbers
- 2:41um so precision is if this is
- 2:44any representation um it's the number of
- 2:47bits we're throwing at the problem so
- 2:49precision oh i want more precision
- 2:50throw more bits of the problem oh okay
- 2:52float only has this number of bits well
- 2:54double that or quadruple that
- 2:55you'll get more precision the more bits
- 2:57you have the closer you can get to the
- 2:59actual number you're trying to get to
- 3:01accuracy is the distance between any
- 3:04number that you have coded and the
- 3:06original
- 3:06actual number so his example high
- 3:09precision
- 3:10a lot of bits permits high accuracy but
- 3:12doesn't guarantee it
- 3:13here's an example here float pi equals
- 3:163.14
- 3:17well there's going to be lots of bits
- 3:2020 you know 32 bits associated for that
- 3:23but you're still pretty far in terms of
- 3:25accuracy from the actual value of pi
- 3:27now you could have done more i could
- 3:28have initialized it to a closer number
- 3:30but that
- 3:30typical example says i got a lot of bits
- 3:33but it isn't
- 3:34close to the actual number so that has
- 3:36high precision
- 3:37low accuracy in that case
- 3:40now let's talk about rounding so
- 3:42rounding is this complicated thing
- 3:44uh where you're doing some calculation
- 3:48and then you figure out okay well i'm
- 3:49not exactly at
- 3:50one or the other i'm not exactly at a
- 3:52floating point number so do i go
- 3:54the one above or the one below that's
- 3:56the idea of this
- 3:57and there's some hardware that usually
- 4:00carries with
- 4:00the usually floating point hardware has
- 4:03a couple of extra bits
- 4:04on the far right side that's called
- 4:06rounding bits
- 4:07and those extra bits can be used to
- 4:09figure out where do i really want to go
- 4:11the floating point number above
- 4:13or the one below in terms of that so
- 4:14this is the 32 bits and there's extra
- 4:16bits on the right side of that
- 4:17that's useful so there are modes when
- 4:20i'm in the middle let's say i'm some you
- 4:21know the rounding bits say that there's
- 4:23something in here that's not a zero down
- 4:24here my rounding bits means there
- 4:25actually is something to the right of
- 4:26that
- 4:28you could either there's four rounding
- 4:29modes you could either round
- 4:31always toward positive infinity always
- 4:33to the right on the number line
- 4:35or always to the left of the number line
- 4:36always towards minus infinity
- 4:38or just truncating which actually means
- 4:40always rounding it towards zero
- 4:42or here's the more clever way you could
- 4:45have it in an
- 4:46unbiased way rather than always bigger
- 4:48or always smaller
- 4:49or always towards zero you could
- 4:51sometimes go up and sometimes go down
- 4:54and that means when you're midway well
- 4:56this by the way these examples i'm
- 4:57showing you are all examples
- 4:59in decimal this is not the same as
- 5:01binary but you're obviously working with
- 5:03binary we're working with floating point
- 5:04calculations so in the first example you
- 5:07round up that means the 2.001 well you
- 5:10round toward the next
- 5:11that's in the middle between two and
- 5:12three that would round up to three
- 5:14okay and minus two would round to two we
- 5:16always round to the right
- 5:17the same idea always round left one
- 5:19point nine nine nine rounds to one even
- 5:20though you're almost at two you go back
- 5:22to one you lose it
- 5:24and minus 1.99 also goes down also goes
- 5:26negative
- 5:27so here's the case when you're in the
- 5:29middle so normal rounding
- 5:31um 2.4 again this is in decimal
- 5:34is less than halfway so it round down to
- 5:362. 2.6 we're
- 5:38more than halfway round up 2.5 is
- 5:42exactly in the middle
- 5:43so that was going to round toward the
- 5:45even number so that rounds to
- 5:472 but 3.5 would round to 4. 4
- 5:51nice so think about this i've got some
- 5:55binary equivalent of this but this is
- 5:56now in binary
- 5:57so i've got some binary numbers here
- 5:59okay 0
- 6:011 0 1. here are the rounding areas okay
- 6:05so now
- 6:07what's halfway between this by the way
- 6:09if i have many bits here let's say i
- 6:11have let's say i have three bits let's
- 6:12just say i have three bits
- 6:14those bits would be from zero zero zero
- 6:17to one one
- 6:18one what's halfway in the middle in
- 6:20binary
- 6:21remember what half is this yeah it'd be
- 6:24a one
- 6:26and zeros that's halfway in the middle
- 6:29so when you're that way that's when this
- 6:31k when when the rounding vowel
- 6:33bits are one and all zeros that's
- 6:35halfway in the middle
- 6:36and then you're making a decision in
- 6:37this particular mode which way to go
- 6:40well you round toward the even number so
- 6:43you ask is that an even number no that's
- 6:46a one that's an odd number
- 6:47so therefore i'm going to round not
- 6:49toward that number but toward the one
- 6:50above
- 6:51and that would round up so this would
- 6:53round to 0
- 6:541 1 0. okay
- 6:57if i had 0 1
- 7:01zero and then here's the rounding guys
- 7:03and i were one
- 7:04zero i would round toward the even i
- 7:06would round down
- 7:07which means around down toward the even
- 7:09and i would just cross it off and that
- 7:11would be the closest even
- 7:12does that make sense so you round toward
- 7:14even in binary
- 7:15which means you have to be a one zero
- 7:17zero zero zero zero zero in the rounding
- 7:19bits and you round toward the even
- 7:20number if it's a one there you round up
- 7:22you add one to it if it's a zero
- 7:23you just truncate those rounding bits
- 7:25okay that way we kind of round
- 7:27both ways and that's kind of that's
- 7:29really we really like that as a default
- 7:30mode so that's the default mode for the
- 7:31system
- 7:32because it kind of is fair it doesn't
- 7:33always go make things bigger or make
- 7:35things smaller
- 7:36um however irs if you're calculating my
- 7:38my return
- 7:39can you round up always just don't round
- 7:43the middle so now you know when you see
- 7:46these errors here's this wonderful comic
- 7:48strip i found online
- 7:50uh and here is a robot and it says
- 7:53welcome to the secret robot internet
- 7:55prove that you're a human point one plus
- 7:570.2 0.3004
- 8:01we can represent powers of two perfectly
- 8:03fractional powers are two perfectly in
- 8:05floating point
- 8:05you know because it's zero in the
- 8:06significant and some value of the
- 8:08exponent
- 8:09but but because we're decimal we can't
- 8:12do point one or point two
- 8:13very easily because it's one tenth this
- 8:15ten thing is different there's ten
- 8:16fingers with a problem
- 8:18so this is an example of point three if
- 8:20you type point three into my simulator
- 8:22you get 0.300 and it's not exactly
- 8:26o4 maybe that's a double but in the
- 8:27float in a single precision or 32 bits
- 8:30of a float we get 0.3001 so this is
- 8:32exactly why
- 8:33you would see this error
- 8:37addition let's talk about addition now
- 8:38we're just talking about elements of
- 8:39this
- 8:40rounding and addition and truncating and
- 8:42all those things
- 8:43so addition uh is a little bit harder
- 8:46uh because you can't just add the
- 8:48significance um you have to denormalize
- 8:51to match the exponent you have to kind
- 8:52of line them up okay
- 8:54to match the exponents then you add the
- 8:57significance to get the resulting guy
- 8:59and then you have to keep the same
- 9:00exponent and then normalize to get
- 9:01the end so you have to kind of shift it
- 9:03do something and then shift it back
- 9:04whatever the result is they may cancel
- 9:06each other out to make something
- 9:07you may have a big number and a big
- 9:08number but it turns out that they cancel
- 9:10most of the big stuff and the small
- 9:11values of what's left then you have to
- 9:13renormalize to get that right if the
- 9:14signs
- 9:15differ you just do a subtraction instead
- 9:18this is how you do casting you've seen
- 9:20casting when you do malik
- 9:22when you go to the godfather you have to
- 9:24cast that un
- 9:25you know uninitialized void star space
- 9:28into whatever type you want so that you
- 9:29know
- 9:30you know pointer plus plus knows how
- 9:31much to jump that's the reason we do
- 9:33that
- 9:34and also to check you know whether
- 9:35you're whether you're matching these
- 9:36right and everything ever the pointers
- 9:38work out
- 9:40type tracking obviously
- 9:43here if you've got a floating point
- 9:45number you want to convert to an integer
- 9:46just say
- 9:46int if you remember in python it was it
- 9:49was in
- 9:50parentheses there's a function call to
- 9:51that as a method call to that here you
- 9:53just type cache with the parenthesis in
- 9:54so here's 3.14159 times some float
- 9:58that's obviously a float and now you
- 9:59cast that to an end to figure what the
- 10:01closest energy
- 10:02that might be conversely you do the same
- 10:04thing on the other side if you have an
- 10:05integer number you want to float convert
- 10:07it like that so you say
- 10:08float of that integer expression and now
- 10:10you're able to work with that so if i
- 10:11have f as a float i wanna
- 10:13f plus equal another float well if i
- 10:15have i which is int you have to convert
- 10:17it to a float so
- 10:18that the compiler will be happy so
- 10:21now you might ask yourself well if i go
- 10:23from an into a flow back to an int
- 10:24do i get the same number does this
- 10:26always print true
- 10:30pause the video and think about it talk
- 10:32to your neighbor and then come on back
- 10:35and we're back if i equal equal int a
- 10:39float of i
- 10:40same thing ain't
- 10:43no free lunch there are integers
- 10:48that the float can't handle
- 10:51so there are certain integers in fact i
- 10:54even showed you one i think it was
- 10:55a little bit more than 16 million that
- 10:58has no
- 10:59float equivalent you know you can count
- 11:01up to four billion in in floats and
- 11:03integers
- 11:04unsigned in numbers but floats the first
- 11:06number you can't do we saw that before
- 11:0816 million plus one i can't do that
- 11:11number 16 million plus one is
- 11:13that can't be done so i mean two to the
- 11:15two to the fourteen two to the 24
- 11:18plus one can't be done i showed you
- 11:20before
- 11:21so we do that it's going to snap it to
- 11:24the closest float
- 11:25when you go back to int it's the back
- 11:27end the same end i can do that it'll
- 11:28either be
- 11:29not think about rounding modes that
- 11:31number is two to the 24
- 11:33plus one can't be done so does it snap
- 11:35to the fourth part does it do 22 24
- 11:38or 2 to the 24 plus 2 well this is an
- 11:41even number so it'll round to even
- 11:43so that's an exact case of that rounding
- 11:45mode there that's going to be
- 11:462 to the 20. try it 2 to the 24
- 11:50plus 1. see what it says that'd be
- 11:52really fun to play with
- 11:53okay can't be done sorry
- 11:57double precision so let's talk about
- 11:59double precision could that do it you
- 12:00want to ask yourself look at the
- 12:02double precision could double precision
- 12:03do that as well
- 12:06float to int to float does this always
- 12:09return true
- 12:10pause the video and think about it and
- 12:12come on back
- 12:14and we're back ain't no free lunch
- 12:19what if the float is 1.5 this is easiest
- 12:21one 1.5
- 12:22obviously there's no int that does 1.5
- 12:24and you go back it's going to be
- 12:26the integer representation that'd be
- 12:27either one or two what do you think it
- 12:29does
- 12:301.5 round to even
- 12:33the closest even number is 2. so that
- 12:36should in theory be
- 12:37two when it comes back so that should
- 12:39not be the same that's an obvious one
- 12:41why that doesn't work we'll see the next
- 12:42video
About this transcript
This page contains the full transcript of [CS61C FA20] Lecture 06.5 - Floating Point: Floating Point Discussion by CS 61C Departmental, generated from the public captions YouTube serves with the video. The transcript has 2,502 words across 400 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.