The spelled-out intro to neural networks and backpropagation: building micrograd — Transcript
Full transcript
- 0:00hello my name is andre
- 0:01and i've been training deep neural
- 0:02networks for a bit more than a decade
- 0:04and in this lecture i'd like to show you
- 0:06what neural network training looks like
- 0:08under the hood so in particular we are
- 0:10going to start with a blank jupiter
- 0:12notebook and by the end of this lecture
- 0:14we will define and train in neural net
- 0:16and you'll get to see everything that
- 0:18goes on under the hood and exactly
- 0:20sort of how that works on an intuitive
- 0:21level
- 0:22now specifically what i would like to do
- 0:24is i would like to take you through
- 0:26building of micrograd now micrograd is
- 0:29this library that i released on github
- 0:30about two years ago but at the time i
- 0:32only uploaded the source code and you'd
- 0:34have to go in by yourself and really
- 0:37figure out how it works
- 0:39so in this lecture i will take you
- 0:40through it step by step and kind of
- 0:42comment on all the pieces of it so what
- 0:44is micrograd and why is it interesting
- 0:47good
- 0:48um
- 0:49micrograd is basically an autograd
- 0:51engine autograd is short for automatic
- 0:53gradient and really what it does is it
- 0:55implements backpropagation now
- 0:57backpropagation is this algorithm that
- 0:59allows you to efficiently evaluate the
- 1:01gradient of
- 1:03some kind of a loss function with
- 1:05respect to the weights of a neural
- 1:07network and what that allows us to do
- 1:09then is we can iteratively tune the
- 1:11weights of that neural network to
- 1:12minimize the loss function and therefore
- 1:14improve the accuracy of the network so
- 1:16back propagation would be at the
- 1:18mathematical core of any modern deep
- 1:20neural network library like say pytorch
- 1:22or jaxx
- 1:24so the functionality of microgrant is i
- 1:25think best illustrated by an example so
- 1:27if we just scroll down here
- 1:29you'll see that micrograph basically
- 1:31allows you to build out mathematical
- 1:32expressions
- 1:34and um here what we are doing is we have
- 1:36an expression that we're building out
- 1:37where you have two inputs a and b
- 1:40and you'll see that a and b are negative
- 1:43four and two but we are wrapping those
- 1:46values into this value object that we
- 1:48are going to build out as part of
- 1:49micrograd
- 1:51so this value object will wrap the
- 1:53numbers themselves
- 1:54and then we are going to build out a
- 1:56mathematical expression here where a and
- 1:58b are transformed into c d and
- 2:01eventually e f and g
- 2:03and i'm showing some of the functions
- 2:05some of the functionality of micrograph
- 2:07and the operations that it supports so
- 2:08you can add two value objects you can
- 2:11multiply them you can raise them to a
- 2:13constant power you can offset by one
- 2:15negate squash at zero
- 2:18square divide by constant divide by it
- 2:21etc
- 2:22and so we're building out an expression
- 2:24graph with with these two inputs a and b
- 2:27and we're creating an output value of g
- 2:30and micrograd will in the background
- 2:32build out this entire mathematical
- 2:34expression so it will for example know
- 2:36that c is also a value
- 2:38c was a result of an addition operation
- 2:41and the
- 2:42child nodes of c are a and b because the
- 2:46and will maintain pointers to a and b
- 2:48value objects so we'll basically know
- 2:50exactly how all of this is laid out
- 2:53and then not only can we do what we call
- 2:55the forward pass where we actually look
- 2:57at the value of g of course that's
- 2:58pretty straightforward we will access
- 3:00that using the dot data attribute and so
- 3:03the output of the forward pass the value
- 3:06of g is 24.7 it turns out but the big
- 3:09deal is that we can also take this g
- 3:11value object and we can call that
- 3:13backward
- 3:14and this will basically uh initialize
- 3:16back propagation at the node g
- 3:19and what backpropagation is going to do
- 3:21is it's going to start at g and it's
- 3:23going to go backwards through that
- 3:25expression graph and it's going to
- 3:26recursively apply the chain rule from
- 3:28calculus
- 3:30and what that allows us to do then is
- 3:32we're going to evaluate basically the
- 3:34derivative of g with respect to all the
- 3:36internal nodes
- 3:38like e d and c but also with respect to
- 3:40the inputs a and b
- 3:43and then we can actually query this
- 3:45derivative of g with respect to a for
- 3:47example that's a dot grad in this case
- 3:50it happens to be 138 and the derivative
- 3:52of g with respect to b
- 3:54which also happens to be here 645
- 3:57and this derivative we'll see soon is
- 3:59very important information because it's
- 4:01telling us how a and b are affecting g
- 4:04through this mathematical expression so
- 4:06in particular
- 4:08a dot grad is 138 so if we slightly
- 4:11nudge a and make it slightly larger
- 4:14138 is telling us that g will grow and
- 4:18the slope of that growth is going to be
- 4:19138
- 4:20and the slope of growth of b is going to
- 4:22be 645. so that's going to tell us about
- 4:25how g will respond if a and b get
- 4:27tweaked a tiny amount in a positive
- 4:29direction
- 4:31okay
- 4:33now you might be confused about what
- 4:34this expression is that we built out
- 4:36here and this expression by the way is
- 4:38completely meaningless i just made it up
- 4:40i'm just flexing about the kinds of
- 4:42operations that are supported by
- 4:43micrograd
- 4:44what we actually really care about are
- 4:46neural networks but it turns out that
- 4:48neural networks are just mathematical
- 4:49expressions just like this one but
- 4:51actually slightly bit less crazy even
- 4:54neural networks are just a mathematical
- 4:56expression they take the input data as
- 4:59an input and they take the weights of a
- 5:00neural network as an input and it's a
- 5:02mathematical expression and the output
- 5:04are your predictions of your neural net
- 5:06or the loss function we'll see this in a
- 5:08bit but basically neural networks just
- 5:10happen to be a certain class of
- 5:12mathematical expressions
- 5:13but back propagation is actually
- 5:15significantly more general it doesn't
- 5:17actually care about neural networks at
- 5:18all it only tells us about arbitrary
- 5:20mathematical expressions and then we
- 5:22happen to use that machinery for
- 5:24training of neural networks now one more
- 5:26note i would like to make at this stage
- 5:28is that as you see here micrograd is a
- 5:30scalar valued auto grant engine so it's
- 5:32working on the you know level of
- 5:34individual scalars like negative four
- 5:36and two and we're taking neural nets and
- 5:37we're breaking them down all the way to
- 5:39these atoms of individual scalars and
- 5:41all the little pluses and times and it's
- 5:43just excessive and so obviously you
- 5:45would never be doing any of this in
- 5:47production it's really just put down for
- 5:48pedagogical reasons because it allows us
- 5:50to not have to deal with these
- 5:52n-dimensional tensors that you would use
- 5:54in modern deep neural network library so
- 5:56this is really done so that you
- 5:58understand and refactor out back
- 6:00propagation and chain rule and
- 6:02understanding of neurologic training
- 6:04and then if you actually want to train
- 6:06bigger networks you have to be using
- 6:08these tensors but none of the math
- 6:09changes this is done purely for
- 6:11efficiency we are basically taking scale
- 6:13value
- 6:14all the scale values we're packaging
- 6:16them up into tensors which are just
- 6:17arrays of these scalars and then because
- 6:20we have these large arrays we're making
- 6:22operations on those large arrays that
- 6:24allows us to take advantage of the
- 6:26parallelism in a computer and all those
- 6:28operations can be done in parallel and
- 6:30then the whole thing runs faster but
- 6:32really none of the math changes and
- 6:33that's done purely for efficiency so i
- 6:35don't think that it's pedagogically
- 6:36useful to be dealing with tensors from
- 6:38scratch uh and i think and that's why i
- 6:40fundamentally wrote micrograd because
- 6:42you can understand how things work uh at
- 6:44the fundamental level and then you can
- 6:46speed it up later okay so here's the fun
- 6:48part my claim is that micrograd is what
- 6:51you need to train your networks and
- 6:52everything else is just efficiency so
- 6:54you'd think that micrograd would be a
- 6:56very complex piece of code and that
- 6:58turns out to not be the case
- 7:01so if we just go to micrograd
- 7:03and you'll see that there's only two
- 7:05files here in micrograd this is the
- 7:07actual engine it doesn't know anything
- 7:09about neural nuts and this is the entire
- 7:10neural nets library
- 7:12on top of micrograd so engine and nn.pi
- 7:17so the actual backpropagation autograd
- 7:19engine
- 7:21that gives you the power of neural
- 7:22networks is literally
- 7:26100 lines of code of like very simple
- 7:28python
- 7:30which we'll understand by the end of
- 7:31this lecture
- 7:32and then nn.pi
- 7:33this neural network library built on top
- 7:35of the autograd engine
- 7:37um is like a joke it's like
- 7:40we have to define what is a neuron and
- 7:42then we have to define what is the layer
- 7:44of neurons and then we define what is a
- 7:46multi-layer perceptron which is just a
- 7:47sequence of layers of neurons and so
- 7:50it's just a total joke
- 7:52so basically
- 7:53there's a lot of power that comes from
- 7:55only 150 lines of code
- 7:57and that's all you need to understand to
- 7:59understand neural network training and
- 8:00everything else is just efficiency and
- 8:02of course there's a lot to efficiency
- 8:05but fundamentally that's all that's
- 8:07happening okay so now let's dive right
- 8:09in and implement micrograph step by step
- 8:11the first thing i'd like to do is i'd
- 8:12like to make sure that you have a very
- 8:13good understanding intuitively of what a
- 8:16derivative is and exactly what
- 8:18information it gives you so let's start
- 8:20with some basic imports that i copy
- 8:22paste in every jupiter notebook always
- 8:25and let's define a function a scalar
- 8:27valued function
- 8:28f of x
- 8:30as follows
- 8:31so i just make this up randomly i just
- 8:33want to scale a valid function that
- 8:34takes a single scalar x and returns a
- 8:36single scalar y
- 8:38and we can call this function of course
- 8:40so we can pass in say 3.0 and get 20
- 8:42back
- 8:43now we can also plot this function to
- 8:45get a sense of its shape you can tell
- 8:47from the mathematical expression that
- 8:48this is probably a parabola it's a
- 8:50quadratic
- 8:51and so if we just uh create a set of um
- 8:56um
- 8:57scale values that we can feed in using
- 8:59for example a range from negative five
- 9:01to five in steps of 0.25
- 9:03so this is so axis is just from negative
- 9:065 to 5 not including 5 in steps of 0.25
- 9:11and we can actually call this function
- 9:12on this numpy array as well so we get a
- 9:14set of y's if we call f on axis
- 9:17and these y's are basically
- 9:20also applying a function on every one of
- 9:23these elements independently
- 9:25and we can plot this using matplotlib so
- 9:28plt.plot x's and y's and we get a nice
- 9:31parabola so previously here we fed in
- 9:333.0 somewhere here and we received 20
- 9:36back which is here the y coordinate so
- 9:39now i'd like to think through
- 9:40what is the derivative
- 9:42of this function at any single input
- 9:44point x
- 9:45right so what is the derivative at
- 9:47different points x of this function now
- 9:49if you remember back to your calculus
- 9:51class you've probably derived
- 9:52derivatives so we take this mathematical
- 9:54expression 3x squared minus 4x plus 5
- 9:57and you would write out on a piece of
- 9:58paper and you would you know apply the
- 9:59product rule and all the other rules and
- 10:01derive the mathematical expression of
- 10:03the great derivative of the original
- 10:05function and then you could plug in
- 10:06different texts and see what the
- 10:08derivative is
- 10:09we're not going to actually do that
- 10:11because no one in neural networks
- 10:13actually writes out the expression for
- 10:15the neural net it would be a massive
- 10:16expression um it would be you know
- 10:18thousands tens of thousands of terms no
- 10:20one actually derives the derivative of
- 10:22course and so we're not going to take
- 10:24this kind of like a symbolic approach
- 10:26instead what i'd like to do is i'd like
- 10:27to look at the definition of derivative
- 10:29and just make sure that we really
- 10:30understand what derivative is measuring
- 10:32what it's telling you about the function
- 10:34and so if we just look up derivative
- 10:42we see that
- 10:43okay so this is not a very good
- 10:44definition of derivative this is a
- 10:46definition of what it means to be
- 10:47differentiable
- 10:48but if you remember from your calculus
- 10:50it is the limit as h goes to zero of f
- 10:52of x plus h minus f of x over h so
- 10:55basically what it's saying is if you
- 10:58slightly bump up you're at some point x
- 11:00that you're interested in or a and if
- 11:02you slightly bump up
- 11:04you know you slightly increase it by
- 11:06small number h
- 11:08how does the function respond with what
- 11:09sensitivity does it respond what is the
- 11:11slope at that point does the function go
- 11:13up or does it go down and by how much
- 11:16and that's the slope of that function
- 11:18the
- 11:18the slope of that response at that point
- 11:21and so we can basically evaluate
- 11:23the derivative here numerically by
- 11:26taking a very small h of course the
- 11:28definition would ask us to take h to
- 11:30zero we're just going to pick a very
- 11:31small h 0.001
- 11:34and let's say we're interested in point
- 11:353.0 so we can look at f of x of course
- 11:37as 20
- 11:38and now f of x plus h
- 11:40so if we slightly nudge x in a positive
- 11:42direction how is the function going to
- 11:44respond
- 11:45and just looking at this do you expect
- 11:47do you expect f of x plus h to be
- 11:49slightly greater than 20 or do you
- 11:51expect to be slightly lower than 20
- 11:54and since this 3 is here and this is 20
- 11:57if we slightly go positively the
- 11:59function will respond positively so
- 12:01you'd expect this to be slightly greater
- 12:03than 20. and now by how much it's
- 12:05telling you the
- 12:06sort of the
- 12:07the strength of that slope right the the
- 12:09size of the slope so f of x plus h minus
- 12:12f of x this is how much the function
- 12:14responded
- 12:16in the positive direction and we have to
- 12:17normalize by the
- 12:19run so we have the rise over run to get
- 12:22the slope so this of course is just a
- 12:24numerical approximation of the slope
- 12:26because we have to make age very very
- 12:28small to converge to the exact amount
- 12:32now if i'm doing too many zeros
- 12:35at some point
- 12:36i'm gonna get an incorrect answer
- 12:38because we're using floating point
- 12:39arithmetic and the representations of
- 12:41all these numbers in computer memory is
- 12:43finite and at some point we get into
- 12:45trouble
- 12:46so we can converse towards the right
- 12:47answer with this approach
- 12:50but basically um at 3 the slope is 14.
- 12:54and you can see that by taking 3x
- 12:56squared minus 4x plus 5 and
- 12:58differentiating it in our head
- 13:00so 3x squared would be
- 13:026 x minus 4
- 13:04and then we plug in x equals 3 so that's
- 13:0718 minus 4 is 14. so this is correct
- 13:10so that's
- 13:12at 3. now how about the slope at say
- 13:15negative 3
- 13:17would you expect would you expect for
- 13:19the slope
- 13:20now telling the exact value is really
- 13:22hard but what is the sign of that slope
- 13:24so at negative three
- 13:26if we slightly go in the positive
- 13:28direction at x the function would
- 13:30actually go down and so that tells you
- 13:32that the slope would be negative so
- 13:33we'll get a slight number below
- 13:36below 20. and so if we take the slope we
- 13:39expect something negative
- 13:40negative 22. okay
- 13:43and at some point here of course the
- 13:45slope would be zero now for this
- 13:47specific function i looked it up
- 13:48previously and it's at point two over
- 13:51three
- 13:52so at roughly two over three
- 13:54uh that's somewhere here
- 13:55um
- 13:57this derivative be zero
- 13:59so basically at that precise point
- 14:03yeah
- 14:04at that precise point if we nudge in a
- 14:06positive direction the function doesn't
- 14:07respond this stays the same almost and
- 14:09so that's why the slope is zero okay now
- 14:11let's look at a bit more complex case
- 14:14so we're going to start you know
- 14:15complexifying a bit so now we have a
- 14:18function
- 14:19here
- 14:20with output variable d
- 14:22that is a function of three scalar
- 14:24inputs a b and c
- 14:26so a b and c are some specific values
- 14:28three inputs into our expression graph
- 14:30and a single output d
- 14:32and so if we just print d we get four
- 14:36and now what i have to do is i'd like to
- 14:38again look at the derivatives of d with
- 14:40respect to a b and c
- 14:42and uh think through uh again just the
- 14:44intuition of what this derivative is
- 14:46telling us
- 14:47so in order to evaluate this derivative
- 14:49we're going to get a bit hacky here
- 14:52we're going to again have a very small
- 14:53value of h
- 14:55and then we're going to fix the inputs
- 14:57at some
- 14:58values that we're interested in
- 15:00so these are the this is the point abc
- 15:02at which we're going to be evaluating
- 15:04the the
- 15:05derivative of d with respect to all a b
- 15:07and c at that point
- 15:09so there are the inputs and now we have
- 15:11d1 is that expression
- 15:13and then we're going to for example look
- 15:15at the derivative of d with respect to a
- 15:17so we'll take a and we'll bump it by h
- 15:19and then we'll get d2 to be the exact
- 15:22same function
- 15:23and now we're going to print um
- 15:26you know f1
- 15:28d1 is d1
- 15:31d2 is d2
- 15:32and print slope
- 15:35so the derivative or slope
- 15:37here will be um
- 15:39of course
- 15:41d2
- 15:42minus d1 divide h
- 15:44so d2 minus d1 is how much the function
- 15:47increased
- 15:48uh when we bumped
- 15:50the uh
- 15:51the specific input that we're interested
- 15:53in by a tiny amount
- 15:55and
- 15:56this is then normalized by h
- 15:59to get the slope
- 16:02so
- 16:03um
- 16:05yeah
- 16:06so this so if i just run this we're
- 16:08going to print
- 16:10d1
- 16:12which we know is four
- 16:15now d2 will be bumped a will be bumped
- 16:18by h
- 16:20so let's just think through
- 16:22a little bit uh what d2 will be uh
- 16:26printed out here
- 16:27in particular
- 16:29d1 will be four
- 16:31will d2 be a number slightly greater
- 16:33than four or slightly lower than four
- 16:35and that's going to tell us the sl the
- 16:37the sign of the derivative
- 16:40so
- 16:42we're bumping a by h
- 16:45b as minus three c is ten
- 16:48so you can just intuitively think
- 16:49through this derivative and what it's
- 16:51doing a will be slightly more positive
- 16:54and but b is a negative number
- 16:57so if a is slightly more positive
- 17:00because b is negative three
- 17:03we're actually going to be adding less
- 17:06to d
- 17:08so you'd actually expect that the value
- 17:10of the function will go down
- 17:13so let's just see this
- 17:16yeah and so we went from 4
- 17:18to 3.9996
- 17:20and that tells you that the slope will
- 17:22be negative
- 17:23and then
- 17:24uh will be a negative number
- 17:26because we went down
- 17:27and then
- 17:29the exact number of slope will be
- 17:31exact amount of slope is negative 3.
- 17:33and you can also convince yourself that
- 17:35negative 3 is the right answer
- 17:36mathematically and analytically because
- 17:39if you have a times b plus c and you are
- 17:41you know you have calculus then
- 17:43differentiating a times b plus c with
- 17:46respect to a gives you just b
- 17:48and indeed the value of b is negative 3
- 17:50which is the derivative that we have so
- 17:52you can tell that that's correct
- 17:54so now if we do this with b
- 17:57so if we bump b by a little bit in a
- 17:59positive direction we'd get different
- 18:02slopes so what is the influence of b on
- 18:04the output d
- 18:06so if we bump b by a tiny amount in a
- 18:08positive direction then because a is
- 18:10positive
- 18:11we'll be adding more to d
- 18:13right
- 18:14so um and now what is the what is the
- 18:17sensitivity what is the slope of that
- 18:18addition
- 18:19and it might not surprise you that this
- 18:21should be
- 18:222
- 18:24and y is a 2 because d of d
- 18:27by db differentiating with respect to b
- 18:30would be would give us a
- 18:31and the value of a is two so that's also
- 18:34working well
- 18:35and then if c gets bumped a tiny amount
- 18:37in h
- 18:38by h
- 18:39then of course a times b is unaffected
- 18:41and now c becomes slightly bit higher
- 18:44what does that do to the function it
- 18:45makes it slightly bit higher because
- 18:47we're simply adding c
- 18:48and it makes it slightly bit higher by
- 18:50the exact same amount that we added to c
- 18:53and so that tells you that the slope is
- 18:55one
- 18:56that will be the
- 18:59the rate at which
- 19:01d will increase as we scale
- 19:04c
- 19:05okay so we now have some intuitive sense
- 19:06of what this derivative is telling you
- 19:08about the function and we'd like to move
- 19:10to neural networks now as i mentioned
- 19:11neural networks will be pretty massive
- 19:13expressions mathematical expressions so
- 19:15we need some data structures that
- 19:16maintain these expressions and that's
- 19:17what we're going to start to build out
- 19:19now
- 19:20so we're going to
- 19:22build out this value object that i
- 19:24showed you in the readme page of
- 19:26micrograd
- 19:27so let me copy paste a skeleton of the
- 19:30first very simple value object
- 19:33so class value takes a single
- 19:36scalar value that it wraps and keeps
- 19:38track of
- 19:39and that's it so
- 19:41we can for example do value of 2.0 and
- 19:43then we can
- 19:45get we can look at its content and
- 19:48python will internally
- 19:50use the wrapper function
- 19:52to uh return
- 19:54uh this string oops
- 19:56like that
- 19:58so this is a value object with data
- 20:00equals two that we're creating here
- 20:03now we'd like to do is like we'd like to
- 20:04be able to
- 20:07have not just like two values
- 20:10but we'd like to do a bluffy right we'd
- 20:12like to add them
- 20:13so currently you would get an error
- 20:15because python doesn't know how to add
- 20:17two value objects so we have to tell it
- 20:21so here's
- 20:22addition
- 20:26so you have to basically use these
- 20:27special double underscore methods in
- 20:29python to define these operators for
- 20:31these objects so if we call um
- 20:35the uh if we use this plus operator
- 20:39python will internally call a dot add of
- 20:43b
- 20:43that's what will happen internally and
- 20:45so b will be the other and
- 20:48self will be a
- 20:50and so we see that what we're going to
- 20:52return is a new value object and it's
- 20:54just it's going to be wrapping
- 20:56the plus of
- 20:58their data
- 20:59but remember now because data is the
- 21:02actual like numbered python number so
- 21:04this operator here is just the typical
- 21:06floating point plus addition now it's
- 21:09not an addition of value objects
- 21:11and will return a new value so now a
- 21:14plus b should work and it should print
- 21:16value of
- 21:17negative one
- 21:18because that's two plus minus three
- 21:20there we go
- 21:21okay let's now implement multiply
- 21:24just so we can recreate this expression
- 21:25here
- 21:26so multiply i think it won't surprise
- 21:28you will be fairly similar
- 21:31so instead of add we're going to be
- 21:33using mul
- 21:34and then here of course we want to do
- 21:36times
- 21:36and so now we can create a c value
- 21:38object which will be 10.0 and now we
- 21:41should be able to do a times b well
- 21:44let's just do a times b first
- 21:46um
- 21:47[Music]
- 21:48that's value of negative six now
- 21:50and by the way i skipped over this a
- 21:52little bit suppose that i didn't have
- 21:53the wrapper function here
- 21:55then it's just that you'll get some kind
- 21:57of an ugly expression so what wrapper is
- 21:59doing is it's providing us a way to
- 22:02print out like a nicer looking
- 22:03expression in python
- 22:05uh so we don't just have something
- 22:07cryptic we actually are you know it's
- 22:09value of
- 22:10negative six so this gives us a times
- 22:14and then this we should now be able to
- 22:16add c to it because we've defined and
- 22:18told the python how to do mul and add
- 22:20and so this will call this will
- 22:22basically be equivalent to a dot
- 22:24small
- 22:26of b
- 22:27and then this new value object will be
- 22:29dot add
- 22:31of c
- 22:32and so let's see if that worked
- 22:34yep so that worked well that gave us
- 22:36four which is what we expect from before
- 22:39and i believe we can just call them
- 22:40manually as well there we go so
- 22:44yeah
- 22:45okay so now what we are missing is the
- 22:46connective tissue of this expression as
- 22:49i mentioned we want to keep these
- 22:50expression graphs so we need to know and
- 22:52keep pointers about what values produce
- 22:54what other values
- 22:56so here for example we are going to
- 22:58introduce a new variable which we'll
- 23:00call children and by default it will be
- 23:02an empty tuple
- 23:03and then we're actually going to keep a
- 23:04slightly different variable in the class
- 23:06which we'll call underscore prev which
- 23:08will be the set of children
- 23:11this is how i done i did it in the
- 23:13original micrograd looking at my code
- 23:15here i can't remember exactly the reason
- 23:17i believe it was efficiency but this
- 23:19underscore children will be a tuple for
- 23:20convenience but then when we actually
- 23:22maintain it in the class it will be just
- 23:23this set yeah i believe for efficiency
- 23:27um
- 23:28so now
- 23:29when we are creating a value like this
- 23:31with a constructor children will be
- 23:33empty and prep will be the empty set but
- 23:36when we're creating a value through
- 23:37addition or multiplication we're going
- 23:39to feed in the children of this value
- 23:42which in this case is self and other
- 23:46so those are the children
- 23:48here
- 23:50so now we can do d dot prev
- 23:52and we'll see that the children of the
- 23:55we now know are this value of negative 6
- 23:58and value of 10 and this of course is
- 24:00the value resulting from a times b and
- 24:03the c value which is 10.
- 24:06now the last piece of information we
- 24:08don't know so we know that the children
- 24:10of every single value but we don't know
- 24:12what operation created this value
- 24:14so we need one more element here let's
- 24:16call it underscore pop
- 24:19and by default this is the empty set for
- 24:21leaves
- 24:22and then we'll just maintain it here
- 24:25and now the operation will be just a
- 24:27simple string and in the case of
- 24:29addition it's plus in the case of
- 24:31multiplication is times
- 24:33so now we
- 24:35not just have d dot pref we also have a
- 24:37d dot up
- 24:38and we know that d was produced by an
- 24:40addition of those two values and so now
- 24:42we have the full
- 24:44mathematical expression uh and we're
- 24:46building out this data structure and we
- 24:47know exactly how each value came to be
- 24:49by word expression and from what other
- 24:51values
- 24:54now because these expressions are about
- 24:56to get quite a bit larger we'd like a
- 24:58way to nicely visualize these
- 25:00expressions that we're building out so
- 25:02for that i'm going to copy paste a bunch
- 25:03of slightly scary code that's going to
- 25:06visualize this these expression graphs
- 25:08for us
- 25:09so here's the code and i'll explain it
- 25:11in a bit but first let me just show you
- 25:13what this code does
- 25:14basically what it does is it creates a
- 25:16new function drawdot that we can call on
- 25:19some root node
- 25:20and then it's going to visualize it so
- 25:22if we call drawdot on d
- 25:24which is this final value here that is a
- 25:27times b plus c
- 25:29it creates something like this so this
- 25:31is d
- 25:32and you see that this is a times b
- 25:34creating an integrated value plus c
- 25:36gives us this output node d
- 25:40so that's dried out of d
- 25:42and i'm not going to go through this in
- 25:44complete detail you can take a look at
- 25:46graphless and its api uh graphis is a
- 25:48open source graph visualization software
- 25:51and what we're doing here is we're
- 25:52building out this graph and graphis
- 25:54api and
- 25:56you can basically see that trace is this
- 25:58helper function that enumerates all of
- 26:00the nodes and edges in the graph
- 26:02so that just builds a set of all the
- 26:04nodes and edges and then we iterate for
- 26:06all the nodes and we create special node
- 26:08objects
- 26:08for them in
- 26:11using dot node
- 26:13and then we also create edges using dot
- 26:15dot edge
- 26:16and the only thing that's like slightly
- 26:18tricky here is you'll notice that i
- 26:20basically add these fake nodes which are
- 26:22these operation nodes so for example
- 26:24this node here is just like a plus node
- 26:27and
- 26:28i create these
- 26:31special op nodes here
- 26:34and i connect them accordingly so these
- 26:37nodes of course are not actual
- 26:39nodes in the original graph
- 26:41they're not actually a value object the
- 26:43only value objects here are the things
- 26:46in squares those are actual value
- 26:48objects or representations thereof and
- 26:50these op nodes are just created in this
- 26:52drawdot routine so that it looks nice
- 26:55let's also add labels to these graphs
- 26:57just so we know what variables are where
- 26:59so let's create a special underscore
- 27:01label
- 27:02um
- 27:03or let's just do label
- 27:05equals empty by default and save it in
- 27:08each node
- 27:11and then here we're going to do label as
- 27:13a
- 27:15label is the
- 27:17label a c
- 27:22and then
- 27:24let's create a special um
- 27:27e equals a times b
- 27:30and e dot label will be e
- 27:34it's kind of naughty
- 27:35and e will be e plus c
- 27:38and a d dot label will be
- 27:40d
- 27:42okay so nothing really changes i just
- 27:44added this new e function
- 27:46a new e variable
- 27:48and then here when we are
- 27:50printing this
- 27:51i'm going to print the label here so
- 27:54this will be a percent s
- 27:56bar
- 27:56and this will be end.label
- 28:01and so now
- 28:03we have the label on the left here so it
- 28:05says a b creating e and then e plus c
- 28:07creates d
- 28:08just like we have it here
- 28:10and finally let's make this expression
- 28:12just one layer deeper
- 28:14so d will not be the final output node
- 28:17instead after d we are going to create a
- 28:20new value object
- 28:21called f we're going to start running
- 28:23out of variables soon f will be negative
- 28:252.0
- 28:27and its label will of course just be f
- 28:30and then l capital l will be the output
- 28:34of our graph
- 28:35and l will be p times f
- 28:38okay
- 28:38so l will be negative eight is the
- 28:40output
- 28:42so
- 28:44now we don't just draw a d we draw l
- 28:50okay
- 28:52and somehow the label of
- 28:54l was undefined oops all that label has
- 28:56to be explicitly sort of given to it
- 28:59there we go so l is the output
- 29:01so let's quickly recap what we've done
- 29:03so far
- 29:04we are able to build out mathematical
- 29:05expressions using only plus and times so
- 29:08far
- 29:09they are scalar valued along the way
- 29:11and we can do this forward pass
- 29:14and build out a mathematical expression
- 29:16so we have multiple inputs here a b c
- 29:18and f
- 29:19going into a mathematical expression
- 29:21that produces a single output l
- 29:24and this here is visualizing the forward
- 29:26pass so the output of the forward pass
- 29:28is negative eight that's the value
- 29:31now what we'd like to do next is we'd
- 29:33like to run back propagation
- 29:35and in back propagation we are going to
- 29:37start here at the end and we're going to
- 29:39reverse
- 29:40and calculate the gradient along along
- 29:43all these intermediate values
- 29:45and really what we're computing for
- 29:46every single value here
- 29:48um we're going to compute the derivative
- 29:50of that node with respect to l
- 29:55so
- 29:56the derivative of l with respect to l is
- 29:58just uh one
- 30:00and then we're going to derive what is
- 30:01the derivative of l with respect to f
- 30:03with respect to d with respect to c with
- 30:06respect to e
- 30:07with respect to b and with respect to a
- 30:10and in the neural network setting you'd
- 30:12be very interested in the derivative of
- 30:13basically this loss function l
- 30:16with respect to the weights of a neural
- 30:18network
- 30:19and here of course we have just these
- 30:20variables a b c and f
- 30:22but some of these will eventually
- 30:23represent the weights of a neural net
- 30:25and so we'll need to know how those
- 30:27weights are impacting
- 30:29the loss function so we'll be interested
- 30:31basically in the derivative of the
- 30:32output with respect to some of its leaf
- 30:34nodes and those leaf nodes will be the
- 30:36weights of the neural net
- 30:38and the other leaf nodes of course will
- 30:39be the data itself but usually we will
- 30:41not want or use the derivative of the
- 30:44loss function with respect to data
- 30:45because the data is fixed but the
- 30:47weights will be iterated on
- 30:50using the gradient information so next
- 30:52we are going to create a variable inside
- 30:54the value class that maintains the
- 30:57derivative of l with respect to that
- 30:59value
- 31:00and we will call this variable grad
- 31:03so there's a data and there's a
- 31:05self.grad
- 31:07and initially it will be zero and
- 31:09remember that zero is basically means no
- 31:12effect so at initialization we're
- 31:14assuming that every value does not
- 31:16impact does not affect the out the
- 31:18output
- 31:19right because if the gradient is zero
- 31:21that means that changing this variable
- 31:23is not changing the loss function
- 31:25so by default we assume that the
- 31:27gradient is zero
- 31:28and then
- 31:31now that we have grad and it's 0.0
- 31:36we are going to be able to visualize it
- 31:38here after data so here grad is 0.4 f
- 31:42and this will be in that graph
- 31:45and now we are going to be showing both
- 31:47the data and the grad
- 31:50initialized at zero
- 31:53and we are just about getting ready to
- 31:55calculate the back propagation
- 31:57and of course this grad again as i
- 31:58mentioned is representing
- 32:00the derivative of the output in this
- 32:02case l with respect to this value so
- 32:05with respect to so this is the
- 32:06derivative of l with respect to f with
- 32:08respect to d and so on so let's now fill
- 32:11in those gradients and actually do back
- 32:12propagation manually so let's start
- 32:14filling in these gradients and start all
- 32:16the way at the end as i mentioned here
- 32:18first we are interested to fill in this
- 32:20gradient here so what is the derivative
- 32:22of l with respect to l
- 32:25in other words if i change l by a tiny
- 32:27amount of h
- 32:29how much does
- 32:30l change
- 32:32it changes by h so it's proportional and
- 32:35therefore derivative will be one
- 32:37we can of course measure these or
- 32:39estimate these numerical gradients
- 32:40numerically just like we've seen before
- 32:43so if i take this expression
- 32:45and i create a def lol function here
- 32:49and put this here now the reason i'm
- 32:51creating a gating function hello here is
- 32:53because i don't want to pollute or mess
- 32:55up the global scope here this is just
- 32:57kind of like a little staging area and
- 32:58as you know in python all of these will
- 33:00be local variables to this function so
- 33:02i'm not changing any of the global scope
- 33:04here
- 33:05so here l1 will be l
- 33:10and then copy pasting this expression
- 33:13we're going to add a small amount h
- 33:17in for example a
- 33:20right and this would be measuring the
- 33:22derivative of l with respect to a
- 33:25so here this will be l2
- 33:28and then we want to print this
- 33:29derivative so print
- 33:31l2 minus l1 which is how much l changed
- 33:35and then normalize it by h so this is
- 33:37the rise over run
- 33:39and we have to be careful because l is a
- 33:41value node so we actually want its data
- 33:45um
- 33:46so that these are floats dividing by h
- 33:48and this should print the derivative of
- 33:50l with respect to a because a is the one
- 33:53that we bumped a little bit by h
- 33:55so what is the
- 33:57derivative of l with respect to a
- 33:59it's six
- 34:01okay and obviously
- 34:03if we change l by h
- 34:06then that would be
- 34:09here effectively
- 34:12this looks really awkward but changing l
- 34:14by h
- 34:16you see the derivative here is 1. um
- 34:20that's kind of like the base case of
- 34:23what we are doing here
- 34:24so basically we cannot come up here and
- 34:26we can manually set l.grad to one this
- 34:29is our manual back propagation
- 34:31l dot grad is one and let's redraw
- 34:35and we'll see that we filled in grad as
- 34:371 for l
- 34:39we're now going to continue the back
- 34:40propagation so let's here look at the
- 34:42derivatives of l with respect to d and f
- 34:45let's do a d first
- 34:47so what we are interested in if i create
- 34:49a markdown on here is we'd like to know
- 34:51basically we have that l is d times f
- 34:54and we'd like to know what is uh d
- 34:57l by d d
- 35:00what is that
- 35:01and if you know your calculus uh l is d
- 35:03times f so what is d l by d d it would
- 35:06be f
- 35:08and if you don't believe me we can also
- 35:10just derive it because the proof would
- 35:11be fairly straightforward uh we go to
- 35:14the
- 35:15definition of the derivative which is f
- 35:18of x plus h minus f of x divide h
- 35:22as a limit limit of h goes to zero of
- 35:24this kind of expression so when we have
- 35:26l is d times f
- 35:28then increasing d by h
- 35:31would give us the output of b plus h
- 35:33times f
- 35:35that's basically f of x plus h right
- 35:38minus d times f
- 35:42and then divide h and symbolically
- 35:44expanding out here we would have
- 35:46basically d times f plus h times f minus
- 35:50t times f divide h
- 35:52and then you see how the df minus df
- 35:54cancels so you're left with h times f
- 35:57divide h
- 35:58which is f
- 35:59so in the limit as h goes to zero of
- 36:03you know
- 36:04derivative
- 36:06definition we just get f in the case of
- 36:09d times f
- 36:12so
- 36:13symmetrically
- 36:14dl by d
- 36:15f will just be d
- 36:18so what we have is that f dot grad
- 36:21we see now is just the value of d
- 36:24which is 4.
- 36:28and we see that
- 36:30d dot grad
- 36:31is just uh the value of f
- 36:36and so the value of f is negative two
- 36:41so we'll set those manually
- 36:45let me erase this markdown node and then
- 36:47let's redraw what we have
- 36:50okay
- 36:51and let's just make sure that these were
- 36:53correct so we seem to think that dl by
- 36:56dd is negative two so let's double check
- 36:59um let me erase this plus h from before
- 37:02and now we want the derivative with
- 37:03respect to f
- 37:05so let's just come here when i create f
- 37:06and let's do a plus h here and this
- 37:08should print the derivative of l with
- 37:10respect to f so we expect to see four
- 37:14yeah and this is four up to floating
- 37:16point
- 37:17funkiness
- 37:18and then dl by dd
- 37:21should be f which is negative two
- 37:25grad is negative two
- 37:26so if we again come here and we change d
- 37:31d dot data plus equals h right here
- 37:35so we expect so we've added a little h
- 37:37and then we see how l changed and we
- 37:40expect to print
- 37:42uh negative two
- 37:44there we go
- 37:47so we've numerically verified what we're
- 37:49doing here is what kind of like an
- 37:50inline gradient check gradient check is
- 37:53when we
- 37:54are deriving this like back propagation
- 37:56and getting the derivative with respect
- 37:57to all the intermediate results and then
- 38:00numerical gradient is just you know
- 38:03estimating it using small step size
- 38:06now we're getting to the crux of
- 38:08backpropagation so this will be the most
- 38:10important node to understand because if
- 38:12you understand the gradient for this
- 38:14node you understand all of back
- 38:16propagation and all of training of
- 38:17neural nets basically
- 38:19so we need to derive dl by bc
- 38:23in other words the derivative of l with
- 38:24respect to c
- 38:26because we've computed all these other
- 38:27gradients already
- 38:29now we're coming here and we're
- 38:30continuing the back propagation manually
- 38:33so we want dl by dc and then we'll also
- 38:36derive dl by de
- 38:38now here's the problem
- 38:40how do we derive dl
- 38:41by dc
- 38:44we actually know the derivative l with
- 38:46respect to d so we know how l assessed
- 38:48it to d
- 38:50but how is l sensitive to c so if we
- 38:53wiggle c how does that impact l
- 38:55through d
- 38:58so we know dl by dc
- 39:01and we also here know how c impacts d
- 39:04and so just very intuitively if you know
- 39:06the impact that c is having on d and the
- 39:09impact that d is having on l
- 39:11then you should be able to somehow put
- 39:12that information together to figure out
- 39:14how c impacts l
- 39:16and indeed this is what we can actually
- 39:18do so in particular we know just
- 39:20concentrating on d first let's look at
- 39:22how what is the derivative basically of
- 39:24d with respect to c so in other words
- 39:27what is dd by dc
- 39:31so here we know that d is c times c plus
- 39:34e
- 39:35that's what we know and now we're
- 39:37interested in dd by dc
- 39:39if you just know your calculus again and
- 39:41you remember that differentiating c plus
- 39:43e with respect to c you know that that
- 39:45gives you
- 39:461.0
- 39:47and we can also go back to the basics
- 39:49and derive this because again we can go
- 39:51to our f of x plus h minus f of x
- 39:54divide by h
- 39:56that's the definition of a derivative as
- 39:58h goes to zero
- 40:00and so here
- 40:01focusing on c and its effect on d
- 40:04we can basically do the f of x plus h
- 40:06will be
- 40:07c is incremented by h plus e
- 40:10that's the first evaluation of our
- 40:12function minus
- 40:14c plus e
- 40:16and then divide h
- 40:18and so what is this
- 40:19uh just expanding this out this will be
- 40:21c plus h plus e minus c minus e
- 40:25divide h and then you see here how c
- 40:27minus c cancels e minus e cancels we're
- 40:30left with h over h which is 1.0
- 40:33and so
- 40:35by symmetry also d d by d
- 40:38e
- 40:39will be 1.0 as well
- 40:42so basically the derivative of a sum
- 40:44expression is very simple and and this
- 40:46is the local derivative so i call this
- 40:49the local derivative because we have the
- 40:51final output value all the way at the
- 40:52end of this graph and we're now like a
- 40:54small node here
- 40:55and this is a little plus node
- 40:58and it the little plus node doesn't know
- 41:00anything about the rest of the graph
- 41:02that it's embedded in all it knows is
- 41:04that it did a plus it took a c and an e
- 41:07added them and created d
- 41:09and this plus note also knows the local
- 41:11influence of c on d or rather rather the
- 41:14derivative of d with respect to c and it
- 41:16also
- 41:17knows the derivative of d with respect
- 41:18to e but that's not what we want that's
- 41:21just a local derivative what we actually
- 41:23want is d l by d c and l could l is here
- 41:27just one step away but in a general case
- 41:30this little plus note is could be
- 41:32embedded in like a massive graph
- 41:34so
- 41:35again we know how l impacts d and now we
- 41:38know how c and e impact d how do we put
- 41:41that information together to write dl by
- 41:43dc and the answer of course is the chain
- 41:46rule in calculus
- 41:47and so um
- 41:50i pulled up a chain rule here from
- 41:51kapedia
- 41:52and
- 41:53i'm going to go through this very
- 41:54briefly so chain rule
- 41:57wikipedia sometimes can be very
- 41:58confusing and calculus can
- 42:00can be very confusing like this is the
- 42:02way i
- 42:03learned
- 42:05chain rule and it was very confusing
- 42:06like what is happening it's just
- 42:08complicated so i like this expression
- 42:10much better
- 42:12if a variable z depends on a variable y
- 42:15which itself depends on the variable x
- 42:18then z depends on x as well obviously
- 42:20through the intermediate variable y
- 42:22in this case the chain rule is expressed
- 42:24as
- 42:25if you want dz by dx
- 42:28then you take the dz by dy and you
- 42:30multiply it by d y by dx
- 42:33so the chain rule fundamentally is
- 42:34telling you
- 42:36how
- 42:37we chain these
- 42:39uh derivatives together
- 42:41correctly so to differentiate through a
- 42:44function composition
- 42:46we have to apply a multiplication
- 42:48of
- 42:49those derivatives
- 42:51so that's really what chain rule is
- 42:53telling us
- 42:54and there's a nice little intuitive
- 42:56explanation here which i also think is
- 42:58kind of cute the chain rule says that
- 42:59knowing the instantaneous rate of change
- 43:01of z with respect to y and y relative to
- 43:03x allows one to calculate the
- 43:04instantaneous rate of change of z
- 43:06relative to x
- 43:07as a product of those two rates of
- 43:09change
- 43:10simply the product of those two
- 43:12so here's a good one
- 43:14if a car travels twice as fast as
- 43:16bicycle and the bicycle is four times as
- 43:18fast as walking man
- 43:19then the car travels two times four
- 43:22eight times as fast as demand
- 43:25and so this makes it very clear that the
- 43:27correct thing to do sort of
- 43:29is to multiply
- 43:30so
- 43:31cars twice as fast as bicycle and
- 43:33bicycle is four times as fast as man
- 43:36so the car will be eight times as fast
- 43:38as the man and so we can take these
- 43:42intermediate rates of change if you will
- 43:44and multiply them together
- 43:46and that justifies the
- 43:48chain rule intuitively so have a look at
- 43:50chain rule about here really what it
- 43:52means for us is there's a very simple
- 43:54recipe for deriving what we want which
- 43:56is dl by dc
- 43:59and what we have so far
- 44:01is we know
- 44:03want
- 44:05and we know
- 44:07what is the
- 44:08impact of d on l so we know d l by
- 44:12d d the derivative of l with respect to
- 44:14d d we know that that's negative two
- 44:17and now because of this local
- 44:19reasoning that we've done here we know
- 44:21dd by d
- 44:23c
- 44:24so how does c impact d and in
- 44:27particular this is a plus node so the
- 44:29local derivative is simply 1.0 it's very
- 44:32simple
- 44:33and so
- 44:34the chain rule tells us that dl by dc
- 44:37going through this intermediate variable
- 44:40will just be simply d l by
- 44:44d
- 44:44times
- 44:49dd by dc
- 44:51that's chain rule
- 44:53so this is identical to what's happening
- 44:55here
- 44:56except
- 44:58z is rl
- 44:59y is our d and x is rc
- 45:03so we literally just have to multiply
- 45:05these
- 45:06and because
- 45:10these local derivatives like dd by dc
- 45:12are just one
- 45:14we basically just copy over dl by dd
- 45:17because this is just times one
- 45:19so what does it do so because dl by dd
- 45:22is negative two what is dl by dc
- 45:25well it's the local gradient 1.0 times
- 45:29dl by dd which is negative two
- 45:31so literally what a plus node does you
- 45:33can look at it that way is it literally
- 45:35just routes the gradient
- 45:37because the plus nodes local derivatives
- 45:39are just one and so in the chain rule
- 45:41one times
- 45:43dl by dd
- 45:45is um
- 45:47is uh is just dl by dd and so that
- 45:50derivative just gets routed to both c
- 45:53and to e in this case
- 45:55so basically um we have that that grad
- 45:59or let's start with c since that's the
- 46:01one we looked at
- 46:02is
- 46:03negative two times one
- 46:06negative two
- 46:08and in the same way by symmetry e that
- 46:11grad will be negative two that's the
- 46:13claim so we can set those
- 46:16we can redraw
- 46:19and you see how we just assign negative
- 46:20to negative two so this backpropagating
- 46:23signal which is carrying the information
- 46:25of like what is the derivative of l with
- 46:26respect to all the intermediate nodes
- 46:28we can imagine it almost like flowing
- 46:30backwards through the graph and a plus
- 46:32node will simply distribute the
- 46:34derivative to all the leaf nodes sorry
- 46:36to all the children nodes of it
- 46:39so this is the claim and now let's
- 46:40verify it so let me remove the plus h
- 46:43here from before
- 46:45and now instead what we're going to do
- 46:46is we're going to increment c so c dot
- 46:48data will be credited by h
- 46:50and when i run this we expect to see
- 46:52negative 2
- 46:54negative 2. and then of course for e
- 46:58so e dot data plus equals h and we
- 47:01expect to see negative 2.
- 47:03simple
- 47:07so those are the derivatives of these
- 47:09internal nodes
- 47:11and now we're going to recurse our way
- 47:13backwards again
- 47:15and we're again going to apply the chain
- 47:17rule so here we go our second
- 47:19application of chain rule and we will
- 47:20apply it all the way through the graph
- 47:22we just happen to only have one more
- 47:24node remaining
- 47:25we have that d l
- 47:27by d e
- 47:28as we have just calculated is negative
- 47:30two so we know that
- 47:32so we know the derivative of l with
- 47:33respect to e
- 47:36and now we want dl
- 47:39by
- 47:40da
- 47:41right
- 47:42and the chain rule is telling us that
- 47:44that's just dl by de
- 47:48negative 2
- 47:50times the local gradient so what is the
- 47:52local gradient basically d e
- 47:55by d a
- 47:56we have to look at that
- 48:00so i'm a little times node
- 48:02inside a massive graph
- 48:04and i only know that i did a times b and
- 48:06i produced an e
- 48:09so now what is d e by d a and d e by d b
- 48:12that's the only thing that i sort of
- 48:14know about that's my local gradient
- 48:17so
- 48:17because we have that e's a times b we're
- 48:20asking what is d e by d a
- 48:24and of course we just did that here we
- 48:26had a
- 48:27times so i'm not going to rederive it
- 48:30but if you want to differentiate this
- 48:32with respect to a you'll just get b
- 48:34right the value of b
- 48:36which in this case is negative 3.0
- 48:41so
- 48:41basically we have that dl by da
- 48:45well let me just do it right here we
- 48:47have that a dot grad and we are applying
- 48:49chain rule here
- 48:50is d l by d e which we see here is
- 48:54negative two
- 48:56times
- 48:57what is d e by d a
- 48:59it's the value of b which is negative 3.
- 49:04that's it
- 49:07and then we have b grad is again dl by
- 49:10de
- 49:11which is negative 2
- 49:13just the same way
- 49:14times
- 49:15what is d e by d
- 49:18um db
- 49:19is the value of a which is 2.2.0
- 49:23as the value of a
- 49:25so these are our claimed derivatives
- 49:28let's
- 49:30redraw
- 49:32and we see here that
- 49:33a dot grad turns out to be 6 because
- 49:36that is negative 2 times negative 3
- 49:38and b dot grad is negative 4
- 49:41times sorry is negative 2 times 2 which
- 49:43is negative 4.
- 49:45so those are our claims let's delete
- 49:47this and let's verify them
- 49:50we have
- 49:52a here a dot data plus equals h
- 49:57so the claim is that
- 49:59a dot grad is six
- 50:01let's verify
- 50:03six
- 50:04and we have beta data
- 50:07plus equals h
- 50:08so nudging b by h
- 50:11and looking at what happens
- 50:13we claim it's negative four
- 50:15and indeed it's negative four plus minus
- 50:17again float oddness
- 50:20um
- 50:21and uh
- 50:23that's it this
- 50:24that was the manual
- 50:26back propagation
- 50:28uh all the way from here to all the leaf
- 50:30nodes and we've done it piece by piece
- 50:33and really all we've done is as you saw
- 50:35we iterated through all the nodes one by
- 50:37one and locally applied the chain rule
- 50:39we always know what is the derivative of
- 50:41l with respect to this little output and
- 50:44then we look at how this output was
- 50:45produced this output was produced
- 50:47through some operation and we have the
- 50:49pointers to the children nodes of this
- 50:51operation
- 50:52and so in this little operation we know
- 50:54what the local derivatives are and we
- 50:56just multiply them onto the derivative
- 50:58always
- 50:59so we just go through and recursively
- 51:01multiply on the local derivatives and
- 51:04that's what back propagation is is just
- 51:05a recursive application of chain rule
- 51:08backwards through the computation graph
- 51:10let's see this power in action just very
- 51:12briefly what we're going to do is we're
- 51:14going to
- 51:15nudge our inputs to try to make l go up
- 51:19so in particular what we're doing is we
- 51:21want a.data we're going to change it
- 51:24and if we want l to go up that means we
- 51:26just have to go in the direction of the
- 51:27gradient so
- 51:29a
- 51:30should increase in the direction of
- 51:32gradient by like some small step amount
- 51:34this is the step size
- 51:36and we don't just want this for ba but
- 51:38also for b
- 51:41also for c
- 51:44also for f
- 51:46those are
- 51:47leaf nodes which we usually have control
- 51:49over
- 51:50and if we nudge in direction of the
- 51:52gradient we expect a positive influence
- 51:54on l
- 51:55so we expect l to go up
- 51:58positively
- 51:59so it should become less negative it
- 52:01should go up to say negative you know
- 52:03six or something like that
- 52:05uh it's hard to tell exactly and we'd
- 52:08have to rewrite the forward pass so let
- 52:09me just um
- 52:12do that here
- 52:13um
- 52:16this would be the forward pass f would
- 52:18be unchanged this is effectively the
- 52:20forward pass and now if we print l.data
- 52:24we expect because we nudged all the
- 52:27values all the inputs in the rational
- 52:28gradient we expected a less negative l
- 52:30we expect it to go up
- 52:32so maybe it's negative six or so let's
- 52:34see what happens
- 52:36okay negative seven
- 52:38and uh this is basically one step of an
- 52:41optimization that we'll end up running
- 52:43and really does gradient just give us
- 52:46some power because we know how to
- 52:47influence the final outcome and this
- 52:49will be extremely useful for training
- 52:50knowledge as well as you'll see
- 52:52so now i would like to do one more uh
- 52:55example of manual backpropagation using
- 52:58a bit more complex and uh useful example
- 53:02we are going to back propagate through a
- 53:04neuron
- 53:05so
- 53:07we want to eventually build up neural
- 53:08networks and in the simplest case these
- 53:10are multilateral perceptrons as they're
- 53:12called so this is a two layer neural net
- 53:15and it's got these hidden layers made up
- 53:17of neurons and these neurons are fully
- 53:18connected to each other
- 53:20now biologically neurons are very
- 53:21complicated devices but we have very
- 53:23simple mathematical models of them
- 53:26and so this is a very simple
- 53:27mathematical model of a neuron you have
- 53:29some inputs axis
- 53:31and then you have these synapses that
- 53:33have weights on them so
- 53:36the w's are weights
- 53:39and then
- 53:40the synapse interacts with the input to
- 53:42this neuron multiplicatively so what
- 53:44flows to the cell body
- 53:47of this neuron is w times x
- 53:49but there's multiple inputs so there's
- 53:51many w times x's flowing into the cell
- 53:53body
- 53:54the cell body then has also like some
- 53:56bias
- 53:57so this is kind of like the
- 53:59inert innate sort of trigger happiness
- 54:02of this neuron so this bias can make it
- 54:04a bit more trigger happy or a bit less
- 54:06trigger happy regardless of the input
- 54:08but basically we're taking all the w
- 54:10times x
- 54:11of all the inputs adding the bias and
- 54:13then we take it through an activation
- 54:15function
- 54:16and this activation function is usually
- 54:18some kind of a squashing function
- 54:20like a sigmoid or 10h or something like
- 54:22that so as an example
- 54:24we're going to use the 10h in this
- 54:26example
- 54:28numpy has a
- 54:29np.10h
- 54:31so
- 54:32we can call it on a range
- 54:34and we can plot it
- 54:36this is the 10h function and you see
- 54:38that the inputs as they come in
- 54:41get squashed on the y coordinate here so
- 54:44um
- 54:45right at zero we're going to get exactly
- 54:47zero and then as you go more positive in
- 54:49the input
- 54:50then you'll see that the function will
- 54:52only go up to one and then plateau out
- 54:55and so if you pass in very positive
- 54:57inputs we're gonna cap it smoothly at
- 55:00one and on the negative side we're gonna
- 55:02cap it smoothly to negative one
- 55:04so that's 10h
- 55:06and that's the squashing function or an
- 55:08activation function and what comes out
- 55:10of this neuron is just the activation
- 55:12function applied to the dot product of
- 55:14the weights and the
- 55:16inputs
- 55:18so let's
- 55:19write one out
- 55:21um
- 55:22i'm going to copy paste because
- 55:27i don't want to type too much
- 55:28but okay so here we have the inputs
- 55:31x1 x2 so this is a two-dimensional
- 55:33neuron so two inputs are going to come
- 55:34in
- 55:35these are thought out as the weights of
- 55:37this neuron
- 55:38weights w1 w2 and these weights again
- 55:41are the synaptic strengths for each
- 55:43input
- 55:45and this is the bias of the neuron
- 55:47b
- 55:49and now we want to do is according to
- 55:51this model we need to multiply x1 times
- 55:54w1
- 55:55and x2 times w2
- 55:57and then we need to add bias on top of
- 56:00it
- 56:01and it gets a little messy here but all
- 56:03we are trying to do is x1 w1 plus x2 w2
- 56:06plus b
- 56:07and these are multiply here
- 56:09except i'm doing it in small steps so
- 56:12that we actually have pointers to all
- 56:13these intermediate nodes so we have x1
- 56:15w1 variable x times x2 w2 variable and
- 56:19i'm also labeling them
- 56:21so n is now
- 56:23the cell body raw
- 56:25raw
- 56:26activation without
- 56:28the activation function for now
- 56:30and this should be enough to basically
- 56:32plot it so draw dot of n
- 56:37gives us x1 times w1 x2 times w2
- 56:41being added
- 56:43then the bias gets added on top of this
- 56:45and this n
- 56:47is this sum
- 56:49so we're now going to take it through an
- 56:50activation function
- 56:52and let's say we use the 10h
- 56:54so that we produce the output
- 56:56so what we'd like to do here is we'd
- 56:58like to do the output and i'll call it o
- 57:01is um
- 57:03n dot 10h
- 57:05okay but we haven't yet written the 10h
- 57:08now the reason that we need to implement
- 57:09another 10h function here is that
- 57:12tanh is a
- 57:14hyperbolic function and we've only so
- 57:16far implemented a plus and the times and
- 57:18you can't make a 10h out of just pluses
- 57:20and times
- 57:22you also need exponentiation so 10h is
- 57:25this kind of a formula here
- 57:27you can use either one of these and you
- 57:28see that there's exponentiation involved
- 57:30which we have not implemented yet for
- 57:32our low value node here so we're not
- 57:34going to be able to produce 10h yet and
- 57:36we have to go back up and implement
- 57:37something like it
- 57:39now one option here
- 57:42is we could actually implement um
- 57:44exponentiation
- 57:46right and we could return the x of a
- 57:49value instead of a 10h of a value
- 57:52because if we had x then we have
- 57:54everything else that we need so um
- 57:56because we know how to add and we know
- 57:58how to
- 58:00um
- 58:01we know how to add and we know how to
- 58:02multiply so we'd be able to create 10h
- 58:04if we knew how to x
- 58:06but for the purposes of this example i
- 58:08specifically wanted to
- 58:10show you
- 58:11that we don't necessarily need to have
- 58:13the most atomic pieces
- 58:15in
- 58:16um
- 58:16in this value object we can actually
- 58:19like create functions at arbitrary
- 58:23points of abstraction they can be
- 58:24complicated functions but they can be
- 58:26also very very simple functions like a
- 58:27plus and it's totally up to us the only
- 58:30thing that matters is that we know how
- 58:31to differentiate through any one
- 58:33function so we take some inputs and we
- 58:35make an output the only thing that
- 58:37matters it can be arbitrarily complex
- 58:38function as long as you know how to
- 58:41create the local derivative if you know
- 58:43the local derivative of how the inputs
- 58:44impact the output then that's all you
- 58:46need so we're going to cluster up
- 58:49all of this expression and we're not
- 58:51going to break it down to its atomic
- 58:52pieces we're just going to directly
- 58:54implement tanh
- 58:55so let's do that
- 58:57depth nh
- 58:59and then out will be a value
- 59:02of
- 59:03and we need this expression here so
- 59:05um
- 59:08let me actually
- 59:10copy paste
- 59:14let's grab n which is a cell.theta
- 59:17and then this
- 59:18i believe is the tan h
- 59:21math.x of
- 59:24two
- 59:25no n
- 59:27n minus one over
- 59:28two n plus one
- 59:30maybe i can call this x
- 59:33just so that it matches exactly
- 59:35okay and now
- 59:37this will be t
- 59:40and uh children of this node there's
- 59:42just one child
- 59:44and i'm wrapping it in a tuple so this
- 59:46is a tuple of one object just self
- 59:48and here the name of this operation will
- 59:50be 10h
- 59:52and we're going to return that
- 59:56okay
- 59:58so now valley should be implementing 10h
- 1:00:02and now we can scroll all the way down
- 1:00:03here
- 1:00:04and we can actually do n.10 h and that's
- 1:00:06going to return the tanhd
- 1:00:09output of n
- 1:00:11and now we should be able to draw it out
- 1:00:12of o not of n
- 1:00:14so let's see how that worked
- 1:00:18there we go
- 1:00:19n went through 10 h
- 1:00:21to produce this output
- 1:00:24so now tan h is a
- 1:00:26sort of
- 1:00:27our little micro grad supported node
- 1:00:30here as an operation
- 1:00:33and as long as we know the derivative of
- 1:00:3510h
- 1:00:36then we'll be able to back propagate
- 1:00:37through it now let's see this 10h in
- 1:00:39action currently it's not squashing too
- 1:00:41much because the input to it is pretty
- 1:00:43low so if the bias was increased to say
- 1:00:46eight
- 1:00:49then we'll see that what's flowing into
- 1:00:51the 10h now is
- 1:00:53two
- 1:00:54and 10h is squashing it to 0.96 so we're
- 1:00:57already hitting the tail of this 10h and
- 1:00:59it will sort of smoothly go up to 1 and
- 1:01:01then plateau out over there
- 1:01:03okay so now i'm going to do something
- 1:01:04slightly strange i'm going to change
- 1:01:06this bias from 8 to this number
- 1:01:096.88 etc
- 1:01:11and i'm going to do this for specific
- 1:01:13reasons because we're about to start
- 1:01:15back propagation
- 1:01:16and i want to make sure that our numbers
- 1:01:19come out nice they're not like very
- 1:01:21crazy numbers they're nice numbers that
- 1:01:22we can sort of understand in our head
- 1:01:24let me also add a pose label
- 1:01:26o is short for output here
- 1:01:30so that's zero
- 1:01:31okay so
- 1:01:320.88 flows into 10 h comes out 0.7 so on
- 1:01:36so now we're going to do back
- 1:01:37propagation and we're going to fill in
- 1:01:38all the gradients
- 1:01:40so what is the derivative o with respect
- 1:01:43to
- 1:01:44all the
- 1:01:45inputs here and of course in the typical
- 1:01:47neural network setting what we really
- 1:01:48care about the most is the derivative of
- 1:01:51these neurons on the weights
- 1:01:53specifically the w2 and w1 because those
- 1:01:56are the weights that we're going to be
- 1:01:57changing part of the optimization
- 1:01:59and the other thing that we have to
- 1:02:00remember is here we have only a single
- 1:02:02neuron but in the neural natives
- 1:02:03typically have many neurons and they're
- 1:02:04connected
- 1:02:07so this is only like a one small neuron
- 1:02:09a piece of a much bigger puzzle and
- 1:02:10eventually there's a loss function that
- 1:02:12sort of measures the accuracy of the
- 1:02:13neural net and we're back propagating
- 1:02:15with respect to that accuracy and trying
- 1:02:16to increase it
- 1:02:19so let's start off by propagation here
- 1:02:21in the end
- 1:02:22what is the derivative of o with respect
- 1:02:24to o the base case sort of we know
- 1:02:26always is that the gradient is just 1.0
- 1:02:30so let me fill it in
- 1:02:32and then let me
- 1:02:35split out
- 1:02:37the drawing function
- 1:02:40here
- 1:02:43and then here cell
- 1:02:47clear this output here okay
- 1:02:50so now when we draw o we'll see that oh
- 1:02:52that grad is one
- 1:02:53so now we're going to back propagate
- 1:02:55through the tan h
- 1:02:56so to back propagate through 10h we need
- 1:02:58to know the local derivative of 10h
- 1:03:01so if we have that
- 1:03:03o is 10 h of
- 1:03:07n
- 1:03:08then what is d o by d n
- 1:03:12now what you could do is you could come
- 1:03:13here and you could take this expression
- 1:03:15and you could
- 1:03:16do your calculus derivative taking
- 1:03:19um and that would work but we can also
- 1:03:21just scroll down wikipedia here
- 1:03:23into a section that hopefully tells us
- 1:03:26that derivative uh
- 1:03:28d by dx of 10 h of x is
- 1:03:31any of these i like this one 1 minus 10
- 1:03:33h square of x
- 1:03:35so this is 1 minus 10 h
- 1:03:37of x squared
- 1:03:39so basically what this is saying is that
- 1:03:41d o by d n
- 1:03:43is
- 1:03:441 minus 10 h
- 1:03:47of n
- 1:03:48squared
- 1:03:51and we already have 10 h of n that's
- 1:03:52just o
- 1:03:54so it's one minus o squared
- 1:03:56so o is the output here so the output is
- 1:03:59this number
- 1:04:02data
- 1:04:04is this number
- 1:04:06and then
- 1:04:08what this is saying is that do by dn is
- 1:04:101 minus
- 1:04:11this squared so
- 1:04:13one minus of that data squared
- 1:04:16is 0.5 conveniently
- 1:04:18so the local derivative of this 10 h
- 1:04:21operation here is 0.5
- 1:04:24and
- 1:04:25so that would be d o by d n
- 1:04:27so
- 1:04:28we can fill in that in that grad
- 1:04:33is 0.5 we'll just fill in
- 1:04:42so this is exactly 0.5 one half
- 1:04:45so now we're going to continue the back
- 1:04:47propagation
- 1:04:49this is 0.5 and this is a plus node
- 1:04:52so how is backprop going to what is that
- 1:04:55going to do here
- 1:04:56and if you remember our previous example
- 1:04:58a plus is just a distributor of gradient
- 1:05:01so this gradient will simply flow to
- 1:05:03both of these equally and that's because
- 1:05:05the local derivative of this operation
- 1:05:07is one for every one of its nodes so 1
- 1:05:10times 0.5 is 0.5
- 1:05:12so therefore we know that
- 1:05:14this node here which we called this
- 1:05:18its grad is just 0.5
- 1:05:21and we know that b dot grad is also 0.5
- 1:05:24so let's set those and let's draw
- 1:05:28so 0.5
- 1:05:30continuing we have another plus
- 1:05:320.5 again we'll just distribute it so
- 1:05:340.5 will flow to both of these
- 1:05:37so we can set
- 1:05:39theirs
- 1:05:43x2w2 as well that grad is 0.5
- 1:05:47and let's redraw pluses are my favorite
- 1:05:50uh operations to back propagate through
- 1:05:51because
- 1:05:53it's very simple
- 1:05:55so now it's flowing into these
- 1:05:56expressions is 0.5 and so really again
- 1:05:58keep in mind what the derivative is
- 1:05:59telling us at every point in time along
- 1:06:01here this is saying that
- 1:06:04if we want the output of this neuron to
- 1:06:06increase
- 1:06:08then
- 1:06:08the influence on these expressions is
- 1:06:10positive on the output both of them are
- 1:06:13positive
- 1:06:16contribution to the output
- 1:06:20so now back propagating to x2 and w2
- 1:06:23first
- 1:06:24this is a times node so we know that the
- 1:06:26local derivative is you know the other
- 1:06:28term
- 1:06:28so if we want to calculate x2.grad
- 1:06:32then
- 1:06:33can you think through what it's going to
- 1:06:34be
- 1:06:40so x2.grad will be
- 1:06:42w2.data
- 1:06:44times this x2w2
- 1:06:48by grad right
- 1:06:51and
- 1:06:52w2.grad will be
- 1:06:55x2 that data times x2w2.grad
- 1:07:01right so that's the local piece of chain
- 1:07:03rule
- 1:07:07let's set them and let's redraw
- 1:07:09so here we see that the gradient on our
- 1:07:11weight 2 is 0 because x2 data was 0
- 1:07:15right but x2 will have the gradient 0.5
- 1:07:18because data here was 1.
- 1:07:20and so what's interesting here right is
- 1:07:22because the input x2 was 0 then because
- 1:07:25of the way the times works
- 1:07:28of course this gradient will be zero and
- 1:07:30think about intuitively why that is
- 1:07:33derivative always tells us the influence
- 1:07:35of
- 1:07:36this on the final output if i wiggle w2
- 1:07:39how is the output changing
- 1:07:41it's not changing because we're
- 1:07:42multiplying by zero
- 1:07:44so because it's not changing there's no
- 1:07:46derivative and zero is the correct
- 1:07:47answer
- 1:07:48because we're
- 1:07:49squashing it at zero
- 1:07:52and let's do it here point five should
- 1:07:54come here and flow through this times
- 1:07:57and so we'll have that x1.grad is
- 1:08:01can you think through a little bit what
- 1:08:03what
- 1:08:04this should be
- 1:08:07the local derivative of times
- 1:08:09with respect to x1 is going to be w1
- 1:08:12so w1 is data times
- 1:08:15x1 w1 dot grad
- 1:08:18and w1.grad will be x1.data times
- 1:08:23x1 w2 w1 with graph
- 1:08:27let's see what those came out to be
- 1:08:29so this is 0.5 so this would be negative
- 1:08:311.5 and this would be 1.
- 1:08:34and we've back propagated through this
- 1:08:36expression these are the actual final
- 1:08:38derivatives so if we want this neuron's
- 1:08:40output to increase
- 1:08:43we know that what's necessary is that
- 1:08:47w2 we have no gradient w2 doesn't
- 1:08:49actually matter to this neuron right now
- 1:08:51but this neuron this weight should uh go
- 1:08:54up
- 1:08:55so if this weight goes up then this
- 1:08:57neuron's output would have gone up and
- 1:08:59proportionally because the gradient is
- 1:09:01one okay so doing the back propagation
- 1:09:03manually is obviously ridiculous so we
- 1:09:05are now going to put an end to this
- 1:09:06suffering and we're going to see how we
- 1:09:08can implement uh the backward pass a bit
- 1:09:11more automatically we're not going to be
- 1:09:12doing all of it manually out here
- 1:09:14it's now pretty obvious to us by example
- 1:09:17how these pluses and times are back
- 1:09:18property ingredients so let's go up to
- 1:09:20the value
- 1:09:22object and we're going to start
- 1:09:24codifying what we've seen
- 1:09:27in the examples below
- 1:09:29so we're going to do this by storing a
- 1:09:31special cell dot backward
- 1:09:34and underscore backward and this will be
- 1:09:37a function which is going to do that
- 1:09:39little piece of chain rule at each
- 1:09:41little node that compute that took
- 1:09:43inputs and produced output uh we're
- 1:09:45going to store
- 1:09:46how we are going to chain the the
- 1:09:49outputs gradient into the inputs
- 1:09:51gradients
- 1:09:52so by default
- 1:09:54this will be a function
- 1:09:55that uh doesn't do anything
- 1:09:58so um
- 1:09:59and you can also see that here in the
- 1:10:01value in micrograb
- 1:10:03so
- 1:10:04with this backward function by default
- 1:10:06doesn't do anything
- 1:10:08this is an empty function
- 1:10:10and that would be sort of the case for
- 1:10:11example for a leaf node for leaf node
- 1:10:13there's nothing to do
- 1:10:15but now if when we're creating these out
- 1:10:18values these out values are an addition
- 1:10:21of self and other
- 1:10:24and so we will want to sell set
- 1:10:27outs backward to be
- 1:10:29the function that propagates the
- 1:10:31gradient
- 1:10:34so
- 1:10:35let's define what should happen
- 1:10:40and we're going to store it in a closure
- 1:10:42let's define what should happen when we
- 1:10:44call
- 1:10:45outs grad
- 1:10:47for in addition
- 1:10:50our job is to take
- 1:10:52outs grad and propagate it into self's
- 1:10:55grad and other grad so basically we want
- 1:10:57to sell self.grad to something
- 1:11:00and we want to set others.grad to
- 1:11:02something
- 1:11:04okay
- 1:11:05and the way we saw below how chain rule
- 1:11:08works we want to take the local
- 1:11:10derivative times
- 1:11:11the
- 1:11:12sort of global derivative i should call
- 1:11:14it which is the derivative of the final
- 1:11:16output of the expression with respect to
- 1:11:18out's data
- 1:11:21with respect to out
- 1:11:22so
- 1:11:24the local derivative of self in an
- 1:11:27addition is 1.0
- 1:11:29so it's just 1.0 times
- 1:11:31outs grad
- 1:11:34that's the chain rule
- 1:11:35and others.grad will be 1.0 times
- 1:11:38outgrad
- 1:11:39and what you basically what you're
- 1:11:40seeing here is that outscrad
- 1:11:42will simply be copied onto selfs grad
- 1:11:45and others grad as we saw happens for an
- 1:11:48addition operation
- 1:11:49so we're going to later call this
- 1:11:51function to propagate the gradient
- 1:11:53having done an addition
- 1:11:55let's now do multiplication we're going
- 1:11:57to also define that backward
- 1:12:02and we're going to set its backward to
- 1:12:04be backward
- 1:12:07and we want to chain outgrad into
- 1:12:11self.grad
- 1:12:14and others.grad
- 1:12:17and this will be a little piece of chain
- 1:12:18rule for multiplication
- 1:12:20so we'll have
- 1:12:21so what should this be
- 1:12:23can you think through
- 1:12:28so what is the local derivative
- 1:12:30here the local derivative was
- 1:12:32others.data
- 1:12:35and then
- 1:12:36oops others.data and the times of that
- 1:12:39grad that's channel
- 1:12:42and here we have self.data times of that
- 1:12:44grad
- 1:12:45that's what we've been doing
- 1:12:49and finally here for 10 h
- 1:12:51left backward
- 1:12:54and then we want to set out backwards to
- 1:12:57be just backward
- 1:13:00and here we need to
- 1:13:02back propagate we have out that grad and
- 1:13:04we want to chain it into self.grad
- 1:13:09and salt.grad will be
- 1:13:11the local derivative of this operation
- 1:13:13that we've done here which is 10h
- 1:13:16and so we saw that the local the
- 1:13:17gradient is 1 minus the tan h of x
- 1:13:20squared which here is t
- 1:13:23that's the local derivative because
- 1:13:25that's t is the output of this 10 h so 1
- 1:13:27minus t squared is the local derivative
- 1:13:30and then gradient um
- 1:13:32has to be multiplied because of the
- 1:13:33chain rule
- 1:13:34so outgrad is chained through the local
- 1:13:36gradient into salt.grad
- 1:13:39and that should be basically it so we're
- 1:13:41going to redefine our value node
- 1:13:44we're going to swing all the way down
- 1:13:46here
- 1:13:48and we're going to
- 1:13:49redefine
- 1:13:51our expression
- 1:13:52make sure that all the grads are zero
- 1:13:55okay
- 1:13:56but now we don't have to do this
- 1:13:57manually anymore
- 1:13:59we are going to basically be calling the
- 1:14:01dot backward in the right order
- 1:14:04so
- 1:14:05first we want to call os
- 1:14:07dot backwards
- 1:14:14so o was the outcome of 10h
- 1:14:17right so calling all that those who's
- 1:14:20backward
- 1:14:22will be
- 1:14:23this function this is what it will do
- 1:14:26now we have to be careful because
- 1:14:29there's a times out.grad
- 1:14:31and out.grad remember is initialized to
- 1:14:34zero
- 1:14:38so here we see grad zero so as a base
- 1:14:41case we need to set both.grad to 1.0
- 1:14:46to initialize this with 1
- 1:14:53and then once this is 1 we can call oda
- 1:14:56backward
- 1:14:57and what that should do is it should
- 1:14:58propagate this grad through 10h
- 1:15:02so the local derivative times
- 1:15:04the global derivative which is
- 1:15:05initialized at one so
- 1:15:08this should
- 1:15:11um
- 1:15:15a dope
- 1:15:17so i thought about redoing it but i
- 1:15:19figured i should just leave the error in
- 1:15:20here because it's pretty funny why is
- 1:15:22anti-object not callable
- 1:15:24uh it's because
- 1:15:27i screwed up we're trying to save these
- 1:15:29functions so this is correct
- 1:15:31this here
- 1:15:33we don't want to call the function
- 1:15:34because that returns none these
- 1:15:36functions return none we just want to
- 1:15:38store the function
- 1:15:39so let me redefine the value object
- 1:15:42and then we're going to come back in
- 1:15:43redefine the expression draw a dot
- 1:15:46everything is great o dot grad is one
- 1:15:50o dot grad is one and now
- 1:15:53now this should work of course
- 1:15:55okay so all that backward should
- 1:15:58this grant should now be 0.5 if we
- 1:16:00redraw and if everything went correctly
- 1:16:030.5 yay
- 1:16:05okay so now we need to call ns.grad
- 1:16:10and it's not awkward sorry
- 1:16:13ends backward
- 1:16:14so that seems to have worked
- 1:16:17so instead backward routed the gradient
- 1:16:21to both of these so this is looking
- 1:16:22great
- 1:16:24now we could of course called uh called
- 1:16:26b grad
- 1:16:27beat up backwards sorry
- 1:16:30what's gonna happen
- 1:16:32well b doesn't have it backward b is
- 1:16:34backward
- 1:16:35because b is a leaf node
- 1:16:37b's backward is by initialization the
- 1:16:40empty function
- 1:16:41so nothing would happen but we can call
- 1:16:44call it on it
- 1:16:45but when we call
- 1:16:48this one
- 1:16:50it's backward
- 1:16:53then we expect this 0.5 to get further
- 1:16:56routed
- 1:16:57right so there we go 0.5.5
- 1:17:00and then finally
- 1:17:02we want to call
- 1:17:05it here on x2 w2
- 1:17:10and on x1 w1
- 1:17:16do both of those
- 1:17:17and there we go
- 1:17:19so we get 0 0.5 negative 1.5 and 1
- 1:17:23exactly as we did before but now
- 1:17:26we've done it through
- 1:17:28calling that backward um
- 1:17:30sort of manually
- 1:17:32so we have the lamp one last piece to
- 1:17:34get rid of which is us calling
- 1:17:36underscore backward manually so let's
- 1:17:38think through what we are actually doing
- 1:17:40um
- 1:17:41we've laid out a mathematical expression
- 1:17:43and now we're trying to go backwards
- 1:17:44through that expression
- 1:17:46um so going backwards through the
- 1:17:48expression just means that we never want
- 1:17:50to call a dot backward for any node
- 1:17:54before
- 1:17:55we've done a sort of um everything after
- 1:17:58it
- 1:17:59so we have to do everything after it
- 1:18:01before we're ever going to call that
- 1:18:02backward on any one node we have to get
- 1:18:04all of its full dependencies everything
- 1:18:06that it depends on has to
- 1:18:08propagate to it before we can continue
- 1:18:10back propagation so this ordering of
- 1:18:14graphs can be achieved using something
- 1:18:16called topological sort
- 1:18:17so topological sort
- 1:18:20is basically a laying out of a graph
- 1:18:23such that all the edges go only from
- 1:18:24left to right basically
- 1:18:26so here we have a graph it's a directory
- 1:18:29a cyclic graph a dag
- 1:18:31and this is two different topological
- 1:18:34orders of it i believe where basically
- 1:18:36you'll see that it's laying out of the
- 1:18:37notes such that all the edges go only
- 1:18:39one way from left to right
- 1:18:41and implementing topological sort you
- 1:18:44can look in wikipedia and so on i'm not
- 1:18:46going to go through it in detail
- 1:18:48but basically this is what builds a
- 1:18:51topological graph
- 1:18:54we maintain a set of visited nodes and
- 1:18:56then we are
- 1:18:59going through starting at some root node
- 1:19:02which for us is o that's where we want
- 1:19:03to start the topological sort
- 1:19:05and starting at o we go through all of
- 1:19:08its children and we need to lay them out
- 1:19:10from left to right
- 1:19:12and basically this starts at o
- 1:19:14if it's not visited then it marks it as
- 1:19:17visited and then it iterates through all
- 1:19:19of its children
- 1:19:20and calls build topological on them
- 1:19:24and then uh after it's gone through all
- 1:19:26the children it adds itself
- 1:19:28so basically
- 1:19:29this node that we're going to call it on
- 1:19:31like say o is only going to add itself
- 1:19:34to the topo list after all of the
- 1:19:37children have been processed and that's
- 1:19:39how this function is guaranteeing
- 1:19:41that you're only going to be in the list
- 1:19:43once all your children are in the list
- 1:19:45and that's the invariant that is being
- 1:19:46maintained so if we built upon o and
- 1:19:49then inspect this list
- 1:19:52we're going to see that it ordered our
- 1:19:54value objects
- 1:19:56and the last one
- 1:19:58is the value of 0.707 which is the
- 1:20:00output
- 1:20:01so this is o and then this is n
- 1:20:04and then all the other nodes get laid
- 1:20:07out before it
- 1:20:09so that builds the topological graph and
- 1:20:12really what we're doing now is we're
- 1:20:13just calling dot underscore backward on
- 1:20:16all of the nodes in a topological order
- 1:20:19so if we just reset the gradients
- 1:20:22they're all zero
- 1:20:23what did we do
- 1:20:24we started by
- 1:20:27setting o dot grad
- 1:20:29to b1
- 1:20:31that's the base case
- 1:20:33then we built the topological order
- 1:20:38and then we went for node
- 1:20:41in
- 1:20:42reversed
- 1:20:44of topo
- 1:20:46now
- 1:20:47in in the reverse order because this
- 1:20:49list goes from
- 1:20:50you know we need to go through it in
- 1:20:52reversed order
- 1:20:53so starting at o
- 1:20:56note that backward
- 1:20:58and this should be
- 1:21:01it
- 1:21:03there we go
- 1:21:05those are the correct derivatives
- 1:21:07finally we are going to hide this
- 1:21:08functionality
- 1:21:10so i'm going to
- 1:21:11copy this and we're going to hide it
- 1:21:13inside the valley class because we don't
- 1:21:15want to have all that code lying around
- 1:21:18so instead of an underscore backward
- 1:21:19we're now going to define an actual
- 1:21:21backward so that's backward without the
- 1:21:23underscore
- 1:21:26and that's going to do all the stuff
- 1:21:27that we just arrived
- 1:21:29so let me just clean this up a little
- 1:21:30bit so
- 1:21:32we're first going to
- 1:21:37build a topological graph
- 1:21:38starting at self
- 1:21:41so build topo of self
- 1:21:44will populate the topological order into
- 1:21:46the topo list which is a local variable
- 1:21:49then we set self.grad to be one
- 1:21:52and then for each node in the reversed
- 1:21:55list so starting at us and going to all
- 1:21:57the children
- 1:22:00underscore backward
- 1:22:02and
- 1:22:03that should be it so
- 1:22:06save
- 1:22:08come down here
- 1:22:09redefine
- 1:22:09[Music]
- 1:22:11okay all the grands are zero
- 1:22:13and now what we can do is oh that
- 1:22:15backward without the underscore
- 1:22:17and
- 1:22:21there we go
- 1:22:22and that's uh that's back propagation
- 1:22:26place for one neuron
- 1:22:28now we shouldn't be too happy with
- 1:22:29ourselves actually because we have a bad
- 1:22:32bug um and we have not surfaced the bug
- 1:22:35because of some specific conditions that
- 1:22:36we are we have to think about right now
- 1:22:39so here's the simplest case that shows
- 1:22:42the bug
- 1:22:43say i create a single node a
- 1:22:48and then i create a b that is a plus a
- 1:22:51and then i called backward
- 1:22:54so what's going to happen is a is 3
- 1:22:57and then a b is a plus a so there's two
- 1:23:00arrows on top of each other here
- 1:23:03then we can see that b is of course the
- 1:23:05forward pass works
- 1:23:06b is just
- 1:23:08a plus a which is six
- 1:23:10but the gradient here is not actually
- 1:23:11correct
- 1:23:12that we calculate it automatically
- 1:23:15and that's because
- 1:23:17um
- 1:23:19of course uh
- 1:23:20just doing calculus in your head the
- 1:23:22derivative of b with respect to a
- 1:23:24should be uh two
- 1:23:27one plus one
- 1:23:28it's not one
- 1:23:30intuitively what's happening here right
- 1:23:32so b is the result of a plus a and then
- 1:23:34we call backward on it
- 1:23:36so let's go up and see what that does
- 1:23:42um
- 1:23:43b is a result of addition
- 1:23:45so out as
- 1:23:46b and then when we called backward what
- 1:23:49happened is
- 1:23:50self.grad was set
- 1:23:53to one
- 1:23:54and then other that grad was set to one
- 1:23:57but because we're doing a plus a
- 1:23:59self and other are actually the exact
- 1:24:02same object
- 1:24:03so we are overriding the gradient we are
- 1:24:06setting it to one and then we are
- 1:24:07setting it again to one and that's why
- 1:24:10it stays
- 1:24:11at one
- 1:24:13so that's a problem
- 1:24:14there's another way to see this in a
- 1:24:16little bit more complicated expression
- 1:24:21so here we have
- 1:24:23a and b
- 1:24:25and then uh d will be the multiplication
- 1:24:28of the two and e will be the addition of
- 1:24:30the two
- 1:24:32and
- 1:24:33then we multiply e times d to get f and
- 1:24:35then we called fda backward
- 1:24:37and these gradients if you check will be
- 1:24:39incorrect
- 1:24:40so fundamentally what's happening here
- 1:24:42again is
- 1:24:45basically we're going to see an issue
- 1:24:46anytime we use a variable more than once
- 1:24:49until now in these expressions above
- 1:24:51every variable is used exactly once so
- 1:24:53we didn't see the issue
- 1:24:54but here if a variable is used more than
- 1:24:56once what's going to happen during
- 1:24:57backward pass we're backpropagating from
- 1:25:00f to e to d so far so good but now
- 1:25:03equals it backward and it deposits its
- 1:25:05gradients to a and b but then we come
- 1:25:08back to d
- 1:25:09and call backward and it overwrites
- 1:25:11those gradients at a and b
- 1:25:14so that's obviously a problem
- 1:25:17and the solution here if you look at
- 1:25:19the multivariate case of the chain rule
- 1:25:22and its generalization there
- 1:25:23the solution there is basically that we
- 1:25:26have to accumulate these gradients these
- 1:25:28gradients add
- 1:25:30and so instead of setting those
- 1:25:32gradients
- 1:25:34we can simply do plus equals we need to
- 1:25:37accumulate those gradients
- 1:25:39plus equals plus equals
- 1:25:41plus equals
- 1:25:44plus equals
- 1:25:46and this will be okay remember because
- 1:25:48we are initializing them at zero so they
- 1:25:50start at zero
- 1:25:51and then any
- 1:25:53contribution
- 1:25:54that flows backwards
- 1:25:57will simply add
- 1:25:58so now if we redefine
- 1:26:01this one
- 1:26:03because the plus equals this now works
- 1:26:06because a.grad started at zero and we
- 1:26:08called beta backward we deposit one and
- 1:26:11then we deposit one again and now this
- 1:26:13is two which is correct
- 1:26:14and here this will also work and we'll
- 1:26:16get correct gradients
- 1:26:18because when we call eta backward we
- 1:26:20will deposit the gradients from this
- 1:26:21branch and then we get to back into
- 1:26:23detail backward it will deposit its own
- 1:26:26gradients and then those gradients
- 1:26:28simply add on top of each other and so
- 1:26:30we just accumulate those gradients and
- 1:26:31that fixes the issue okay now before we
- 1:26:34move on let me actually do a bit of
- 1:26:35cleanup here and delete some of these
- 1:26:38some of this intermediate work so
- 1:26:41we're not gonna need any of this now
- 1:26:42that we've derived all of it
- 1:26:44um
- 1:26:45we are going to keep this because i want
- 1:26:48to come back to it
- 1:26:49delete the 10h
- 1:26:51delete our morning example
- 1:26:53delete the step
- 1:26:55delete this keep the code that draws
- 1:26:59and then delete this example
- 1:27:02and leave behind only the definition of
- 1:27:03value
- 1:27:05and now let's come back to this
- 1:27:06non-linearity here that we implemented
- 1:27:08the tanh now i told you that we could
- 1:27:10have broken down 10h into its explicit
- 1:27:13atoms in terms of other expressions if
- 1:27:16we had the x function so if you remember
- 1:27:18tan h is defined like this and we chose
- 1:27:20to develop tan h as a single function
- 1:27:22and we can do that because we know its
- 1:27:24derivative and we can back propagate
- 1:27:26through it
- 1:27:26but we can also break down tan h into
- 1:27:29and express it as a function of x and i
- 1:27:31would like to do that now because i want
- 1:27:33to prove to you that you get all the
- 1:27:34same results and all those ingredients
- 1:27:36but also because it forces us to
- 1:27:38implement a few more expressions it
- 1:27:40forces us to do exponentiation addition
- 1:27:42subtraction division and things like
- 1:27:44that and i think it's a good exercise to
- 1:27:46go through a few more of these
- 1:27:48okay so let's scroll up
- 1:27:50to the definition of value
- 1:27:52and here one thing that we currently
- 1:27:53can't do is we can do like a value of
- 1:27:56say 2.0
- 1:27:58but we can't do you know here for
- 1:28:00example we want to add constant one and
- 1:28:02we can't do something like this
- 1:28:05and we can't do it because it says
- 1:28:06object has no attribute data that's
- 1:28:08because a plus one comes right here to
- 1:28:11add
- 1:28:12and then other is the integer one and
- 1:28:14then here python is trying to access
- 1:28:16one.data and that's not a thing and
- 1:28:18that's because basically one is not a
- 1:28:20value object and we only have addition
- 1:28:22for value objects so as a matter of
- 1:28:24convenience so that we can create
- 1:28:26expressions like this and make them make
- 1:28:28sense
- 1:28:29we can simply do something like this
- 1:28:32basically
- 1:28:33we let other alone if other is an
- 1:28:35instance of value but if it's not an
- 1:28:37instance of value we're going to assume
- 1:28:39that it's a number like an integer float
- 1:28:40and we're going to simply wrap it in in
- 1:28:43value and then other will just become
- 1:28:45value of other and then other will have
- 1:28:46a data attribute and this should work so
- 1:28:49if i just say this predefined value then
- 1:28:51this should work
- 1:28:53there we go okay now let's do the exact
- 1:28:55same thing for multiply because we can't
- 1:28:57do something like this
- 1:28:58again
- 1:28:59for the exact same reason so we just
- 1:29:01have to go to mole and if other is
- 1:29:04not a value then let's wrap it in value
- 1:29:07let's redefine value and now this works
- 1:29:10now here's a kind of unfortunate and not
- 1:29:12obvious part a times two works we saw
- 1:29:15that but two times a is that gonna work
- 1:29:19you'd expect it to right but actually it
- 1:29:21will not
- 1:29:22and the reason it won't is because
- 1:29:24python doesn't know
- 1:29:26like when when you do a times two
- 1:29:28basically um so a times two python will
- 1:29:31go and it will basically do something
- 1:29:32like a dot mul
- 1:29:34of two that's basically what it will
- 1:29:36call but to it 2 times a is the same as
- 1:29:392 dot mol of a
- 1:29:41and it doesn't 2 can't multiply
- 1:29:44value and so it's really confused about
- 1:29:46that
- 1:29:47so instead what happens is in python the
- 1:29:49way this works is you are free to define
- 1:29:51something called the r mold
- 1:29:54and our mole
- 1:29:55is kind of like a fallback so if python
- 1:29:58can't do 2 times a it will check if um
- 1:30:02if by any chance a knows how to multiply
- 1:30:05two and that will be called into our
- 1:30:07mole
- 1:30:08so because python can't do two times a
- 1:30:11it will check is there an our mole in
- 1:30:12value and because there is it will now
- 1:30:15call that
- 1:30:16and what we'll do here is we will swap
- 1:30:18the order of the operands so basically
- 1:30:21two times a will redirect to armel and
- 1:30:23our mole will basically call a times two
- 1:30:26and that's how that will work
- 1:30:28so
- 1:30:29redefining now with armor two times a
- 1:30:31becomes four okay now looking at the
- 1:30:33other elements that we still need we
- 1:30:34need to know how to exponentiate and how
- 1:30:36to divide so let's first the explanation
- 1:30:38to the exponentiation part we're going
- 1:30:40to introduce
- 1:30:41a single
- 1:30:42function x here
- 1:30:45and x is going to mirror 10h in the
- 1:30:47sense that it's a simple single function
- 1:30:49that transforms a single scalar value
- 1:30:51and outputs a single scalar value
- 1:30:53so we pop out the python number we use
- 1:30:56math.x to exponentiate it create a new
- 1:30:58value object
- 1:30:59everything that we've seen before the
- 1:31:00tricky part of course is how do you
- 1:31:02propagate through e to the x
- 1:31:04and
- 1:31:05so here you can potentially pause the
- 1:31:07video and think about what should go
- 1:31:09here
- 1:31:13okay so basically we need to know what
- 1:31:15is the local derivative of e to the x so
- 1:31:18d by d x of e to the x is famously just
- 1:31:21e to the x and we've already just
- 1:31:23calculated e to the x and it's inside
- 1:31:25out that data so we can do up that data
- 1:31:27times
- 1:31:28and
- 1:31:29out that grad that's the chain rule
- 1:31:32so we're just chaining on to the current
- 1:31:33running grad
- 1:31:35and this is what the expression looks
- 1:31:36like it looks a little confusing but
- 1:31:38this is what it is and that's the
- 1:31:40exponentiation
- 1:31:41so redefining we should now be able to
- 1:31:43call a.x
- 1:31:45and
- 1:31:46hopefully the backward pass works as
- 1:31:47well okay and the last thing we'd like
- 1:31:49to do of course is we'd like to be able
- 1:31:50to divide
- 1:31:52now
- 1:31:53i actually will implement something
- 1:31:54slightly more powerful than division
- 1:31:56because division is just a special case
- 1:31:57of something a bit more powerful
- 1:31:59so in particular just by rearranging
- 1:32:02if we have some kind of a b equals
- 1:32:04value of 4.0 here we'd like to basically
- 1:32:07be able to do a divide b and we'd like
- 1:32:09this to be able to give us 0.5
- 1:32:11now division actually can be reshuffled
- 1:32:14as follows if we have a divide b that's
- 1:32:17actually the same as a multiplying one
- 1:32:18over b
- 1:32:19and that's the same as a multiplying b
- 1:32:21to the power of negative one
- 1:32:24and so what i'd like to do instead is i
- 1:32:25basically like to implement the
- 1:32:27operation of x to the k for some
- 1:32:29constant uh k so it's an integer or a
- 1:32:32float um and we would like to be able to
- 1:32:35differentiate this and then as a special
- 1:32:36case uh negative one will be division
- 1:32:40and so i'm doing that just because uh
- 1:32:42it's more general and um yeah you might
- 1:32:45as well do it that way so basically what
- 1:32:46i'm saying is we can redefine
- 1:32:49uh division
- 1:32:51which we will put here somewhere
- 1:32:54yeah we can put it here somewhere what
- 1:32:56i'm saying is that we can redefine
- 1:32:58division so self-divide other
- 1:33:00can actually be rewritten as self times
- 1:33:03other to the power of negative one
- 1:33:05and now
- 1:33:07a value raised to the power of negative
- 1:33:09one we have now defined that
- 1:33:11so
- 1:33:12here's
- 1:33:13so we need to implement the pow function
- 1:33:15where am i going to put the power
- 1:33:17function maybe here somewhere
- 1:33:20this is the skeleton for it
- 1:33:22so this function will be called when we
- 1:33:24try to raise a value to some power and
- 1:33:26other will be that power
- 1:33:28now i'd like to make sure that other is
- 1:33:30only an int or a float usually other is
- 1:33:33some kind of a different value object
- 1:33:35but here other will be forced to be an
- 1:33:37end or a float otherwise the math
- 1:33:40won't work for
- 1:33:42for or try to achieve in the specific
- 1:33:43case that would be a different
- 1:33:45derivative expression if we wanted other
- 1:33:47to be a value
- 1:33:49so here we create the output value which
- 1:33:51is just uh you know this data raised to
- 1:33:53the power of other and other here could
- 1:33:55be for example negative one that's what
- 1:33:56we are hoping to achieve
- 1:33:59and then uh this is the backwards stub
- 1:34:01and this is the fun part which is what
- 1:34:03is the uh chain rule expression here for
- 1:34:07back for um
- 1:34:09back propagating through the power
- 1:34:11function where the power is to the power
- 1:34:13of some kind of a constant
- 1:34:15so this is the exercise and maybe pause
- 1:34:17the video here and see if you can figure
- 1:34:18it out yourself as to what we should put
- 1:34:20here
- 1:34:26okay so
- 1:34:29you can actually go here and look at
- 1:34:30derivative rules as an example and we
- 1:34:32see lots of derivatives that you can
- 1:34:34hopefully know from calculus in
- 1:34:36particular what we're looking for is the
- 1:34:37power rule
- 1:34:39because that's telling us that if we're
- 1:34:40trying to take d by dx of x to the n
- 1:34:42which is what we're doing here
- 1:34:44then that is just n times x to the n
- 1:34:46minus 1
- 1:34:48right
- 1:34:49okay
- 1:34:50so
- 1:34:51that's telling us about the local
- 1:34:53derivative of this power operation
- 1:34:55so all we want here
- 1:34:58basically n is now other
- 1:35:00and self.data is x
- 1:35:03and so this now becomes
- 1:35:06other which is n times
- 1:35:08self.data
- 1:35:10which is now a python in torah float
- 1:35:13it's not a valley object we're accessing
- 1:35:14the data attribute
- 1:35:16raised
- 1:35:17to the power of other minus one or n
- 1:35:19minus one
- 1:35:21i can put brackets around this but this
- 1:35:22doesn't matter because
- 1:35:25power takes precedence over multiply and
- 1:35:27python so that would have been okay
- 1:35:29and that's the local derivative only but
- 1:35:31now we have to chain it and we change
- 1:35:33just simply by multiplying by output
- 1:35:34grad that's chain rule
- 1:35:36and this should technically work
- 1:35:40and we're going to find out soon but now
- 1:35:42if we
- 1:35:43do this this should now work
- 1:35:46and we get 0.5 so the forward pass works
- 1:35:49but does the backward pass work and i
- 1:35:51realize that we actually also have to
- 1:35:52know how to subtract so
- 1:35:54right now a minus b will not work
- 1:35:57to make it work we need one more
- 1:36:00piece of code here
- 1:36:01and
- 1:36:02basically this is the
- 1:36:05subtraction and the way we're going to
- 1:36:06implement subtraction is we're going to
- 1:36:08implement it by addition of a negation
- 1:36:10and then to implement negation we're
- 1:36:12gonna multiply by negative one so just
- 1:36:14again using the stuff we've already
- 1:36:15built and just um expressing it in terms
- 1:36:17of what we have and a minus b is now
- 1:36:20working okay so now let's scroll again
- 1:36:22to this expression here for this neuron
- 1:36:25and let's just
- 1:36:26compute the backward pass here once
- 1:36:28we've defined o
- 1:36:30and let's draw it
- 1:36:32so here's the gradients for all these
- 1:36:33leaf nodes for this two-dimensional
- 1:36:35neuron that has a 10h that we've seen
- 1:36:37before so now what i'd like to do is i'd
- 1:36:39like to break up this 10h
- 1:36:41into this expression here
- 1:36:44so let me copy paste this
- 1:36:46here
- 1:36:47and now instead of we'll preserve the
- 1:36:49label
- 1:36:50and we will change how we define o
- 1:36:53so in particular we're going to
- 1:36:55implement this formula here
- 1:36:56so we need e to the 2x
- 1:36:58minus 1 over e to the x plus 1. so e to
- 1:37:01the 2x we need to take 2 times n and we
- 1:37:04need to exponentiate it that's e to the
- 1:37:07two x and then because we're using it
- 1:37:08twice let's create an intermediate
- 1:37:10variable e
- 1:37:12and then define o as
- 1:37:14e plus one over
- 1:37:16e minus one over e plus one
- 1:37:19e minus one over e plus one
- 1:37:22and that should be it and then we should
- 1:37:24be able to draw that of o
- 1:37:26so now before i run this what do we
- 1:37:29expect to see
- 1:37:30number one we're expecting to see a much
- 1:37:32longer
- 1:37:33graph here because we've broken up 10h
- 1:37:35into a bunch of other operations
- 1:37:37but those operations are mathematically
- 1:37:39equivalent and so what we're expecting
- 1:37:41to see is number one the same
- 1:37:43result here so the forward pass works
- 1:37:45and number two because of that
- 1:37:47mathematical equivalence we expect to
- 1:37:49see the same backward pass and the same
- 1:37:51gradients on these leaf nodes so these
- 1:37:53gradients should be identical
- 1:37:55so let's run this
- 1:37:58so number one let's verify that instead
- 1:38:00of a single 10h node we have now x and
- 1:38:03we have plus we have times negative one
- 1:38:06uh this is the division
- 1:38:08and we end up with the same forward pass
- 1:38:10here
- 1:38:11and then the gradients we have to be
- 1:38:13careful because they're in slightly
- 1:38:14different order potentially the
- 1:38:16gradients for w2x2 should be 0 and 0.5
- 1:38:19w2 and x2 are 0 and 0.5
- 1:38:22and w1 x1 are 1 and negative 1.5
- 1:38:251 and negative 1.5
- 1:38:27so that means that both our forward
- 1:38:28passes and backward passes were correct
- 1:38:31because this turned out to be equivalent
- 1:38:33to
- 1:38:3410h before
- 1:38:35and so the reason i wanted to go through
- 1:38:37this exercise is number one we got to
- 1:38:39practice a few more operations and uh
- 1:38:41writing more backwards passes and number
- 1:38:43two i wanted to illustrate the point
- 1:38:45that
- 1:38:46the um
- 1:38:47the level at which you implement your
- 1:38:49operations is totally up to you you can
- 1:38:51implement backward passes for tiny
- 1:38:53expressions like a single individual
- 1:38:54plus or a single times
- 1:38:56or you can implement them for say
- 1:38:5810h
- 1:39:00which is a kind of a potentially you can
- 1:39:01see it as a composite operation because
- 1:39:03it's made up of all these more atomic
- 1:39:05operations but really all of this is
- 1:39:07kind of like a fake concept all that
- 1:39:08matters is we have some kind of inputs
- 1:39:10and some kind of an output and this
- 1:39:11output is a function of the inputs in
- 1:39:13some way and as long as you can do
- 1:39:14forward pass and the backward pass of
- 1:39:16that little operation it doesn't matter
- 1:39:19what that operation is
- 1:39:21and how composite it is
- 1:39:23if you can write the local gradients you
- 1:39:24can chain the gradient and you can
- 1:39:26continue back propagation so the design
- 1:39:28of what those functions are is
- 1:39:30completely up to you
- 1:39:31so now i would like to show you how you
- 1:39:33can do the exact same thing by using a
- 1:39:35modern deep neural network library like
- 1:39:37for example pytorch which i've roughly
- 1:39:40modeled micrograd
- 1:39:41by
- 1:39:42and so
- 1:39:43pytorch is something you would use in
- 1:39:44production and i'll show you how you can
- 1:39:46do the exact same thing but in pytorch
- 1:39:48api so i'm just going to copy paste it
- 1:39:50in and walk you through it a little bit
- 1:39:52this is what it looks like
- 1:39:54so we're going to import pi torch and
- 1:39:56then we need to define these
- 1:39:59value objects like we have here
- 1:40:01now micrograd is a scalar valued
- 1:40:04engine so we only have scalar values
- 1:40:07like 2.0 but in pi torch everything is
- 1:40:10based around tensors and like i
- 1:40:11mentioned tensors are just n-dimensional
- 1:40:13arrays of scalars
- 1:40:15so that's why things get a little bit
- 1:40:17more complicated here i just need a
- 1:40:19scalar value to tensor a tensor with
- 1:40:21just a single element
- 1:40:23but by default when you work with
- 1:40:25pytorch you would use um
- 1:40:28more complicated tensors like this so if
- 1:40:30i import pytorch
- 1:40:33then i can create tensors like this and
- 1:40:36this tensor for example is a two by
- 1:40:38three array
- 1:40:39of scalar
- 1:40:41scalars
- 1:40:42in a single compact representation so we
- 1:40:45can check its shape we see that it's a
- 1:40:46two by three array
- 1:40:48and so on
- 1:40:49so this is usually what you would work
- 1:40:50with um in the actual libraries so here
- 1:40:54i'm creating
- 1:40:55a tensor that has only a single element
- 1:40:582.0
- 1:41:00and then i'm casting it to be double
- 1:41:03because python is by default using
- 1:41:05double precision for its floating point
- 1:41:07numbers so i'd like everything to be
- 1:41:08identical by default the data type of
- 1:41:12these tensors will be float32 so it's
- 1:41:14only using a single precision float so
- 1:41:16i'm casting it to double
- 1:41:18so that we have float64 just like in
- 1:41:21python
- 1:41:22so i'm casting to double and then we get
- 1:41:24something similar to value of two the
- 1:41:28next thing i have to do is because these
- 1:41:29are leaf nodes by default pytorch
- 1:41:31assumes that they do not require
- 1:41:32gradients so i need to explicitly say
- 1:41:35that all of these nodes require
- 1:41:36gradients
- 1:41:37okay so this is going to construct
- 1:41:39scalar valued one element tensors
- 1:41:43make sure that fighters knows that they
- 1:41:44require gradients now by default these
- 1:41:47are set to false by the way because of
- 1:41:48efficiency reasons because usually you
- 1:41:50would not want gradients for leaf nodes
- 1:41:53like the inputs to the network and this
- 1:41:55is just trying to be efficient in the
- 1:41:57most common cases
- 1:41:59so once we've defined all of our values
- 1:42:01in python we can perform arithmetic just
- 1:42:03like we can here in microgradlend so
- 1:42:06this will just work and then there's a
- 1:42:07torch.10h also
- 1:42:09and when we get back is a tensor again
- 1:42:12and we can
- 1:42:13just like in micrograd it's got a data
- 1:42:15attribute and it's got grant attributes
- 1:42:18so these tensor objects just like in
- 1:42:19micrograd have a dot data and a dot grad
- 1:42:22and
- 1:42:23the only difference here is that we need
- 1:42:25to call it that item because otherwise
- 1:42:28um pi torch
- 1:42:30that item basically takes
- 1:42:32a single tensor of one element and it
- 1:42:34just returns that element stripping out
- 1:42:36the tensor
- 1:42:37so let me just run this and hopefully we
- 1:42:39are going to get this is going to print
- 1:42:41the forward pass
- 1:42:42which is 0.707
- 1:42:44and this will be the gradients which
- 1:42:46hopefully are
- 1:42:480.5 0 negative 1.5 and 1.
- 1:42:51so if we just run this
- 1:42:53there we go
- 1:42:540.7 so the forward pass agrees and then
- 1:42:57point five zero negative one point five
- 1:42:59and one
- 1:43:00so pi torch agrees with us
- 1:43:02and just to show you here basically o
- 1:43:05here's a tensor with a single element
- 1:43:08and it's a double
- 1:43:09and we can call that item on it to just
- 1:43:12get the single number out
- 1:43:14so that's what item does and o is a
- 1:43:16tensor object like i mentioned and it's
- 1:43:18got a backward function just like we've
- 1:43:20implemented
- 1:43:22and then all of these also have a dot
- 1:43:23graph so like x2 for example in the grad
- 1:43:26and it's a tensor and we can pop out the
- 1:43:28individual number with that actin
- 1:43:31so basically
- 1:43:32torches torch can do what we did in
- 1:43:35micrograph is a special case when your
- 1:43:37tensors are all single element tensors
- 1:43:40but the big deal with pytorch is that
- 1:43:42everything is significantly more
- 1:43:43efficient because we are working with
- 1:43:45these tensor objects and we can do lots
- 1:43:47of operations in parallel on all of
- 1:43:49these tensors
- 1:43:51but otherwise what we've built very much
- 1:43:53agrees with the api of pytorch
- 1:43:55okay so now that we have some machinery
- 1:43:57to build out pretty complicated
- 1:43:58mathematical expressions we can also
- 1:44:00start building out neural nets and as i
- 1:44:02mentioned neural nets are just a
- 1:44:03specific class of mathematical
- 1:44:05expressions
- 1:44:07so we're going to start building out a
- 1:44:08neural net piece by piece and eventually
- 1:44:09we'll build out a two-layer multi-layer
- 1:44:12layer perceptron as it's called and i'll
- 1:44:14show you exactly what that means
- 1:44:15let's start with a single individual
- 1:44:17neuron we've implemented one here but
- 1:44:19here i'm going to implement one that
- 1:44:21also subscribes to the pytorch api in
- 1:44:24how it designs its neural network
- 1:44:26modules
- 1:44:27so just like we saw that we can like
- 1:44:28match the api of pytorch
- 1:44:31on the auto grad side we're going to try
- 1:44:33to do that on the neural network modules
- 1:44:35so here's class neuron
- 1:44:38and just for the sake of efficiency i'm
- 1:44:40going to copy paste some sections that
- 1:44:42are relatively straightforward
- 1:44:45so the constructor will take
- 1:44:47number of inputs to this neuron which is
- 1:44:49how many inputs come to a neuron so this
- 1:44:52one for example has three inputs
- 1:44:55and then it's going to create a weight
- 1:44:57there is some random number between
- 1:44:58negative one and one for every one of
- 1:45:00those inputs
- 1:45:01and a bias that controls the overall
- 1:45:03trigger happiness of this neuron
- 1:45:06and then we're going to implement a def
- 1:45:08underscore underscore call
- 1:45:11of self and x some input x
- 1:45:14and really what we don't do here is w
- 1:45:15times x plus b
- 1:45:17where w times x here is a dot product
- 1:45:19specifically
- 1:45:21now if you haven't seen
- 1:45:22call
- 1:45:24let me just return 0.0 here for now the
- 1:45:26way this works now is we can have an x
- 1:45:28which is say like 2.0 3.0 then we can
- 1:45:31initialize a neuron that is
- 1:45:32two-dimensional
- 1:45:33because these are two numbers and then
- 1:45:35we can feed those two numbers into that
- 1:45:37neuron to get an output
- 1:45:39and so when you use this notation n of x
- 1:45:42python will use call
- 1:45:45so currently call just return 0.0
- 1:45:50now we'd like to actually do the forward
- 1:45:52pass of this neuron instead
- 1:45:54so we're going to do here first is we
- 1:45:57need to basically multiply all of the
- 1:45:58elements of w with all of the elements
- 1:46:01of x pairwise we need to multiply them
- 1:46:04so the first thing we're going to do is
- 1:46:05we're going to zip up
- 1:46:07celta w and x
- 1:46:09and in python zip takes two iterators
- 1:46:12and it creates a new iterator that
- 1:46:14iterates over the tuples of the
- 1:46:16corresponding entries
- 1:46:17so for example just to show you we can
- 1:46:20print this list
- 1:46:22and still return 0.0 here
- 1:46:30sorry
- 1:46:34so we see that these w's are paired up
- 1:46:36with the x's w with x
- 1:46:41and now what we want to do is
- 1:46:47for w i x i in
- 1:46:50we want to multiply w times
- 1:46:52w wi times x i
- 1:46:54and then we want to sum all of that
- 1:46:56together
- 1:46:57to come up with an activation
- 1:46:59and add also subnet b on top
- 1:47:02so that's the raw activation and then of
- 1:47:04course we need to pass that through a
- 1:47:05non-linearity so what we're going to be
- 1:47:07returning is act.10h
- 1:47:09and here's out
- 1:47:12so
- 1:47:13now we see that we are getting some
- 1:47:14outputs and we get a different output
- 1:47:16from a neuron each time because we are
- 1:47:17initializing different weights and by
- 1:47:19and biases
- 1:47:21and then to be a bit more efficient here
- 1:47:22actually sum by the way takes a second
- 1:47:25optional parameter which is the start
- 1:47:28and by default the start is zero so
- 1:47:31these elements of this sum will be added
- 1:47:34on top of zero to begin with but
- 1:47:35actually we can just start with cell dot
- 1:47:37b
- 1:47:38and then we just have an expression like
- 1:47:39this
- 1:47:45and then the generator expression here
- 1:47:47must be parenthesized in python
- 1:47:49there we go
- 1:47:53yep so now we can forward a single
- 1:47:55neuron next up we're going to define a
- 1:47:57layer of neurons so here we have a
- 1:47:59schematic for a mlb
- 1:48:02so we see that these mlps each layer
- 1:48:05this is one layer has actually a number
- 1:48:07of neurons and they're not connected to
- 1:48:08each other but all of them are fully
- 1:48:09connected to the input
- 1:48:11so what is a layer of neurons it's just
- 1:48:13it's just a set of neurons evaluated
- 1:48:15independently
- 1:48:16so
- 1:48:17in the interest of time i'm going to do
- 1:48:20something fairly straightforward here
- 1:48:23it's um
- 1:48:25literally a layer is just a list of
- 1:48:27neurons
- 1:48:28and then how many neurons do we have we
- 1:48:30take that as an input argument here how
- 1:48:32many neurons do you want in your layer
- 1:48:34number of outputs in this layer
- 1:48:36and so we just initialize completely
- 1:48:38independent neurons with this given
- 1:48:40dimensionality and when we call on it we
- 1:48:43just independently
- 1:48:44evaluate them so now instead of a neuron
- 1:48:47we can make a layer of neurons they are
- 1:48:49two-dimensional neurons and let's have
- 1:48:51three of them
- 1:48:52and now we see that we have three
- 1:48:53independent evaluations of three
- 1:48:55different neurons
- 1:48:57right
- 1:48:58okay finally let's complete this picture
- 1:49:00and define an entire multi-layer
- 1:49:02perceptron or mlp
- 1:49:04and as we can see here in an mlp these
- 1:49:06layers just feed into each other
- 1:49:07sequentially
- 1:49:09so let's come here and i'm just going to
- 1:49:11copy the code here in interest of time
- 1:49:14so an mlp is very similar
- 1:49:16we're taking the number of inputs
- 1:49:18as before but now instead of taking a
- 1:49:20single n out which is number of neurons
- 1:49:22in a single layer we're going to take a
- 1:49:24list of an outs and this list defines
- 1:49:26the sizes of all the layers that we want
- 1:49:28in our mlp
- 1:49:30so here we just put them all together
- 1:49:31and then iterate over consecutive pairs
- 1:49:34of these sizes and create layer objects
- 1:49:36for them
- 1:49:37and then in the call function we are
- 1:49:39just calling them sequentially so that's
- 1:49:41an mlp really
- 1:49:42and let's actually re-implement this
- 1:49:44picture so we want three input neurons
- 1:49:46and then two layers of four and an
- 1:49:48output unit
- 1:49:49so
- 1:49:50we want
- 1:49:52a three-dimensional input say this is an
- 1:49:54example input we want three inputs into
- 1:49:57two layers of four and one output
- 1:50:00and this of course is an mlp
- 1:50:03and there we go that's a forward pass of
- 1:50:05an mlp
- 1:50:06to make this a little bit nicer you see
- 1:50:08how we have just a single element but
- 1:50:09it's wrapped in a list because layer
- 1:50:11always returns lists
- 1:50:13circle for convenience
- 1:50:15return outs at zero if len out is
- 1:50:18exactly a single element
- 1:50:20else return fullest
- 1:50:22and this will allow us to just get a
- 1:50:23single value out at the last layer that
- 1:50:25only has a single neuron
- 1:50:28and finally we should be able to draw
- 1:50:29dot of n of x
- 1:50:31and
- 1:50:32as you might imagine
- 1:50:34these expressions are now getting
- 1:50:36relatively involved
- 1:50:38so this is an entire mlp that we're
- 1:50:40defining now
- 1:50:45all the way until a single output
- 1:50:48okay
- 1:50:49and so obviously you would never
- 1:50:50differentiate on pen and paper these
- 1:50:52expressions but with micrograd we will
- 1:50:55be able to back propagate all the way
- 1:50:56through this
- 1:50:58and back propagate
- 1:50:59into
- 1:51:00these weights of all these neurons so
- 1:51:02let's see how that works okay so let's
- 1:51:04create ourselves a very simple
- 1:51:06example data set here
- 1:51:08so this data set has four examples
- 1:51:11and so we have four possible
- 1:51:13inputs into the neural net
- 1:51:15and we have four desired targets so we'd
- 1:51:17like the neural net to assign
- 1:51:21or output 1.0 when it's fed this example
- 1:51:24negative one when it's fed these
- 1:51:25examples and one when it's fed this
- 1:51:26example so it's a very simple binary
- 1:51:28classifier neural net basically that we
- 1:51:30would like here
- 1:51:32now let's think what the neural net
- 1:51:33currently thinks about these four
- 1:51:34examples we can just get their
- 1:51:36predictions
- 1:51:37um basically we can just call n of x for
- 1:51:40x in axis
- 1:51:42and then we can
- 1:51:43print
- 1:51:45so these are the outputs of the neural
- 1:51:46net on those four examples
- 1:51:48so
- 1:51:50the first one is 0.91 but we'd like it
- 1:51:52to be one so we should push this one
- 1:51:55higher this one we want to be higher
- 1:51:58this one says 0.88 and we want this to
- 1:52:00be negative one
- 1:52:02this is 0.8 we want it to be negative
- 1:52:04one
- 1:52:05and this one is 0.8 we want it to be one
- 1:52:08so how do we make the neural net and how
- 1:52:10do we tune the weights
- 1:52:12to
- 1:52:12better predict the desired targets
- 1:52:16and the trick used in deep learning to
- 1:52:18achieve this is to
- 1:52:20calculate a single number that somehow
- 1:52:22measures the total performance of your
- 1:52:24neural net and we call this single
- 1:52:25number the loss
- 1:52:28so the loss
- 1:52:29first
- 1:52:31is is a single number that we're going
- 1:52:32to define that basically measures how
- 1:52:34well the neural net is performing right
- 1:52:36now we have the intuitive sense that
- 1:52:37it's not performing very well because
- 1:52:38we're not very much close to this
- 1:52:40so the loss will be high and we'll want
- 1:52:43to minimize the loss
- 1:52:44so in particular in this case what we're
- 1:52:46going to do is we're going to implement
- 1:52:47the mean squared error loss
- 1:52:49so this is doing is we're going to
- 1:52:51basically iterate um
- 1:52:54for y ground truth
- 1:52:56and y output in zip of um
- 1:52:59wise and white red so we're going to
- 1:53:01pair up the
- 1:53:03ground truths with the predictions
- 1:53:06and this zip iterates over tuples of
- 1:53:07them
- 1:53:08and for each
- 1:53:11y ground truth and y output we're going
- 1:53:13to subtract them
- 1:53:16and square them
- 1:53:18so let's first see what these losses are
- 1:53:20these are individual loss components
- 1:53:22and so basically for each
- 1:53:25one of the four
- 1:53:26we are taking the prediction and the
- 1:53:28ground truth we are subtracting them and
- 1:53:30squaring them
- 1:53:32so because
- 1:53:33this one is so close to its target 0.91
- 1:53:36is almost one
- 1:53:38subtracting them gives a very small
- 1:53:40number
- 1:53:41so here we would get like a negative
- 1:53:43point one and then squaring it
- 1:53:45just makes sure
- 1:53:47that regardless of whether we are more
- 1:53:49negative or more positive we always get
- 1:53:51a positive
- 1:53:52number instead of squaring we should we
- 1:53:55could also take for example the absolute
- 1:53:56value we need to discard the sign
- 1:53:59and so you see that the expression is
- 1:54:00ranged so that you only get zero exactly
- 1:54:03when y out is equal to y ground truth
- 1:54:06when those two are equal so your
- 1:54:07prediction is exactly the target you are
- 1:54:09going to get zero
- 1:54:10and if your prediction is not the target
- 1:54:12you are going to get some other number
- 1:54:15so here for example we are way off and
- 1:54:17so that's why the loss is quite high
- 1:54:19and the more off we are the greater the
- 1:54:22loss will be
- 1:54:24so we don't want high loss we want low
- 1:54:26loss
- 1:54:27and so the final loss here will be just
- 1:54:30the sum
- 1:54:32of all of these
- 1:54:33numbers
- 1:54:34so you see that this should be zero
- 1:54:36roughly plus zero roughly
- 1:54:38but plus
- 1:54:39seven
- 1:54:40so loss should be about seven
- 1:54:43here
- 1:54:44and now we want to minimize the loss we
- 1:54:47want the loss to be low
- 1:54:49because if loss is low
- 1:54:51then every one of the predictions is
- 1:54:54equal to its target
- 1:54:56so the loss the lowest it can be is zero
- 1:54:58and the greater it is the worse off the
- 1:55:01neural net is predicting
- 1:55:04so now of course if we do lost that
- 1:55:05backward
- 1:55:07something magical happened when i hit
- 1:55:09enter
- 1:55:10and the magical thing of course that
- 1:55:12happened is that we can look at
- 1:55:14end.layers.neuron and that layers at say
- 1:55:16like the the first layer
- 1:55:18that neurons at zero
- 1:55:22because remember that mlp has the layers
- 1:55:24which is a list
- 1:55:26and each layer has a neurons which is a
- 1:55:28list and that gives us an individual
- 1:55:29neuron
- 1:55:30and then it's got some weights
- 1:55:32and so we can for example look at the
- 1:55:34weights at zero
- 1:55:38um
- 1:55:40oops it's not called weights it's called
- 1:55:42w
- 1:55:44and that's a value but now this value
- 1:55:46also has a groud because of the backward
- 1:55:48pass
- 1:55:50and so we see that because this gradient
- 1:55:52here on this particular weight of this
- 1:55:54particular neuron of this particular
- 1:55:56layer is negative
- 1:55:57we see that its influence on the loss is
- 1:56:00also negative so slightly increasing
- 1:56:02this particular weight of this neuron of
- 1:56:04this layer would make the loss go down
- 1:56:08and we actually have this information
- 1:56:10for every single one of our neurons and
- 1:56:12all their parameters actually it's worth
- 1:56:13looking at also the draw dot loss by the
- 1:56:16way
- 1:56:17so previously we looked at the draw dot
- 1:56:19of a single neural neuron forward pass
- 1:56:21and that was already a large expression
- 1:56:23but what is this expression we actually
- 1:56:25forwarded
- 1:56:27every one of those four examples and
- 1:56:29then we have the loss on top of them
- 1:56:30with the mean squared error
- 1:56:32and so this is a really massive graph
- 1:56:36because this graph that we've built up
- 1:56:38now
- 1:56:39oh my gosh this graph that we've built
- 1:56:41up now
- 1:56:42which is kind of excessive it's
- 1:56:44excessive because it has four forward
- 1:56:46passes of a neural net for every one of
- 1:56:48the examples and then it has the loss on
- 1:56:50top
- 1:56:51and it ends with the value of the loss
- 1:56:53which was 7.12
- 1:56:55and this loss will now back propagate
- 1:56:56through all the four forward passes all
- 1:56:58the way through just every single
- 1:57:00intermediate value of the neural net
- 1:57:03all the way back to of course the
- 1:57:05parameters of the weights which are the
- 1:57:06input
- 1:57:07so these weight parameters here are
- 1:57:10inputs to this neural net
- 1:57:12and
- 1:57:13these numbers here these scalars are
- 1:57:15inputs to the neural net
- 1:57:16so if we went around here
- 1:57:18we'll probably find
- 1:57:20some of these examples this 1.0
- 1:57:22potentially maybe this 1.0 or you know
- 1:57:25some of the others and you'll see that
- 1:57:26they all have gradients as well
- 1:57:28the thing is these gradients on the
- 1:57:30input data are not that useful to us
- 1:57:33and that's because the input data seems
- 1:57:36to be not changeable it's it's a given
- 1:57:38to the problem and so it's a fixed input
- 1:57:40we're not going to be changing it or
- 1:57:42messing with it even though we do have
- 1:57:43gradients for it
- 1:57:46but some of these gradients here
- 1:57:49will be for the neural network
- 1:57:50parameters the ws and the bs and those
- 1:57:53we of course we want to change
- 1:57:55okay so now we're going to want some
- 1:57:58convenience code to gather up all of the
- 1:57:59parameters of the neural net so that we
- 1:58:01can operate on all of them
- 1:58:03simultaneously and every one of them we
- 1:58:05will nudge a tiny amount
- 1:58:08based on the gradient information
- 1:58:10so let's collect the parameters of the
- 1:58:11neural net all in one array
- 1:58:14so let's create a parameters of self
- 1:58:17that just
- 1:58:18returns celta w which is a list
- 1:58:22concatenated with
- 1:58:24a list of self.b
- 1:58:27so this will just return a list
- 1:58:29list plus list just you know gives you a
- 1:58:31list
- 1:58:32so that's parameters of neuron and i'm
- 1:58:35calling it this way because also pi
- 1:58:36torch has a parameters on every single
- 1:58:38and in module
- 1:58:40and uh it does exactly what we're doing
- 1:58:42here it just returns the
- 1:58:44parameter tensors for us as the
- 1:58:46parameter scalars
- 1:58:48now layer is also a module so it will
- 1:58:50have parameters
- 1:58:52itself
- 1:58:54and basically what we want to do here is
- 1:58:56something like this like
- 1:59:00params is here and then for
- 1:59:03neuron in salt out neurons
- 1:59:07we want to get neuron.parameters
- 1:59:10and we want to params.extend
- 1:59:14right so these are the parameters of
- 1:59:16this neuron and then we want to put them
- 1:59:17on top of params so params dot extend
- 1:59:21of peace
- 1:59:22and then we want to return brands
- 1:59:25so this is way too much code so actually
- 1:59:28there's a way to simplify this which is
- 1:59:31return
- 1:59:33p
- 1:59:35for neuron in self
- 1:59:38neurons
- 1:59:39for
- 1:59:41p in neuron dot parameters
- 1:59:45so it's a single list comprehension in
- 1:59:47python you can sort of nest them like
- 1:59:49this and you can um
- 1:59:51then create
- 1:59:52uh the desired
- 1:59:54array so this is these are identical
- 1:59:57we can take this out
- 2:00:00and then let's do the same here
- 2:00:04def parameters
- 2:00:06self
- 2:00:07and return
- 2:00:09a parameter for layer in self dot layers
- 2:00:13for
- 2:00:15p in layer dot parameters
- 2:00:20and that should be good
- 2:00:23now let me pop out this so
- 2:00:26we don't re-initialize our network
- 2:00:28because we need to re-initialize
- 2:00:31our
- 2:00:35okay so unfortunately we will have to
- 2:00:37probably re-initialize the network
- 2:00:38because we just add functionality
- 2:00:41because this class of course we i want
- 2:00:43to get all the and that parameters but
- 2:00:45that's not going to work because this is
- 2:00:47the old class
- 2:00:49okay
- 2:00:50so unfortunately we do have to
- 2:00:52reinitialize the network which will
- 2:00:53change some of the numbers
- 2:00:55but let me do that so that we pick up
- 2:00:57the new api we can now do in the
- 2:00:58parameters
- 2:01:00and these are all the weights and biases
- 2:01:02inside the entire neural net
- 2:01:05so in total this mlp has 41 parameters
- 2:01:11and
- 2:01:12now we'll be able to change them
- 2:01:15if we recalculate the loss here we see
- 2:01:18that unfortunately we have slightly
- 2:01:19different
- 2:01:22predictions and slightly different laws
- 2:01:26but that's okay
- 2:01:28okay so we see that this neurons
- 2:01:31gradient is slightly negative we can
- 2:01:33also look at its data right now
- 2:01:36which is 0.85 so this is the current
- 2:01:38value of this neuron and this is its
- 2:01:40gradient on the loss
- 2:01:43so what we want to do now is we want to
- 2:01:45iterate for every p in
- 2:01:47n dot parameters so for all the 41
- 2:01:49parameters in this neural net
- 2:01:51we actually want to change p data
- 2:01:55slightly
- 2:01:56according to the gradient information
- 2:01:59okay so
- 2:02:00dot dot to do here
- 2:02:02but this will be basically a tiny update
- 2:02:05in this gradient descent scheme in
- 2:02:08gradient descent we are thinking of the
- 2:02:10gradient as a vector pointing in the
- 2:02:13direction
- 2:02:14of
- 2:02:15increased
- 2:02:16loss
- 2:02:19and so
- 2:02:20in gradient descent we are modifying
- 2:02:22p data
- 2:02:24by a small step size in the direction of
- 2:02:26the gradient so the step size as an
- 2:02:28example could be like a very small
- 2:02:29number like 0.01 is the step size times
- 2:02:32p dot grad
- 2:02:35right
- 2:02:36but we have to think through some of the
- 2:02:37signs here
- 2:02:38so uh
- 2:02:40in particular working with this specific
- 2:02:43example here
- 2:02:44we see that if we just left it like this
- 2:02:47then this neuron's value
- 2:02:49would be currently increased by a tiny
- 2:02:51amount of the gradient
- 2:02:53the grain is negative so this value of
- 2:02:56this neuron would go slightly down it
- 2:02:58would become like 0.8 you know four or
- 2:03:00something like that
- 2:03:02but if this neuron's value goes lower
- 2:03:06that would actually
- 2:03:08increase the loss
- 2:03:10that's because
- 2:03:12the derivative of this neuron is
- 2:03:14negative so increasing
- 2:03:16this makes the loss go down so
- 2:03:19increasing it is what we want to do
- 2:03:21instead of decreasing it so basically
- 2:03:23what we're missing here is we're
- 2:03:24actually missing a negative sign
- 2:03:26and again this other interpretation
- 2:03:29and that's because we want to minimize
- 2:03:30the loss we don't want to maximize the
- 2:03:31loss we want to decrease it
- 2:03:33and the other interpretation as i
- 2:03:34mentioned is you can think of the
- 2:03:36gradient vector
- 2:03:37so basically just the vector of all the
- 2:03:39gradients
- 2:03:40as pointing in the direction of
- 2:03:42increasing
- 2:03:44the loss but then we want to decrease it
- 2:03:46so we actually want to go in the
- 2:03:47opposite direction
- 2:03:49and so you can convince yourself that
- 2:03:50this sort of plug does the right thing
- 2:03:51here with the negative because we want
- 2:03:53to minimize the loss
- 2:03:55so if we nudge all the parameters by
- 2:03:57tiny amount
- 2:04:00then we'll see that
- 2:04:02this data will have changed a little bit
- 2:04:04so now this neuron
- 2:04:06is a tiny amount greater
- 2:04:08value so 0.854 went to 0.857
- 2:04:13and that's a good thing because slightly
- 2:04:16increasing this neuron
- 2:04:18uh
- 2:04:18data makes the loss go down according to
- 2:04:21the gradient and so the correct thing
- 2:04:23has happened sign wise
- 2:04:26and so now what we would expect of
- 2:04:27course is that
- 2:04:29because we've changed all these
- 2:04:30parameters we expect that the loss
- 2:04:32should have gone down a bit
- 2:04:35so we want to re-evaluate the loss let
- 2:04:37me basically
- 2:04:39this is just a data definition that
- 2:04:41hasn't changed but the forward pass here
- 2:04:44of the network we can recalculate
- 2:04:49and actually let me do it outside here
- 2:04:51so that we can compare the two loss
- 2:04:52values
- 2:04:54so here if i recalculate the loss
- 2:04:57we'd expect the new loss now to be
- 2:04:59slightly lower than this number so
- 2:05:01hopefully what we're getting now is a
- 2:05:03tiny bit lower than 4.84
- 2:05:064.36
- 2:05:08okay and remember the way we've arranged
- 2:05:10this is that low loss means that our
- 2:05:12predictions are matching the targets so
- 2:05:15our predictions now are probably
- 2:05:16slightly closer to the
- 2:05:18targets and now all we have to do is we
- 2:05:22have to iterate this process
- 2:05:24so again um we've done the forward pass
- 2:05:26and this is the loss
- 2:05:28now we can lost that backward
- 2:05:30let me take these out and we can do a
- 2:05:32step size
- 2:05:34and now we should have a slightly lower
- 2:05:35loss 4.36 goes to 3.9
- 2:05:39and okay so
- 2:05:41we've done the forward pass here's the
- 2:05:43backward pass
- 2:05:44nudge
- 2:05:45and now the loss is 3.66
- 2:05:503.47
- 2:05:52and you get the idea we just continue
- 2:05:54doing this and this is uh gradient
- 2:05:56descent we're just iteratively doing
- 2:05:58forward pass backward pass update
- 2:06:01forward pass backward pass update and
- 2:06:02the neural net is improving its
- 2:06:04predictions
- 2:06:05so here if we look at why pred now
- 2:06:09like red
- 2:06:12we see that um
- 2:06:14this value should be getting closer to
- 2:06:16one
- 2:06:16so this value should be getting more
- 2:06:17positive these should be getting more
- 2:06:19negative and this one should be also
- 2:06:20getting more positive so if we just
- 2:06:22iterate this
- 2:06:23a few more times
- 2:06:26actually we may be able to afford go to
- 2:06:28go a bit faster let's try a slightly
- 2:06:30higher learning rate
- 2:06:34oops okay there we go so now we're at
- 2:06:350.31
- 2:06:39if you go too fast by the way if you try
- 2:06:41to make it too big of a step you may
- 2:06:43actually overstep
- 2:06:47it's overconfidence because again
- 2:06:48remember we don't actually know exactly
- 2:06:50about the loss function the loss
- 2:06:51function has all kinds of structure and
- 2:06:53we only know about the very local
- 2:06:55dependence of all these parameters on
- 2:06:57the loss but if we step too far
- 2:06:59we may step into you know a part of the
- 2:07:01loss that is completely different
- 2:07:03and that can destabilize training and
- 2:07:04make your loss actually blow up even
- 2:07:08so the loss is now 0.04 so actually the
- 2:07:11predictions should be really quite close
- 2:07:13let's take a look
- 2:07:15so you see how this is almost one
- 2:07:17almost negative one almost one we can
- 2:07:19continue going
- 2:07:21uh so
- 2:07:22yep backward
- 2:07:24update
- 2:07:25oops there we go so we went way too fast
- 2:07:28and um
- 2:07:29we actually overstepped
- 2:07:31so we got two uh too eager where are we
- 2:07:34now oops
- 2:07:36okay
- 2:07:37seven e negative nine so this is very
- 2:07:39very low loss
- 2:07:41and the predictions
- 2:07:43are basically perfect
- 2:07:45so somehow we
- 2:07:47basically we were doing way too big
- 2:07:48updates and we briefly exploded but then
- 2:07:50somehow we ended up getting into a
- 2:07:51really good spot so usually this
- 2:07:54learning rate and the tuning of it is a
- 2:07:56subtle art you want to set your learning
- 2:07:58rate if it's too low you're going to
- 2:08:00take way too long to converge but if
- 2:08:02it's too high the whole thing gets
- 2:08:03unstable and you might actually even
- 2:08:05explode the loss
- 2:08:07depending on your loss function
- 2:08:08so finding the step size to be just
- 2:08:10right it's it's a pretty subtle art
- 2:08:12sometimes when you're using sort of
- 2:08:14vanilla gradient descent
- 2:08:15but we happen to get into a good spot we
- 2:08:17can look at
- 2:08:19n-dot parameters
- 2:08:22so this is the setting of weights and
- 2:08:25biases
- 2:08:26that makes our network
- 2:08:29predict
- 2:08:30the desired targets
- 2:08:31very very close
- 2:08:33and
- 2:08:35basically we've successfully trained
- 2:08:37neural net
- 2:08:38okay let's make this a tiny bit more
- 2:08:40respectable and implement an actual
- 2:08:41training loop and what that looks like
- 2:08:43so this is the data definition that
- 2:08:45stays this is the forward pass
- 2:08:47um so
- 2:08:49for uh k in range you know we're going
- 2:08:52to
- 2:08:53take a bunch of steps
- 2:08:57first you do the forward pass
- 2:09:00we validate the loss
- 2:09:03let's re-initialize the neural net from
- 2:09:05scratch
- 2:09:06and here's the data
- 2:09:08and we first do before pass then we do
- 2:09:11the backward pass
- 2:09:19and then we do an update that's gradient
- 2:09:21descent
- 2:09:26and then we should be able to iterate
- 2:09:27this and we should be able to print the
- 2:09:29current step
- 2:09:30the current loss um let's just print the
- 2:09:33sort of
- 2:09:34number of the loss
- 2:09:36and
- 2:09:38that should be it
- 2:09:40and then the learning rate 0.01 is a
- 2:09:42little too small 0.1 we saw is like a
- 2:09:44little bit dangerously too high let's go
- 2:09:46somewhere in between
- 2:09:47and we'll optimize this for
- 2:09:50not 10 steps but let's go for say 20
- 2:09:52steps
- 2:09:54let me erase all of this junk
- 2:09:59and uh let's run the optimization
- 2:10:03and you see how we've actually converged
- 2:10:05slower in a more controlled manner and
- 2:10:08got to a loss that is very low
- 2:10:11so
- 2:10:12i expect white bread to be quite good
- 2:10:15there we go
- 2:10:19um
- 2:10:22and
- 2:10:23that's it
- 2:10:24okay so this is kind of embarrassing but
- 2:10:25we actually have a really terrible bug
- 2:10:28in here and it's a subtle bug and it's a
- 2:10:31very common bug and i can't believe i've
- 2:10:33done it for the 20th time in my life
- 2:10:36especially on camera and i could have
- 2:10:38reshot the whole thing but i think it's
- 2:10:39pretty funny and you know you get to
- 2:10:41appreciate a bit what um working with
- 2:10:44neural nets maybe
- 2:10:45is like sometimes
- 2:10:47we are guilty of
- 2:10:50come bug i've actually tweeted
- 2:10:52the most common neural net mistakes a
- 2:10:54long time ago now
- 2:10:56uh and
- 2:10:57i'm not really
- 2:10:59gonna explain any of these except for we
- 2:11:01are guilty of number three you forgot to
- 2:11:03zero grad
- 2:11:04before that backward what is that
- 2:11:09basically what's happening and it's a
- 2:11:10subtle bug and i'm not sure if you saw
- 2:11:12it
- 2:11:12is that
- 2:11:14all of these
- 2:11:15weights here have a dot data and a dot
- 2:11:17grad
- 2:11:19and that grad starts at zero
- 2:11:22and then we do backward and we fill in
- 2:11:24the gradients
- 2:11:25and then we do an update on the data but
- 2:11:27we don't flush the grad
- 2:11:29it stays there
- 2:11:31so when we do the second
- 2:11:33forward pass and we do backward again
- 2:11:35remember that all the backward
- 2:11:36operations do a plus equals on the grad
- 2:11:39and so these gradients just
- 2:11:41add up and they never get reset to zero
- 2:11:44so basically we didn't zero grad so
- 2:11:47here's how we zero grad before
- 2:11:50backward
- 2:11:51we need to iterate over all the
- 2:11:52parameters
- 2:11:54and we need to make sure that p dot grad
- 2:11:56is set to zero
- 2:11:58we need to reset it to zero just like it
- 2:12:00is in the constructor
- 2:12:02so remember all the way here for all
- 2:12:04these value nodes grad is reset to zero
- 2:12:07and then all these backward passes do a
- 2:12:09plus equals from that grad
- 2:12:11but we need to make sure that
- 2:12:13we reset these graphs to zero so that
- 2:12:15when we do backward
- 2:12:17all of them start at zero and the actual
- 2:12:18backward pass accumulates um
- 2:12:21the loss derivatives into the grads
- 2:12:25so this is zero grad in pytorch
- 2:12:28and uh
- 2:12:30we will slightly get we'll get a
- 2:12:31slightly different optimization let's
- 2:12:33reset the neural net
- 2:12:34the data is the same this is now i think
- 2:12:37correct
- 2:12:38and we get a much more
- 2:12:40you know we get a much more
- 2:12:42slower descent
- 2:12:44we still end up with pretty good results
- 2:12:46and we can continue this a bit more
- 2:12:48to get down lower
- 2:12:50and lower
- 2:12:51and lower
- 2:12:54yeah
- 2:12:56so the only reason that the previous
- 2:12:57thing worked it's extremely buggy um the
- 2:12:59only reason that worked is that
- 2:13:03this is a very very simple problem
- 2:13:05and it's very easy for this neural net
- 2:13:07to fit this data
- 2:13:09and so the grads ended up accumulating
- 2:13:12and it effectively gave us a massive
- 2:13:13step size and it made us converge
- 2:13:16extremely fast
- 2:13:19but basically now we have to do more
- 2:13:20steps to get to very low values of loss
- 2:13:24and get wipe red to be really good we
- 2:13:26can try to
- 2:13:27step a bit greater
- 2:13:34yeah we're gonna get closer and closer
- 2:13:36to one minus one and one
- 2:13:38so
- 2:13:39working with neural nets is sometimes
- 2:13:41tricky because
- 2:13:43uh
- 2:13:44you may have lots of bugs in the code
- 2:13:47and uh your network might actually work
- 2:13:49just like ours worked
- 2:13:51but chances are is that if we had a more
- 2:13:53complex problem then actually this bug
- 2:13:55would have made us not optimize the loss
- 2:13:57very well and we were only able to get
- 2:13:59away with it because
- 2:14:01the problem is very simple
- 2:14:03so let's now bring everything together
- 2:14:04and summarize what we learned
- 2:14:06what are neural nets neural nets are
- 2:14:09these mathematical expressions
- 2:14:11fairly simple mathematical expressions
- 2:14:13in the case of multi-layer perceptron
- 2:14:15that take
- 2:14:16input as the data and they take input
- 2:14:19the weights and the parameters of the
- 2:14:20neural net mathematical expression for
- 2:14:22the forward pass followed by a loss
- 2:14:24function and the loss function tries to
- 2:14:26measure the accuracy of the predictions
- 2:14:29and usually the loss will be low when
- 2:14:31your predictions are matching your
- 2:14:32targets or where the network is
- 2:14:34basically behaving well so we we
- 2:14:37manipulate the loss function so that
- 2:14:38when the loss is low the network is
- 2:14:40doing what you want it to do on your
- 2:14:42problem
- 2:14:44and then we backward the loss
- 2:14:46use backpropagation to get the gradient
- 2:14:48and then we know how to tune all the
- 2:14:50parameters to decrease the loss locally
- 2:14:52but then we have to iterate that process
- 2:14:54many times in what's called the gradient
- 2:14:55descent
- 2:14:56so we simply follow the gradient
- 2:14:58information and that minimizes the loss
- 2:15:01and the loss is arranged so that when
- 2:15:02the loss is minimized the network is
- 2:15:04doing what you want it to do
- 2:15:06and yeah so we just have a blob of
- 2:15:09neural stuff and we can make it do
- 2:15:11arbitrary things and that's what gives
- 2:15:13neural nets their power um
- 2:15:15it's you know this is a very tiny
- 2:15:16network with 41 parameters
- 2:15:19but you can build significantly more
- 2:15:20complicated neural nets with billions
- 2:15:24at this point almost trillions of
- 2:15:25parameters and it's a massive blob of
- 2:15:28neural tissue simulated neural tissue
- 2:15:31roughly speaking
- 2:15:32and you can make it do extremely complex
- 2:15:34problems and these neurons then have all
- 2:15:37kinds of very fascinating emergent
- 2:15:39properties
- 2:15:40in
- 2:15:41when you try to make them do
- 2:15:43significantly hard problems as in the
- 2:15:45case of gpt for example
- 2:15:47we have massive amounts of text from the
- 2:15:49internet and we're trying to get a
- 2:15:51neural net to predict to take like a few
- 2:15:53words and try to predict the next word
- 2:15:55in a sequence that's the learning
- 2:15:56problem
- 2:15:57and it turns out that when you train
- 2:15:58this on all of internet the neural net
- 2:16:00actually has like really remarkable
- 2:16:02emergent properties but that neural net
- 2:16:04would have hundreds of billions of
- 2:16:05parameters
- 2:16:07but it works on fundamentally the exact
- 2:16:09same principles
- 2:16:10the neural net of course will be a bit
- 2:16:12more complex but otherwise the
- 2:16:15value in the gradient is there
- 2:16:17and would be identical and the gradient
- 2:16:19descent would be there and would be
- 2:16:21basically identical but people usually
- 2:16:23use slightly different updates this is a
- 2:16:25very simple stochastic gradient descent
- 2:16:27update
- 2:16:28um
- 2:16:29and the loss function would not be mean
- 2:16:30squared error they would be using
- 2:16:32something called the cross-entropy loss
- 2:16:34for predicting the next token so there's
- 2:16:36a few more details but fundamentally the
- 2:16:37neural network setup and neural network
- 2:16:39training is identical and pervasive and
- 2:16:42now you understand intuitively
- 2:16:44how that works under the hood in the
- 2:16:46beginning of this video i told you that
- 2:16:47by the end of it you would understand
- 2:16:48everything in micrograd and then we'd
- 2:16:50slowly build it up let me briefly prove
- 2:16:52that to you
- 2:16:54so i'm going to step through all the
- 2:16:55code that is in micrograd as of today
- 2:16:57actually potentially some of the code
- 2:16:59will change by the time you watch this
- 2:17:00video because i intend to continue
- 2:17:01developing micrograd
- 2:17:03but let's look at what we have so far at
- 2:17:05least init.pi is empty when you go to
- 2:17:07engine.pi that has the value
- 2:17:10everything here you should mostly
- 2:17:11recognize so we have the data.grad
- 2:17:13attributes we have the backward function
- 2:17:15uh we have the previous set of children
- 2:17:17and the operation that produced this
- 2:17:19value
- 2:17:20we have addition multiplication and
- 2:17:22raising to a scalar power
- 2:17:25we have the relu non-linearity which is
- 2:17:27slightly different type of nonlinearity
- 2:17:28than 10h that we used in this video
- 2:17:30both of them are non-linearities and
- 2:17:32notably 10h is not actually present in
- 2:17:34micrograd as of right now but i intend
- 2:17:37to add it later
- 2:17:38with the backward which is identical and
- 2:17:40then all of these other operations which
- 2:17:42are built up on top of operations here
- 2:17:45so values should be very recognizable
- 2:17:47except for the non-linearity used in
- 2:17:48this video
- 2:17:50um there's no massive difference between
- 2:17:52relu and 10h and sigmoid and these other
- 2:17:54non-linearities they're all roughly
- 2:17:55equivalent and can be used in mlps so i
- 2:17:58use 10h because it's a bit smoother and
- 2:18:00because it's a little bit more
- 2:18:01complicated than relu and therefore it's
- 2:18:03stressed a little bit more the
- 2:18:05local gradients and working with those
- 2:18:07derivatives which i thought would be
- 2:18:09useful
- 2:18:10and then that pi is the neural networks
- 2:18:12library as i mentioned so you should
- 2:18:14recognize identical implementation of
- 2:18:16neuron layer and mlp
- 2:18:18notably or not so much
- 2:18:20we have a class module here there is a
- 2:18:22parent class of all these modules i did
- 2:18:24that because there's an nn.module class
- 2:18:27in pytorch and so this exactly matches
- 2:18:29that api and end.module and pytorch has
- 2:18:31also a zero grad which i've refactored
- 2:18:33out here
- 2:18:36so that's the end of micrograd really
- 2:18:38then there's a test
- 2:18:40which you'll see
- 2:18:41basically creates
- 2:18:42two chunks of code one in micrograd and
- 2:18:45one in pi torch and we'll make sure that
- 2:18:47the forward and the backward pass agree
- 2:18:49identically
- 2:18:50for a slightly less complicated
- 2:18:51expression a slightly more complicated
- 2:18:53expression everything
- 2:18:55agrees so we agree with pytorch on all
- 2:18:57of these operations
- 2:18:58and finally there's a demo.ipymb here
- 2:19:01and it's a bit more complicated binary
- 2:19:03classification demo than the one i
- 2:19:04covered in this lecture so we only had a
- 2:19:07tiny data set of four examples um here
- 2:19:09we have a bit more complicated example
- 2:19:11with lots of blue points and lots of red
- 2:19:13points and we're trying to again build a
- 2:19:15binary classifier to distinguish uh two
- 2:19:18dimensional points as red or blue
- 2:19:20it's a bit more complicated mlp here
- 2:19:22with it's a bigger mlp
- 2:19:24the loss is a bit more complicated
- 2:19:26because
- 2:19:27it supports batches
- 2:19:29so because our dataset was so tiny we
- 2:19:31always did a forward pass on the entire
- 2:19:32data set of four examples but when your
- 2:19:35data set is like a million examples what
- 2:19:37we usually do in practice is we chair we
- 2:19:39basically pick out some random subset we
- 2:19:41call that a batch and then we only
- 2:19:43process the batch forward backward and
- 2:19:45update so we don't have to forward the
- 2:19:47entire training set
- 2:19:49so this supports batching because
- 2:19:51there's a lot more examples here
- 2:19:53we do a forward pass the loss is
- 2:19:55slightly more different this is a max
- 2:19:57margin loss that i implement here
- 2:20:00the one that we used was the mean
- 2:20:01squared error loss because it's the
- 2:20:03simplest one
- 2:20:04there's also the binary cross entropy
- 2:20:06loss all of them can be used for binary
- 2:20:08classification and don't make too much
- 2:20:10of a difference in the simple examples
- 2:20:11that we looked at so far
- 2:20:13there's something called l2
- 2:20:14regularization used here this has to do
- 2:20:17with generalization of the neural net
- 2:20:19and controls the overfitting in machine
- 2:20:21learning setting but i did not cover
- 2:20:23these concepts and concepts in this
- 2:20:24video potentially later
- 2:20:26and the training loop you should
- 2:20:27recognize so forward backward with zero
- 2:20:31grad
- 2:20:32and update and so on you'll notice that
- 2:20:35in the update here the learning rate is
- 2:20:36scaled as a function of number of
- 2:20:38iterations and it
- 2:20:40shrinks
- 2:20:41and this is something called learning
- 2:20:43rate decay so in the beginning you have
- 2:20:44a high learning rate and as the network
- 2:20:47sort of stabilizes near the end you
- 2:20:49bring down the learning rate to get some
- 2:20:50of the fine details in the end
- 2:20:53and in the end we see the decision
- 2:20:54surface of the neural net and we see
- 2:20:56that it learns to separate out the red
- 2:20:58and the blue area based on the data
- 2:21:00points
- 2:21:01so that's the slightly more complicated
- 2:21:03example and then we'll demo that hyper
- 2:21:05ymb that you're free to go over
- 2:21:07but yeah as of today that is micrograd i
- 2:21:10also wanted to show you a little bit of
- 2:21:11real stuff so that you get to see how
- 2:21:13this is actually implemented in
- 2:21:14production grade library like by torch
- 2:21:16uh so in particular i wanted to show i
- 2:21:18wanted to find and show you the backward
- 2:21:20pass for 10h in pytorch so here in
- 2:21:23micrograd we see that the backward
- 2:21:25password 10h is one minus t square
- 2:21:28where t is the output of the tanh of x
- 2:21:33times of that grad which is the chain
- 2:21:34rule so we're looking for something that
- 2:21:36looks like this
- 2:21:38now
- 2:21:39i went to pytorch um which has an open
- 2:21:42source github codebase and uh i looked
- 2:21:45through a lot of its code
- 2:21:47and honestly i i i spent about 15
- 2:21:49minutes and i couldn't find 10h
- 2:21:51and that's because these libraries
- 2:21:53unfortunately they grow in size and
- 2:21:55entropy and if you just search for 10h
- 2:21:57you get apparently 2 800 results and 400
- 2:22:01and 406 files so i don't know what these
- 2:22:04files are doing honestly
- 2:22:07and why there are so many mentions of
- 2:22:0910h but unfortunately these libraries
- 2:22:11are quite complex they're meant to be
- 2:22:12used not really inspected um
- 2:22:15eventually i did stumble on someone
- 2:22:18who tries to change the 10 h backward
- 2:22:21code for some reason
- 2:22:22and someone here pointed to the cpu
- 2:22:24kernel and the kuda kernel for 10 inch
- 2:22:26backward
- 2:22:27so this so basically depends on if
- 2:22:29you're using pi torch on a cpu device or
- 2:22:31on a gpu which these are different
- 2:22:33devices and i haven't covered this but
- 2:22:35this is the 10 h backwards kernel
- 2:22:37for uh cpu
- 2:22:40and the reason it's so large is that
- 2:22:43number one this is like if you're using
- 2:22:45a complex type which we haven't even
- 2:22:46talked about if you're using a specific
- 2:22:48data type of b-float 16 which we haven't
- 2:22:50talked about
- 2:22:52and then if you're not then this is the
- 2:22:54kernel and deep here we see something
- 2:22:57that resembles our backward pass so they
- 2:23:00have a times one minus
- 2:23:02b square uh so this b
- 2:23:05b here must be the output of the 10h and
- 2:23:07this is the health.grad so here we found
- 2:23:10it
- 2:23:11uh deep inside
- 2:23:14pi torch from this location for some
- 2:23:15reason inside binaryops kernel when 10h
- 2:23:18is not actually a binary op
- 2:23:21and then this is the gpu kernel
- 2:23:25we're not complex
- 2:23:26we're
- 2:23:27here and here we go with one line of
- 2:23:29code
- 2:23:30so we did find it but basically
- 2:23:33unfortunately these codepieces are very
- 2:23:34large and
- 2:23:36micrograd is very very simple but if you
- 2:23:38actually want to use real stuff uh
- 2:23:40finding the code for it you'll actually
- 2:23:41find that difficult
- 2:23:43i also wanted to show you a little
- 2:23:45example here where pytorch is showing
- 2:23:47you how can you can register a new type
- 2:23:49of function that you want to add to
- 2:23:51pytorch as a lego building block
- 2:23:53so here if you want to for example add a
- 2:23:55gender polynomial 3
- 2:23:59here's how you could do it you will
- 2:24:00register it as a class that
- 2:24:03subclasses storage.org that function
- 2:24:06and then you have to tell pytorch how to
- 2:24:07forward your new function
- 2:24:10and how to backward through it
- 2:24:12so as long as you can do the forward
- 2:24:14pass of this little function piece that
- 2:24:15you want to add and as long as you know
- 2:24:17the the local derivative the local
- 2:24:19gradients which are implemented in the
- 2:24:20backward pi torch will be able to back
- 2:24:22propagate through your function and then
- 2:24:24you can use this as a lego block in a
- 2:24:26larger lego castle of all the different
- 2:24:28lego blocks that pytorch already has
- 2:24:31and so that's the only thing you have to
- 2:24:32tell pytorch and everything would just
- 2:24:33work and you can register new types of
- 2:24:35functions
- 2:24:36in this way following this example
- 2:24:38and that is everything that i wanted to
- 2:24:40cover in this lecture
- 2:24:41so i hope you enjoyed building out
- 2:24:42micrograd with me i hope you find it
- 2:24:44interesting insightful
- 2:24:46and
- 2:24:47yeah i will post a lot of the links
- 2:24:50that are related to this video in the
- 2:24:51video description below i will also
- 2:24:53probably post a link to a discussion
- 2:24:55forum
- 2:24:56or discussion group where you can ask
- 2:24:58questions related to this video and then
- 2:25:00i can answer or someone else can answer
- 2:25:02your questions and i may also do a
- 2:25:04follow-up video that answers some of the
- 2:25:06most common questions
- 2:25:08but for now that's it i hope you enjoyed
- 2:25:10it if you did then please like and
- 2:25:11subscribe so that youtube knows to
- 2:25:13feature this video to more people
- 2:25:15and that's it for now i'll see you later
- 2:25:22now here's the problem
- 2:25:24we know
- 2:25:25dl by
- 2:25:28wait what is the problem
- 2:25:31and that's everything i wanted to cover
- 2:25:33in this lecture
- 2:25:34so i hope
- 2:25:35you enjoyed us building up microcraft
- 2:25:38micro crab
- 2:25:42okay now let's do the exact same thing
- 2:25:43for multiply because we can't do
- 2:25:44something like a times two
- 2:25:47oops
- 2:25:50i know what happened there
About this transcript
This page contains the full transcript of The spelled-out intro to neural networks and backpropagation: building micrograd by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 24,352 words across 4,183 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.