YouTube2Text

The spelled-out intro to neural networks and backpropagation: building micrograd — Transcript

by Andrej Karpathy · 24,352 words · 4,183 segments · language en · Watch on YouTube

Full transcript

  1. 0:00hello my name is andre
  2. 0:01and i've been training deep neural
  3. 0:02networks for a bit more than a decade
  4. 0:04and in this lecture i'd like to show you
  5. 0:06what neural network training looks like
  6. 0:08under the hood so in particular we are
  7. 0:10going to start with a blank jupiter
  8. 0:12notebook and by the end of this lecture
  9. 0:14we will define and train in neural net
  10. 0:16and you'll get to see everything that
  11. 0:18goes on under the hood and exactly
  12. 0:20sort of how that works on an intuitive
  13. 0:21level
  14. 0:22now specifically what i would like to do
  15. 0:24is i would like to take you through
  16. 0:26building of micrograd now micrograd is
  17. 0:29this library that i released on github
  18. 0:30about two years ago but at the time i
  19. 0:32only uploaded the source code and you'd
  20. 0:34have to go in by yourself and really
  21. 0:37figure out how it works
  22. 0:39so in this lecture i will take you
  23. 0:40through it step by step and kind of
  24. 0:42comment on all the pieces of it so what
  25. 0:44is micrograd and why is it interesting
  26. 0:47good
  27. 0:48um
  28. 0:49micrograd is basically an autograd
  29. 0:51engine autograd is short for automatic
  30. 0:53gradient and really what it does is it
  31. 0:55implements backpropagation now
  32. 0:57backpropagation is this algorithm that
  33. 0:59allows you to efficiently evaluate the
  34. 1:01gradient of
  35. 1:03some kind of a loss function with
  36. 1:05respect to the weights of a neural
  37. 1:07network and what that allows us to do
  38. 1:09then is we can iteratively tune the
  39. 1:11weights of that neural network to
  40. 1:12minimize the loss function and therefore
  41. 1:14improve the accuracy of the network so
  42. 1:16back propagation would be at the
  43. 1:18mathematical core of any modern deep
  44. 1:20neural network library like say pytorch
  45. 1:22or jaxx
  46. 1:24so the functionality of microgrant is i
  47. 1:25think best illustrated by an example so
  48. 1:27if we just scroll down here
  49. 1:29you'll see that micrograph basically
  50. 1:31allows you to build out mathematical
  51. 1:32expressions
  52. 1:34and um here what we are doing is we have
  53. 1:36an expression that we're building out
  54. 1:37where you have two inputs a and b
  55. 1:40and you'll see that a and b are negative
  56. 1:43four and two but we are wrapping those
  57. 1:46values into this value object that we
  58. 1:48are going to build out as part of
  59. 1:49micrograd
  60. 1:51so this value object will wrap the
  61. 1:53numbers themselves
  62. 1:54and then we are going to build out a
  63. 1:56mathematical expression here where a and
  64. 1:58b are transformed into c d and
  65. 2:01eventually e f and g
  66. 2:03and i'm showing some of the functions
  67. 2:05some of the functionality of micrograph
  68. 2:07and the operations that it supports so
  69. 2:08you can add two value objects you can
  70. 2:11multiply them you can raise them to a
  71. 2:13constant power you can offset by one
  72. 2:15negate squash at zero
  73. 2:18square divide by constant divide by it
  74. 2:21etc
  75. 2:22and so we're building out an expression
  76. 2:24graph with with these two inputs a and b
  77. 2:27and we're creating an output value of g
  78. 2:30and micrograd will in the background
  79. 2:32build out this entire mathematical
  80. 2:34expression so it will for example know
  81. 2:36that c is also a value
  82. 2:38c was a result of an addition operation
  83. 2:41and the
  84. 2:42child nodes of c are a and b because the
  85. 2:46and will maintain pointers to a and b
  86. 2:48value objects so we'll basically know
  87. 2:50exactly how all of this is laid out
  88. 2:53and then not only can we do what we call
  89. 2:55the forward pass where we actually look
  90. 2:57at the value of g of course that's
  91. 2:58pretty straightforward we will access
  92. 3:00that using the dot data attribute and so
  93. 3:03the output of the forward pass the value
  94. 3:06of g is 24.7 it turns out but the big
  95. 3:09deal is that we can also take this g
  96. 3:11value object and we can call that
  97. 3:13backward
  98. 3:14and this will basically uh initialize
  99. 3:16back propagation at the node g
  100. 3:19and what backpropagation is going to do
  101. 3:21is it's going to start at g and it's
  102. 3:23going to go backwards through that
  103. 3:25expression graph and it's going to
  104. 3:26recursively apply the chain rule from
  105. 3:28calculus
  106. 3:30and what that allows us to do then is
  107. 3:32we're going to evaluate basically the
  108. 3:34derivative of g with respect to all the
  109. 3:36internal nodes
  110. 3:38like e d and c but also with respect to
  111. 3:40the inputs a and b
  112. 3:43and then we can actually query this
  113. 3:45derivative of g with respect to a for
  114. 3:47example that's a dot grad in this case
  115. 3:50it happens to be 138 and the derivative
  116. 3:52of g with respect to b
  117. 3:54which also happens to be here 645
  118. 3:57and this derivative we'll see soon is
  119. 3:59very important information because it's
  120. 4:01telling us how a and b are affecting g
  121. 4:04through this mathematical expression so
  122. 4:06in particular
  123. 4:08a dot grad is 138 so if we slightly
  124. 4:11nudge a and make it slightly larger
  125. 4:14138 is telling us that g will grow and
  126. 4:18the slope of that growth is going to be
  127. 4:19138
  128. 4:20and the slope of growth of b is going to
  129. 4:22be 645. so that's going to tell us about
  130. 4:25how g will respond if a and b get
  131. 4:27tweaked a tiny amount in a positive
  132. 4:29direction
  133. 4:31okay
  134. 4:33now you might be confused about what
  135. 4:34this expression is that we built out
  136. 4:36here and this expression by the way is
  137. 4:38completely meaningless i just made it up
  138. 4:40i'm just flexing about the kinds of
  139. 4:42operations that are supported by
  140. 4:43micrograd
  141. 4:44what we actually really care about are
  142. 4:46neural networks but it turns out that
  143. 4:48neural networks are just mathematical
  144. 4:49expressions just like this one but
  145. 4:51actually slightly bit less crazy even
  146. 4:54neural networks are just a mathematical
  147. 4:56expression they take the input data as
  148. 4:59an input and they take the weights of a
  149. 5:00neural network as an input and it's a
  150. 5:02mathematical expression and the output
  151. 5:04are your predictions of your neural net
  152. 5:06or the loss function we'll see this in a
  153. 5:08bit but basically neural networks just
  154. 5:10happen to be a certain class of
  155. 5:12mathematical expressions
  156. 5:13but back propagation is actually
  157. 5:15significantly more general it doesn't
  158. 5:17actually care about neural networks at
  159. 5:18all it only tells us about arbitrary
  160. 5:20mathematical expressions and then we
  161. 5:22happen to use that machinery for
  162. 5:24training of neural networks now one more
  163. 5:26note i would like to make at this stage
  164. 5:28is that as you see here micrograd is a
  165. 5:30scalar valued auto grant engine so it's
  166. 5:32working on the you know level of
  167. 5:34individual scalars like negative four
  168. 5:36and two and we're taking neural nets and
  169. 5:37we're breaking them down all the way to
  170. 5:39these atoms of individual scalars and
  171. 5:41all the little pluses and times and it's
  172. 5:43just excessive and so obviously you
  173. 5:45would never be doing any of this in
  174. 5:47production it's really just put down for
  175. 5:48pedagogical reasons because it allows us
  176. 5:50to not have to deal with these
  177. 5:52n-dimensional tensors that you would use
  178. 5:54in modern deep neural network library so
  179. 5:56this is really done so that you
  180. 5:58understand and refactor out back
  181. 6:00propagation and chain rule and
  182. 6:02understanding of neurologic training
  183. 6:04and then if you actually want to train
  184. 6:06bigger networks you have to be using
  185. 6:08these tensors but none of the math
  186. 6:09changes this is done purely for
  187. 6:11efficiency we are basically taking scale
  188. 6:13value
  189. 6:14all the scale values we're packaging
  190. 6:16them up into tensors which are just
  191. 6:17arrays of these scalars and then because
  192. 6:20we have these large arrays we're making
  193. 6:22operations on those large arrays that
  194. 6:24allows us to take advantage of the
  195. 6:26parallelism in a computer and all those
  196. 6:28operations can be done in parallel and
  197. 6:30then the whole thing runs faster but
  198. 6:32really none of the math changes and
  199. 6:33that's done purely for efficiency so i
  200. 6:35don't think that it's pedagogically
  201. 6:36useful to be dealing with tensors from
  202. 6:38scratch uh and i think and that's why i
  203. 6:40fundamentally wrote micrograd because
  204. 6:42you can understand how things work uh at
  205. 6:44the fundamental level and then you can
  206. 6:46speed it up later okay so here's the fun
  207. 6:48part my claim is that micrograd is what
  208. 6:51you need to train your networks and
  209. 6:52everything else is just efficiency so
  210. 6:54you'd think that micrograd would be a
  211. 6:56very complex piece of code and that
  212. 6:58turns out to not be the case
  213. 7:01so if we just go to micrograd
  214. 7:03and you'll see that there's only two
  215. 7:05files here in micrograd this is the
  216. 7:07actual engine it doesn't know anything
  217. 7:09about neural nuts and this is the entire
  218. 7:10neural nets library
  219. 7:12on top of micrograd so engine and nn.pi
  220. 7:17so the actual backpropagation autograd
  221. 7:19engine
  222. 7:21that gives you the power of neural
  223. 7:22networks is literally
  224. 7:26100 lines of code of like very simple
  225. 7:28python
  226. 7:30which we'll understand by the end of
  227. 7:31this lecture
  228. 7:32and then nn.pi
  229. 7:33this neural network library built on top
  230. 7:35of the autograd engine
  231. 7:37um is like a joke it's like
  232. 7:40we have to define what is a neuron and
  233. 7:42then we have to define what is the layer
  234. 7:44of neurons and then we define what is a
  235. 7:46multi-layer perceptron which is just a
  236. 7:47sequence of layers of neurons and so
  237. 7:50it's just a total joke
  238. 7:52so basically
  239. 7:53there's a lot of power that comes from
  240. 7:55only 150 lines of code
  241. 7:57and that's all you need to understand to
  242. 7:59understand neural network training and
  243. 8:00everything else is just efficiency and
  244. 8:02of course there's a lot to efficiency
  245. 8:05but fundamentally that's all that's
  246. 8:07happening okay so now let's dive right
  247. 8:09in and implement micrograph step by step
  248. 8:11the first thing i'd like to do is i'd
  249. 8:12like to make sure that you have a very
  250. 8:13good understanding intuitively of what a
  251. 8:16derivative is and exactly what
  252. 8:18information it gives you so let's start
  253. 8:20with some basic imports that i copy
  254. 8:22paste in every jupiter notebook always
  255. 8:25and let's define a function a scalar
  256. 8:27valued function
  257. 8:28f of x
  258. 8:30as follows
  259. 8:31so i just make this up randomly i just
  260. 8:33want to scale a valid function that
  261. 8:34takes a single scalar x and returns a
  262. 8:36single scalar y
  263. 8:38and we can call this function of course
  264. 8:40so we can pass in say 3.0 and get 20
  265. 8:42back
  266. 8:43now we can also plot this function to
  267. 8:45get a sense of its shape you can tell
  268. 8:47from the mathematical expression that
  269. 8:48this is probably a parabola it's a
  270. 8:50quadratic
  271. 8:51and so if we just uh create a set of um
  272. 8:56um
  273. 8:57scale values that we can feed in using
  274. 8:59for example a range from negative five
  275. 9:01to five in steps of 0.25
  276. 9:03so this is so axis is just from negative
  277. 9:065 to 5 not including 5 in steps of 0.25
  278. 9:11and we can actually call this function
  279. 9:12on this numpy array as well so we get a
  280. 9:14set of y's if we call f on axis
  281. 9:17and these y's are basically
  282. 9:20also applying a function on every one of
  283. 9:23these elements independently
  284. 9:25and we can plot this using matplotlib so
  285. 9:28plt.plot x's and y's and we get a nice
  286. 9:31parabola so previously here we fed in
  287. 9:333.0 somewhere here and we received 20
  288. 9:36back which is here the y coordinate so
  289. 9:39now i'd like to think through
  290. 9:40what is the derivative
  291. 9:42of this function at any single input
  292. 9:44point x
  293. 9:45right so what is the derivative at
  294. 9:47different points x of this function now
  295. 9:49if you remember back to your calculus
  296. 9:51class you've probably derived
  297. 9:52derivatives so we take this mathematical
  298. 9:54expression 3x squared minus 4x plus 5
  299. 9:57and you would write out on a piece of
  300. 9:58paper and you would you know apply the
  301. 9:59product rule and all the other rules and
  302. 10:01derive the mathematical expression of
  303. 10:03the great derivative of the original
  304. 10:05function and then you could plug in
  305. 10:06different texts and see what the
  306. 10:08derivative is
  307. 10:09we're not going to actually do that
  308. 10:11because no one in neural networks
  309. 10:13actually writes out the expression for
  310. 10:15the neural net it would be a massive
  311. 10:16expression um it would be you know
  312. 10:18thousands tens of thousands of terms no
  313. 10:20one actually derives the derivative of
  314. 10:22course and so we're not going to take
  315. 10:24this kind of like a symbolic approach
  316. 10:26instead what i'd like to do is i'd like
  317. 10:27to look at the definition of derivative
  318. 10:29and just make sure that we really
  319. 10:30understand what derivative is measuring
  320. 10:32what it's telling you about the function
  321. 10:34and so if we just look up derivative
  322. 10:42we see that
  323. 10:43okay so this is not a very good
  324. 10:44definition of derivative this is a
  325. 10:46definition of what it means to be
  326. 10:47differentiable
  327. 10:48but if you remember from your calculus
  328. 10:50it is the limit as h goes to zero of f
  329. 10:52of x plus h minus f of x over h so
  330. 10:55basically what it's saying is if you
  331. 10:58slightly bump up you're at some point x
  332. 11:00that you're interested in or a and if
  333. 11:02you slightly bump up
  334. 11:04you know you slightly increase it by
  335. 11:06small number h
  336. 11:08how does the function respond with what
  337. 11:09sensitivity does it respond what is the
  338. 11:11slope at that point does the function go
  339. 11:13up or does it go down and by how much
  340. 11:16and that's the slope of that function
  341. 11:18the
  342. 11:18the slope of that response at that point
  343. 11:21and so we can basically evaluate
  344. 11:23the derivative here numerically by
  345. 11:26taking a very small h of course the
  346. 11:28definition would ask us to take h to
  347. 11:30zero we're just going to pick a very
  348. 11:31small h 0.001
  349. 11:34and let's say we're interested in point
  350. 11:353.0 so we can look at f of x of course
  351. 11:37as 20
  352. 11:38and now f of x plus h
  353. 11:40so if we slightly nudge x in a positive
  354. 11:42direction how is the function going to
  355. 11:44respond
  356. 11:45and just looking at this do you expect
  357. 11:47do you expect f of x plus h to be
  358. 11:49slightly greater than 20 or do you
  359. 11:51expect to be slightly lower than 20
  360. 11:54and since this 3 is here and this is 20
  361. 11:57if we slightly go positively the
  362. 11:59function will respond positively so
  363. 12:01you'd expect this to be slightly greater
  364. 12:03than 20. and now by how much it's
  365. 12:05telling you the
  366. 12:06sort of the
  367. 12:07the strength of that slope right the the
  368. 12:09size of the slope so f of x plus h minus
  369. 12:12f of x this is how much the function
  370. 12:14responded
  371. 12:16in the positive direction and we have to
  372. 12:17normalize by the
  373. 12:19run so we have the rise over run to get
  374. 12:22the slope so this of course is just a
  375. 12:24numerical approximation of the slope
  376. 12:26because we have to make age very very
  377. 12:28small to converge to the exact amount
  378. 12:32now if i'm doing too many zeros
  379. 12:35at some point
  380. 12:36i'm gonna get an incorrect answer
  381. 12:38because we're using floating point
  382. 12:39arithmetic and the representations of
  383. 12:41all these numbers in computer memory is
  384. 12:43finite and at some point we get into
  385. 12:45trouble
  386. 12:46so we can converse towards the right
  387. 12:47answer with this approach
  388. 12:50but basically um at 3 the slope is 14.
  389. 12:54and you can see that by taking 3x
  390. 12:56squared minus 4x plus 5 and
  391. 12:58differentiating it in our head
  392. 13:00so 3x squared would be
  393. 13:026 x minus 4
  394. 13:04and then we plug in x equals 3 so that's
  395. 13:0718 minus 4 is 14. so this is correct
  396. 13:10so that's
  397. 13:12at 3. now how about the slope at say
  398. 13:15negative 3
  399. 13:17would you expect would you expect for
  400. 13:19the slope
  401. 13:20now telling the exact value is really
  402. 13:22hard but what is the sign of that slope
  403. 13:24so at negative three
  404. 13:26if we slightly go in the positive
  405. 13:28direction at x the function would
  406. 13:30actually go down and so that tells you
  407. 13:32that the slope would be negative so
  408. 13:33we'll get a slight number below
  409. 13:36below 20. and so if we take the slope we
  410. 13:39expect something negative
  411. 13:40negative 22. okay
  412. 13:43and at some point here of course the
  413. 13:45slope would be zero now for this
  414. 13:47specific function i looked it up
  415. 13:48previously and it's at point two over
  416. 13:51three
  417. 13:52so at roughly two over three
  418. 13:54uh that's somewhere here
  419. 13:55um
  420. 13:57this derivative be zero
  421. 13:59so basically at that precise point
  422. 14:03yeah
  423. 14:04at that precise point if we nudge in a
  424. 14:06positive direction the function doesn't
  425. 14:07respond this stays the same almost and
  426. 14:09so that's why the slope is zero okay now
  427. 14:11let's look at a bit more complex case
  428. 14:14so we're going to start you know
  429. 14:15complexifying a bit so now we have a
  430. 14:18function
  431. 14:19here
  432. 14:20with output variable d
  433. 14:22that is a function of three scalar
  434. 14:24inputs a b and c
  435. 14:26so a b and c are some specific values
  436. 14:28three inputs into our expression graph
  437. 14:30and a single output d
  438. 14:32and so if we just print d we get four
  439. 14:36and now what i have to do is i'd like to
  440. 14:38again look at the derivatives of d with
  441. 14:40respect to a b and c
  442. 14:42and uh think through uh again just the
  443. 14:44intuition of what this derivative is
  444. 14:46telling us
  445. 14:47so in order to evaluate this derivative
  446. 14:49we're going to get a bit hacky here
  447. 14:52we're going to again have a very small
  448. 14:53value of h
  449. 14:55and then we're going to fix the inputs
  450. 14:57at some
  451. 14:58values that we're interested in
  452. 15:00so these are the this is the point abc
  453. 15:02at which we're going to be evaluating
  454. 15:04the the
  455. 15:05derivative of d with respect to all a b
  456. 15:07and c at that point
  457. 15:09so there are the inputs and now we have
  458. 15:11d1 is that expression
  459. 15:13and then we're going to for example look
  460. 15:15at the derivative of d with respect to a
  461. 15:17so we'll take a and we'll bump it by h
  462. 15:19and then we'll get d2 to be the exact
  463. 15:22same function
  464. 15:23and now we're going to print um
  465. 15:26you know f1
  466. 15:28d1 is d1
  467. 15:31d2 is d2
  468. 15:32and print slope
  469. 15:35so the derivative or slope
  470. 15:37here will be um
  471. 15:39of course
  472. 15:41d2
  473. 15:42minus d1 divide h
  474. 15:44so d2 minus d1 is how much the function
  475. 15:47increased
  476. 15:48uh when we bumped
  477. 15:50the uh
  478. 15:51the specific input that we're interested
  479. 15:53in by a tiny amount
  480. 15:55and
  481. 15:56this is then normalized by h
  482. 15:59to get the slope
  483. 16:02so
  484. 16:03um
  485. 16:05yeah
  486. 16:06so this so if i just run this we're
  487. 16:08going to print
  488. 16:10d1
  489. 16:12which we know is four
  490. 16:15now d2 will be bumped a will be bumped
  491. 16:18by h
  492. 16:20so let's just think through
  493. 16:22a little bit uh what d2 will be uh
  494. 16:26printed out here
  495. 16:27in particular
  496. 16:29d1 will be four
  497. 16:31will d2 be a number slightly greater
  498. 16:33than four or slightly lower than four
  499. 16:35and that's going to tell us the sl the
  500. 16:37the sign of the derivative
  501. 16:40so
  502. 16:42we're bumping a by h
  503. 16:45b as minus three c is ten
  504. 16:48so you can just intuitively think
  505. 16:49through this derivative and what it's
  506. 16:51doing a will be slightly more positive
  507. 16:54and but b is a negative number
  508. 16:57so if a is slightly more positive
  509. 17:00because b is negative three
  510. 17:03we're actually going to be adding less
  511. 17:06to d
  512. 17:08so you'd actually expect that the value
  513. 17:10of the function will go down
  514. 17:13so let's just see this
  515. 17:16yeah and so we went from 4
  516. 17:18to 3.9996
  517. 17:20and that tells you that the slope will
  518. 17:22be negative
  519. 17:23and then
  520. 17:24uh will be a negative number
  521. 17:26because we went down
  522. 17:27and then
  523. 17:29the exact number of slope will be
  524. 17:31exact amount of slope is negative 3.
  525. 17:33and you can also convince yourself that
  526. 17:35negative 3 is the right answer
  527. 17:36mathematically and analytically because
  528. 17:39if you have a times b plus c and you are
  529. 17:41you know you have calculus then
  530. 17:43differentiating a times b plus c with
  531. 17:46respect to a gives you just b
  532. 17:48and indeed the value of b is negative 3
  533. 17:50which is the derivative that we have so
  534. 17:52you can tell that that's correct
  535. 17:54so now if we do this with b
  536. 17:57so if we bump b by a little bit in a
  537. 17:59positive direction we'd get different
  538. 18:02slopes so what is the influence of b on
  539. 18:04the output d
  540. 18:06so if we bump b by a tiny amount in a
  541. 18:08positive direction then because a is
  542. 18:10positive
  543. 18:11we'll be adding more to d
  544. 18:13right
  545. 18:14so um and now what is the what is the
  546. 18:17sensitivity what is the slope of that
  547. 18:18addition
  548. 18:19and it might not surprise you that this
  549. 18:21should be
  550. 18:222
  551. 18:24and y is a 2 because d of d
  552. 18:27by db differentiating with respect to b
  553. 18:30would be would give us a
  554. 18:31and the value of a is two so that's also
  555. 18:34working well
  556. 18:35and then if c gets bumped a tiny amount
  557. 18:37in h
  558. 18:38by h
  559. 18:39then of course a times b is unaffected
  560. 18:41and now c becomes slightly bit higher
  561. 18:44what does that do to the function it
  562. 18:45makes it slightly bit higher because
  563. 18:47we're simply adding c
  564. 18:48and it makes it slightly bit higher by
  565. 18:50the exact same amount that we added to c
  566. 18:53and so that tells you that the slope is
  567. 18:55one
  568. 18:56that will be the
  569. 18:59the rate at which
  570. 19:01d will increase as we scale
  571. 19:04c
  572. 19:05okay so we now have some intuitive sense
  573. 19:06of what this derivative is telling you
  574. 19:08about the function and we'd like to move
  575. 19:10to neural networks now as i mentioned
  576. 19:11neural networks will be pretty massive
  577. 19:13expressions mathematical expressions so
  578. 19:15we need some data structures that
  579. 19:16maintain these expressions and that's
  580. 19:17what we're going to start to build out
  581. 19:19now
  582. 19:20so we're going to
  583. 19:22build out this value object that i
  584. 19:24showed you in the readme page of
  585. 19:26micrograd
  586. 19:27so let me copy paste a skeleton of the
  587. 19:30first very simple value object
  588. 19:33so class value takes a single
  589. 19:36scalar value that it wraps and keeps
  590. 19:38track of
  591. 19:39and that's it so
  592. 19:41we can for example do value of 2.0 and
  593. 19:43then we can
  594. 19:45get we can look at its content and
  595. 19:48python will internally
  596. 19:50use the wrapper function
  597. 19:52to uh return
  598. 19:54uh this string oops
  599. 19:56like that
  600. 19:58so this is a value object with data
  601. 20:00equals two that we're creating here
  602. 20:03now we'd like to do is like we'd like to
  603. 20:04be able to
  604. 20:07have not just like two values
  605. 20:10but we'd like to do a bluffy right we'd
  606. 20:12like to add them
  607. 20:13so currently you would get an error
  608. 20:15because python doesn't know how to add
  609. 20:17two value objects so we have to tell it
  610. 20:21so here's
  611. 20:22addition
  612. 20:26so you have to basically use these
  613. 20:27special double underscore methods in
  614. 20:29python to define these operators for
  615. 20:31these objects so if we call um
  616. 20:35the uh if we use this plus operator
  617. 20:39python will internally call a dot add of
  618. 20:43b
  619. 20:43that's what will happen internally and
  620. 20:45so b will be the other and
  621. 20:48self will be a
  622. 20:50and so we see that what we're going to
  623. 20:52return is a new value object and it's
  624. 20:54just it's going to be wrapping
  625. 20:56the plus of
  626. 20:58their data
  627. 20:59but remember now because data is the
  628. 21:02actual like numbered python number so
  629. 21:04this operator here is just the typical
  630. 21:06floating point plus addition now it's
  631. 21:09not an addition of value objects
  632. 21:11and will return a new value so now a
  633. 21:14plus b should work and it should print
  634. 21:16value of
  635. 21:17negative one
  636. 21:18because that's two plus minus three
  637. 21:20there we go
  638. 21:21okay let's now implement multiply
  639. 21:24just so we can recreate this expression
  640. 21:25here
  641. 21:26so multiply i think it won't surprise
  642. 21:28you will be fairly similar
  643. 21:31so instead of add we're going to be
  644. 21:33using mul
  645. 21:34and then here of course we want to do
  646. 21:36times
  647. 21:36and so now we can create a c value
  648. 21:38object which will be 10.0 and now we
  649. 21:41should be able to do a times b well
  650. 21:44let's just do a times b first
  651. 21:46um
  652. 21:47[Music]
  653. 21:48that's value of negative six now
  654. 21:50and by the way i skipped over this a
  655. 21:52little bit suppose that i didn't have
  656. 21:53the wrapper function here
  657. 21:55then it's just that you'll get some kind
  658. 21:57of an ugly expression so what wrapper is
  659. 21:59doing is it's providing us a way to
  660. 22:02print out like a nicer looking
  661. 22:03expression in python
  662. 22:05uh so we don't just have something
  663. 22:07cryptic we actually are you know it's
  664. 22:09value of
  665. 22:10negative six so this gives us a times
  666. 22:14and then this we should now be able to
  667. 22:16add c to it because we've defined and
  668. 22:18told the python how to do mul and add
  669. 22:20and so this will call this will
  670. 22:22basically be equivalent to a dot
  671. 22:24small
  672. 22:26of b
  673. 22:27and then this new value object will be
  674. 22:29dot add
  675. 22:31of c
  676. 22:32and so let's see if that worked
  677. 22:34yep so that worked well that gave us
  678. 22:36four which is what we expect from before
  679. 22:39and i believe we can just call them
  680. 22:40manually as well there we go so
  681. 22:44yeah
  682. 22:45okay so now what we are missing is the
  683. 22:46connective tissue of this expression as
  684. 22:49i mentioned we want to keep these
  685. 22:50expression graphs so we need to know and
  686. 22:52keep pointers about what values produce
  687. 22:54what other values
  688. 22:56so here for example we are going to
  689. 22:58introduce a new variable which we'll
  690. 23:00call children and by default it will be
  691. 23:02an empty tuple
  692. 23:03and then we're actually going to keep a
  693. 23:04slightly different variable in the class
  694. 23:06which we'll call underscore prev which
  695. 23:08will be the set of children
  696. 23:11this is how i done i did it in the
  697. 23:13original micrograd looking at my code
  698. 23:15here i can't remember exactly the reason
  699. 23:17i believe it was efficiency but this
  700. 23:19underscore children will be a tuple for
  701. 23:20convenience but then when we actually
  702. 23:22maintain it in the class it will be just
  703. 23:23this set yeah i believe for efficiency
  704. 23:27um
  705. 23:28so now
  706. 23:29when we are creating a value like this
  707. 23:31with a constructor children will be
  708. 23:33empty and prep will be the empty set but
  709. 23:36when we're creating a value through
  710. 23:37addition or multiplication we're going
  711. 23:39to feed in the children of this value
  712. 23:42which in this case is self and other
  713. 23:46so those are the children
  714. 23:48here
  715. 23:50so now we can do d dot prev
  716. 23:52and we'll see that the children of the
  717. 23:55we now know are this value of negative 6
  718. 23:58and value of 10 and this of course is
  719. 24:00the value resulting from a times b and
  720. 24:03the c value which is 10.
  721. 24:06now the last piece of information we
  722. 24:08don't know so we know that the children
  723. 24:10of every single value but we don't know
  724. 24:12what operation created this value
  725. 24:14so we need one more element here let's
  726. 24:16call it underscore pop
  727. 24:19and by default this is the empty set for
  728. 24:21leaves
  729. 24:22and then we'll just maintain it here
  730. 24:25and now the operation will be just a
  731. 24:27simple string and in the case of
  732. 24:29addition it's plus in the case of
  733. 24:31multiplication is times
  734. 24:33so now we
  735. 24:35not just have d dot pref we also have a
  736. 24:37d dot up
  737. 24:38and we know that d was produced by an
  738. 24:40addition of those two values and so now
  739. 24:42we have the full
  740. 24:44mathematical expression uh and we're
  741. 24:46building out this data structure and we
  742. 24:47know exactly how each value came to be
  743. 24:49by word expression and from what other
  744. 24:51values
  745. 24:54now because these expressions are about
  746. 24:56to get quite a bit larger we'd like a
  747. 24:58way to nicely visualize these
  748. 25:00expressions that we're building out so
  749. 25:02for that i'm going to copy paste a bunch
  750. 25:03of slightly scary code that's going to
  751. 25:06visualize this these expression graphs
  752. 25:08for us
  753. 25:09so here's the code and i'll explain it
  754. 25:11in a bit but first let me just show you
  755. 25:13what this code does
  756. 25:14basically what it does is it creates a
  757. 25:16new function drawdot that we can call on
  758. 25:19some root node
  759. 25:20and then it's going to visualize it so
  760. 25:22if we call drawdot on d
  761. 25:24which is this final value here that is a
  762. 25:27times b plus c
  763. 25:29it creates something like this so this
  764. 25:31is d
  765. 25:32and you see that this is a times b
  766. 25:34creating an integrated value plus c
  767. 25:36gives us this output node d
  768. 25:40so that's dried out of d
  769. 25:42and i'm not going to go through this in
  770. 25:44complete detail you can take a look at
  771. 25:46graphless and its api uh graphis is a
  772. 25:48open source graph visualization software
  773. 25:51and what we're doing here is we're
  774. 25:52building out this graph and graphis
  775. 25:54api and
  776. 25:56you can basically see that trace is this
  777. 25:58helper function that enumerates all of
  778. 26:00the nodes and edges in the graph
  779. 26:02so that just builds a set of all the
  780. 26:04nodes and edges and then we iterate for
  781. 26:06all the nodes and we create special node
  782. 26:08objects
  783. 26:08for them in
  784. 26:11using dot node
  785. 26:13and then we also create edges using dot
  786. 26:15dot edge
  787. 26:16and the only thing that's like slightly
  788. 26:18tricky here is you'll notice that i
  789. 26:20basically add these fake nodes which are
  790. 26:22these operation nodes so for example
  791. 26:24this node here is just like a plus node
  792. 26:27and
  793. 26:28i create these
  794. 26:31special op nodes here
  795. 26:34and i connect them accordingly so these
  796. 26:37nodes of course are not actual
  797. 26:39nodes in the original graph
  798. 26:41they're not actually a value object the
  799. 26:43only value objects here are the things
  800. 26:46in squares those are actual value
  801. 26:48objects or representations thereof and
  802. 26:50these op nodes are just created in this
  803. 26:52drawdot routine so that it looks nice
  804. 26:55let's also add labels to these graphs
  805. 26:57just so we know what variables are where
  806. 26:59so let's create a special underscore
  807. 27:01label
  808. 27:02um
  809. 27:03or let's just do label
  810. 27:05equals empty by default and save it in
  811. 27:08each node
  812. 27:11and then here we're going to do label as
  813. 27:13a
  814. 27:15label is the
  815. 27:17label a c
  816. 27:22and then
  817. 27:24let's create a special um
  818. 27:27e equals a times b
  819. 27:30and e dot label will be e
  820. 27:34it's kind of naughty
  821. 27:35and e will be e plus c
  822. 27:38and a d dot label will be
  823. 27:40d
  824. 27:42okay so nothing really changes i just
  825. 27:44added this new e function
  826. 27:46a new e variable
  827. 27:48and then here when we are
  828. 27:50printing this
  829. 27:51i'm going to print the label here so
  830. 27:54this will be a percent s
  831. 27:56bar
  832. 27:56and this will be end.label
  833. 28:01and so now
  834. 28:03we have the label on the left here so it
  835. 28:05says a b creating e and then e plus c
  836. 28:07creates d
  837. 28:08just like we have it here
  838. 28:10and finally let's make this expression
  839. 28:12just one layer deeper
  840. 28:14so d will not be the final output node
  841. 28:17instead after d we are going to create a
  842. 28:20new value object
  843. 28:21called f we're going to start running
  844. 28:23out of variables soon f will be negative
  845. 28:252.0
  846. 28:27and its label will of course just be f
  847. 28:30and then l capital l will be the output
  848. 28:34of our graph
  849. 28:35and l will be p times f
  850. 28:38okay
  851. 28:38so l will be negative eight is the
  852. 28:40output
  853. 28:42so
  854. 28:44now we don't just draw a d we draw l
  855. 28:50okay
  856. 28:52and somehow the label of
  857. 28:54l was undefined oops all that label has
  858. 28:56to be explicitly sort of given to it
  859. 28:59there we go so l is the output
  860. 29:01so let's quickly recap what we've done
  861. 29:03so far
  862. 29:04we are able to build out mathematical
  863. 29:05expressions using only plus and times so
  864. 29:08far
  865. 29:09they are scalar valued along the way
  866. 29:11and we can do this forward pass
  867. 29:14and build out a mathematical expression
  868. 29:16so we have multiple inputs here a b c
  869. 29:18and f
  870. 29:19going into a mathematical expression
  871. 29:21that produces a single output l
  872. 29:24and this here is visualizing the forward
  873. 29:26pass so the output of the forward pass
  874. 29:28is negative eight that's the value
  875. 29:31now what we'd like to do next is we'd
  876. 29:33like to run back propagation
  877. 29:35and in back propagation we are going to
  878. 29:37start here at the end and we're going to
  879. 29:39reverse
  880. 29:40and calculate the gradient along along
  881. 29:43all these intermediate values
  882. 29:45and really what we're computing for
  883. 29:46every single value here
  884. 29:48um we're going to compute the derivative
  885. 29:50of that node with respect to l
  886. 29:55so
  887. 29:56the derivative of l with respect to l is
  888. 29:58just uh one
  889. 30:00and then we're going to derive what is
  890. 30:01the derivative of l with respect to f
  891. 30:03with respect to d with respect to c with
  892. 30:06respect to e
  893. 30:07with respect to b and with respect to a
  894. 30:10and in the neural network setting you'd
  895. 30:12be very interested in the derivative of
  896. 30:13basically this loss function l
  897. 30:16with respect to the weights of a neural
  898. 30:18network
  899. 30:19and here of course we have just these
  900. 30:20variables a b c and f
  901. 30:22but some of these will eventually
  902. 30:23represent the weights of a neural net
  903. 30:25and so we'll need to know how those
  904. 30:27weights are impacting
  905. 30:29the loss function so we'll be interested
  906. 30:31basically in the derivative of the
  907. 30:32output with respect to some of its leaf
  908. 30:34nodes and those leaf nodes will be the
  909. 30:36weights of the neural net
  910. 30:38and the other leaf nodes of course will
  911. 30:39be the data itself but usually we will
  912. 30:41not want or use the derivative of the
  913. 30:44loss function with respect to data
  914. 30:45because the data is fixed but the
  915. 30:47weights will be iterated on
  916. 30:50using the gradient information so next
  917. 30:52we are going to create a variable inside
  918. 30:54the value class that maintains the
  919. 30:57derivative of l with respect to that
  920. 30:59value
  921. 31:00and we will call this variable grad
  922. 31:03so there's a data and there's a
  923. 31:05self.grad
  924. 31:07and initially it will be zero and
  925. 31:09remember that zero is basically means no
  926. 31:12effect so at initialization we're
  927. 31:14assuming that every value does not
  928. 31:16impact does not affect the out the
  929. 31:18output
  930. 31:19right because if the gradient is zero
  931. 31:21that means that changing this variable
  932. 31:23is not changing the loss function
  933. 31:25so by default we assume that the
  934. 31:27gradient is zero
  935. 31:28and then
  936. 31:31now that we have grad and it's 0.0
  937. 31:36we are going to be able to visualize it
  938. 31:38here after data so here grad is 0.4 f
  939. 31:42and this will be in that graph
  940. 31:45and now we are going to be showing both
  941. 31:47the data and the grad
  942. 31:50initialized at zero
  943. 31:53and we are just about getting ready to
  944. 31:55calculate the back propagation
  945. 31:57and of course this grad again as i
  946. 31:58mentioned is representing
  947. 32:00the derivative of the output in this
  948. 32:02case l with respect to this value so
  949. 32:05with respect to so this is the
  950. 32:06derivative of l with respect to f with
  951. 32:08respect to d and so on so let's now fill
  952. 32:11in those gradients and actually do back
  953. 32:12propagation manually so let's start
  954. 32:14filling in these gradients and start all
  955. 32:16the way at the end as i mentioned here
  956. 32:18first we are interested to fill in this
  957. 32:20gradient here so what is the derivative
  958. 32:22of l with respect to l
  959. 32:25in other words if i change l by a tiny
  960. 32:27amount of h
  961. 32:29how much does
  962. 32:30l change
  963. 32:32it changes by h so it's proportional and
  964. 32:35therefore derivative will be one
  965. 32:37we can of course measure these or
  966. 32:39estimate these numerical gradients
  967. 32:40numerically just like we've seen before
  968. 32:43so if i take this expression
  969. 32:45and i create a def lol function here
  970. 32:49and put this here now the reason i'm
  971. 32:51creating a gating function hello here is
  972. 32:53because i don't want to pollute or mess
  973. 32:55up the global scope here this is just
  974. 32:57kind of like a little staging area and
  975. 32:58as you know in python all of these will
  976. 33:00be local variables to this function so
  977. 33:02i'm not changing any of the global scope
  978. 33:04here
  979. 33:05so here l1 will be l
  980. 33:10and then copy pasting this expression
  981. 33:13we're going to add a small amount h
  982. 33:17in for example a
  983. 33:20right and this would be measuring the
  984. 33:22derivative of l with respect to a
  985. 33:25so here this will be l2
  986. 33:28and then we want to print this
  987. 33:29derivative so print
  988. 33:31l2 minus l1 which is how much l changed
  989. 33:35and then normalize it by h so this is
  990. 33:37the rise over run
  991. 33:39and we have to be careful because l is a
  992. 33:41value node so we actually want its data
  993. 33:45um
  994. 33:46so that these are floats dividing by h
  995. 33:48and this should print the derivative of
  996. 33:50l with respect to a because a is the one
  997. 33:53that we bumped a little bit by h
  998. 33:55so what is the
  999. 33:57derivative of l with respect to a
  1000. 33:59it's six
  1001. 34:01okay and obviously
  1002. 34:03if we change l by h
  1003. 34:06then that would be
  1004. 34:09here effectively
  1005. 34:12this looks really awkward but changing l
  1006. 34:14by h
  1007. 34:16you see the derivative here is 1. um
  1008. 34:20that's kind of like the base case of
  1009. 34:23what we are doing here
  1010. 34:24so basically we cannot come up here and
  1011. 34:26we can manually set l.grad to one this
  1012. 34:29is our manual back propagation
  1013. 34:31l dot grad is one and let's redraw
  1014. 34:35and we'll see that we filled in grad as
  1015. 34:371 for l
  1016. 34:39we're now going to continue the back
  1017. 34:40propagation so let's here look at the
  1018. 34:42derivatives of l with respect to d and f
  1019. 34:45let's do a d first
  1020. 34:47so what we are interested in if i create
  1021. 34:49a markdown on here is we'd like to know
  1022. 34:51basically we have that l is d times f
  1023. 34:54and we'd like to know what is uh d
  1024. 34:57l by d d
  1025. 35:00what is that
  1026. 35:01and if you know your calculus uh l is d
  1027. 35:03times f so what is d l by d d it would
  1028. 35:06be f
  1029. 35:08and if you don't believe me we can also
  1030. 35:10just derive it because the proof would
  1031. 35:11be fairly straightforward uh we go to
  1032. 35:14the
  1033. 35:15definition of the derivative which is f
  1034. 35:18of x plus h minus f of x divide h
  1035. 35:22as a limit limit of h goes to zero of
  1036. 35:24this kind of expression so when we have
  1037. 35:26l is d times f
  1038. 35:28then increasing d by h
  1039. 35:31would give us the output of b plus h
  1040. 35:33times f
  1041. 35:35that's basically f of x plus h right
  1042. 35:38minus d times f
  1043. 35:42and then divide h and symbolically
  1044. 35:44expanding out here we would have
  1045. 35:46basically d times f plus h times f minus
  1046. 35:50t times f divide h
  1047. 35:52and then you see how the df minus df
  1048. 35:54cancels so you're left with h times f
  1049. 35:57divide h
  1050. 35:58which is f
  1051. 35:59so in the limit as h goes to zero of
  1052. 36:03you know
  1053. 36:04derivative
  1054. 36:06definition we just get f in the case of
  1055. 36:09d times f
  1056. 36:12so
  1057. 36:13symmetrically
  1058. 36:14dl by d
  1059. 36:15f will just be d
  1060. 36:18so what we have is that f dot grad
  1061. 36:21we see now is just the value of d
  1062. 36:24which is 4.
  1063. 36:28and we see that
  1064. 36:30d dot grad
  1065. 36:31is just uh the value of f
  1066. 36:36and so the value of f is negative two
  1067. 36:41so we'll set those manually
  1068. 36:45let me erase this markdown node and then
  1069. 36:47let's redraw what we have
  1070. 36:50okay
  1071. 36:51and let's just make sure that these were
  1072. 36:53correct so we seem to think that dl by
  1073. 36:56dd is negative two so let's double check
  1074. 36:59um let me erase this plus h from before
  1075. 37:02and now we want the derivative with
  1076. 37:03respect to f
  1077. 37:05so let's just come here when i create f
  1078. 37:06and let's do a plus h here and this
  1079. 37:08should print the derivative of l with
  1080. 37:10respect to f so we expect to see four
  1081. 37:14yeah and this is four up to floating
  1082. 37:16point
  1083. 37:17funkiness
  1084. 37:18and then dl by dd
  1085. 37:21should be f which is negative two
  1086. 37:25grad is negative two
  1087. 37:26so if we again come here and we change d
  1088. 37:31d dot data plus equals h right here
  1089. 37:35so we expect so we've added a little h
  1090. 37:37and then we see how l changed and we
  1091. 37:40expect to print
  1092. 37:42uh negative two
  1093. 37:44there we go
  1094. 37:47so we've numerically verified what we're
  1095. 37:49doing here is what kind of like an
  1096. 37:50inline gradient check gradient check is
  1097. 37:53when we
  1098. 37:54are deriving this like back propagation
  1099. 37:56and getting the derivative with respect
  1100. 37:57to all the intermediate results and then
  1101. 38:00numerical gradient is just you know
  1102. 38:03estimating it using small step size
  1103. 38:06now we're getting to the crux of
  1104. 38:08backpropagation so this will be the most
  1105. 38:10important node to understand because if
  1106. 38:12you understand the gradient for this
  1107. 38:14node you understand all of back
  1108. 38:16propagation and all of training of
  1109. 38:17neural nets basically
  1110. 38:19so we need to derive dl by bc
  1111. 38:23in other words the derivative of l with
  1112. 38:24respect to c
  1113. 38:26because we've computed all these other
  1114. 38:27gradients already
  1115. 38:29now we're coming here and we're
  1116. 38:30continuing the back propagation manually
  1117. 38:33so we want dl by dc and then we'll also
  1118. 38:36derive dl by de
  1119. 38:38now here's the problem
  1120. 38:40how do we derive dl
  1121. 38:41by dc
  1122. 38:44we actually know the derivative l with
  1123. 38:46respect to d so we know how l assessed
  1124. 38:48it to d
  1125. 38:50but how is l sensitive to c so if we
  1126. 38:53wiggle c how does that impact l
  1127. 38:55through d
  1128. 38:58so we know dl by dc
  1129. 39:01and we also here know how c impacts d
  1130. 39:04and so just very intuitively if you know
  1131. 39:06the impact that c is having on d and the
  1132. 39:09impact that d is having on l
  1133. 39:11then you should be able to somehow put
  1134. 39:12that information together to figure out
  1135. 39:14how c impacts l
  1136. 39:16and indeed this is what we can actually
  1137. 39:18do so in particular we know just
  1138. 39:20concentrating on d first let's look at
  1139. 39:22how what is the derivative basically of
  1140. 39:24d with respect to c so in other words
  1141. 39:27what is dd by dc
  1142. 39:31so here we know that d is c times c plus
  1143. 39:34e
  1144. 39:35that's what we know and now we're
  1145. 39:37interested in dd by dc
  1146. 39:39if you just know your calculus again and
  1147. 39:41you remember that differentiating c plus
  1148. 39:43e with respect to c you know that that
  1149. 39:45gives you
  1150. 39:461.0
  1151. 39:47and we can also go back to the basics
  1152. 39:49and derive this because again we can go
  1153. 39:51to our f of x plus h minus f of x
  1154. 39:54divide by h
  1155. 39:56that's the definition of a derivative as
  1156. 39:58h goes to zero
  1157. 40:00and so here
  1158. 40:01focusing on c and its effect on d
  1159. 40:04we can basically do the f of x plus h
  1160. 40:06will be
  1161. 40:07c is incremented by h plus e
  1162. 40:10that's the first evaluation of our
  1163. 40:12function minus
  1164. 40:14c plus e
  1165. 40:16and then divide h
  1166. 40:18and so what is this
  1167. 40:19uh just expanding this out this will be
  1168. 40:21c plus h plus e minus c minus e
  1169. 40:25divide h and then you see here how c
  1170. 40:27minus c cancels e minus e cancels we're
  1171. 40:30left with h over h which is 1.0
  1172. 40:33and so
  1173. 40:35by symmetry also d d by d
  1174. 40:38e
  1175. 40:39will be 1.0 as well
  1176. 40:42so basically the derivative of a sum
  1177. 40:44expression is very simple and and this
  1178. 40:46is the local derivative so i call this
  1179. 40:49the local derivative because we have the
  1180. 40:51final output value all the way at the
  1181. 40:52end of this graph and we're now like a
  1182. 40:54small node here
  1183. 40:55and this is a little plus node
  1184. 40:58and it the little plus node doesn't know
  1185. 41:00anything about the rest of the graph
  1186. 41:02that it's embedded in all it knows is
  1187. 41:04that it did a plus it took a c and an e
  1188. 41:07added them and created d
  1189. 41:09and this plus note also knows the local
  1190. 41:11influence of c on d or rather rather the
  1191. 41:14derivative of d with respect to c and it
  1192. 41:16also
  1193. 41:17knows the derivative of d with respect
  1194. 41:18to e but that's not what we want that's
  1195. 41:21just a local derivative what we actually
  1196. 41:23want is d l by d c and l could l is here
  1197. 41:27just one step away but in a general case
  1198. 41:30this little plus note is could be
  1199. 41:32embedded in like a massive graph
  1200. 41:34so
  1201. 41:35again we know how l impacts d and now we
  1202. 41:38know how c and e impact d how do we put
  1203. 41:41that information together to write dl by
  1204. 41:43dc and the answer of course is the chain
  1205. 41:46rule in calculus
  1206. 41:47and so um
  1207. 41:50i pulled up a chain rule here from
  1208. 41:51kapedia
  1209. 41:52and
  1210. 41:53i'm going to go through this very
  1211. 41:54briefly so chain rule
  1212. 41:57wikipedia sometimes can be very
  1213. 41:58confusing and calculus can
  1214. 42:00can be very confusing like this is the
  1215. 42:02way i
  1216. 42:03learned
  1217. 42:05chain rule and it was very confusing
  1218. 42:06like what is happening it's just
  1219. 42:08complicated so i like this expression
  1220. 42:10much better
  1221. 42:12if a variable z depends on a variable y
  1222. 42:15which itself depends on the variable x
  1223. 42:18then z depends on x as well obviously
  1224. 42:20through the intermediate variable y
  1225. 42:22in this case the chain rule is expressed
  1226. 42:24as
  1227. 42:25if you want dz by dx
  1228. 42:28then you take the dz by dy and you
  1229. 42:30multiply it by d y by dx
  1230. 42:33so the chain rule fundamentally is
  1231. 42:34telling you
  1232. 42:36how
  1233. 42:37we chain these
  1234. 42:39uh derivatives together
  1235. 42:41correctly so to differentiate through a
  1236. 42:44function composition
  1237. 42:46we have to apply a multiplication
  1238. 42:48of
  1239. 42:49those derivatives
  1240. 42:51so that's really what chain rule is
  1241. 42:53telling us
  1242. 42:54and there's a nice little intuitive
  1243. 42:56explanation here which i also think is
  1244. 42:58kind of cute the chain rule says that
  1245. 42:59knowing the instantaneous rate of change
  1246. 43:01of z with respect to y and y relative to
  1247. 43:03x allows one to calculate the
  1248. 43:04instantaneous rate of change of z
  1249. 43:06relative to x
  1250. 43:07as a product of those two rates of
  1251. 43:09change
  1252. 43:10simply the product of those two
  1253. 43:12so here's a good one
  1254. 43:14if a car travels twice as fast as
  1255. 43:16bicycle and the bicycle is four times as
  1256. 43:18fast as walking man
  1257. 43:19then the car travels two times four
  1258. 43:22eight times as fast as demand
  1259. 43:25and so this makes it very clear that the
  1260. 43:27correct thing to do sort of
  1261. 43:29is to multiply
  1262. 43:30so
  1263. 43:31cars twice as fast as bicycle and
  1264. 43:33bicycle is four times as fast as man
  1265. 43:36so the car will be eight times as fast
  1266. 43:38as the man and so we can take these
  1267. 43:42intermediate rates of change if you will
  1268. 43:44and multiply them together
  1269. 43:46and that justifies the
  1270. 43:48chain rule intuitively so have a look at
  1271. 43:50chain rule about here really what it
  1272. 43:52means for us is there's a very simple
  1273. 43:54recipe for deriving what we want which
  1274. 43:56is dl by dc
  1275. 43:59and what we have so far
  1276. 44:01is we know
  1277. 44:03want
  1278. 44:05and we know
  1279. 44:07what is the
  1280. 44:08impact of d on l so we know d l by
  1281. 44:12d d the derivative of l with respect to
  1282. 44:14d d we know that that's negative two
  1283. 44:17and now because of this local
  1284. 44:19reasoning that we've done here we know
  1285. 44:21dd by d
  1286. 44:23c
  1287. 44:24so how does c impact d and in
  1288. 44:27particular this is a plus node so the
  1289. 44:29local derivative is simply 1.0 it's very
  1290. 44:32simple
  1291. 44:33and so
  1292. 44:34the chain rule tells us that dl by dc
  1293. 44:37going through this intermediate variable
  1294. 44:40will just be simply d l by
  1295. 44:44d
  1296. 44:44times
  1297. 44:49dd by dc
  1298. 44:51that's chain rule
  1299. 44:53so this is identical to what's happening
  1300. 44:55here
  1301. 44:56except
  1302. 44:58z is rl
  1303. 44:59y is our d and x is rc
  1304. 45:03so we literally just have to multiply
  1305. 45:05these
  1306. 45:06and because
  1307. 45:10these local derivatives like dd by dc
  1308. 45:12are just one
  1309. 45:14we basically just copy over dl by dd
  1310. 45:17because this is just times one
  1311. 45:19so what does it do so because dl by dd
  1312. 45:22is negative two what is dl by dc
  1313. 45:25well it's the local gradient 1.0 times
  1314. 45:29dl by dd which is negative two
  1315. 45:31so literally what a plus node does you
  1316. 45:33can look at it that way is it literally
  1317. 45:35just routes the gradient
  1318. 45:37because the plus nodes local derivatives
  1319. 45:39are just one and so in the chain rule
  1320. 45:41one times
  1321. 45:43dl by dd
  1322. 45:45is um
  1323. 45:47is uh is just dl by dd and so that
  1324. 45:50derivative just gets routed to both c
  1325. 45:53and to e in this case
  1326. 45:55so basically um we have that that grad
  1327. 45:59or let's start with c since that's the
  1328. 46:01one we looked at
  1329. 46:02is
  1330. 46:03negative two times one
  1331. 46:06negative two
  1332. 46:08and in the same way by symmetry e that
  1333. 46:11grad will be negative two that's the
  1334. 46:13claim so we can set those
  1335. 46:16we can redraw
  1336. 46:19and you see how we just assign negative
  1337. 46:20to negative two so this backpropagating
  1338. 46:23signal which is carrying the information
  1339. 46:25of like what is the derivative of l with
  1340. 46:26respect to all the intermediate nodes
  1341. 46:28we can imagine it almost like flowing
  1342. 46:30backwards through the graph and a plus
  1343. 46:32node will simply distribute the
  1344. 46:34derivative to all the leaf nodes sorry
  1345. 46:36to all the children nodes of it
  1346. 46:39so this is the claim and now let's
  1347. 46:40verify it so let me remove the plus h
  1348. 46:43here from before
  1349. 46:45and now instead what we're going to do
  1350. 46:46is we're going to increment c so c dot
  1351. 46:48data will be credited by h
  1352. 46:50and when i run this we expect to see
  1353. 46:52negative 2
  1354. 46:54negative 2. and then of course for e
  1355. 46:58so e dot data plus equals h and we
  1356. 47:01expect to see negative 2.
  1357. 47:03simple
  1358. 47:07so those are the derivatives of these
  1359. 47:09internal nodes
  1360. 47:11and now we're going to recurse our way
  1361. 47:13backwards again
  1362. 47:15and we're again going to apply the chain
  1363. 47:17rule so here we go our second
  1364. 47:19application of chain rule and we will
  1365. 47:20apply it all the way through the graph
  1366. 47:22we just happen to only have one more
  1367. 47:24node remaining
  1368. 47:25we have that d l
  1369. 47:27by d e
  1370. 47:28as we have just calculated is negative
  1371. 47:30two so we know that
  1372. 47:32so we know the derivative of l with
  1373. 47:33respect to e
  1374. 47:36and now we want dl
  1375. 47:39by
  1376. 47:40da
  1377. 47:41right
  1378. 47:42and the chain rule is telling us that
  1379. 47:44that's just dl by de
  1380. 47:48negative 2
  1381. 47:50times the local gradient so what is the
  1382. 47:52local gradient basically d e
  1383. 47:55by d a
  1384. 47:56we have to look at that
  1385. 48:00so i'm a little times node
  1386. 48:02inside a massive graph
  1387. 48:04and i only know that i did a times b and
  1388. 48:06i produced an e
  1389. 48:09so now what is d e by d a and d e by d b
  1390. 48:12that's the only thing that i sort of
  1391. 48:14know about that's my local gradient
  1392. 48:17so
  1393. 48:17because we have that e's a times b we're
  1394. 48:20asking what is d e by d a
  1395. 48:24and of course we just did that here we
  1396. 48:26had a
  1397. 48:27times so i'm not going to rederive it
  1398. 48:30but if you want to differentiate this
  1399. 48:32with respect to a you'll just get b
  1400. 48:34right the value of b
  1401. 48:36which in this case is negative 3.0
  1402. 48:41so
  1403. 48:41basically we have that dl by da
  1404. 48:45well let me just do it right here we
  1405. 48:47have that a dot grad and we are applying
  1406. 48:49chain rule here
  1407. 48:50is d l by d e which we see here is
  1408. 48:54negative two
  1409. 48:56times
  1410. 48:57what is d e by d a
  1411. 48:59it's the value of b which is negative 3.
  1412. 49:04that's it
  1413. 49:07and then we have b grad is again dl by
  1414. 49:10de
  1415. 49:11which is negative 2
  1416. 49:13just the same way
  1417. 49:14times
  1418. 49:15what is d e by d
  1419. 49:18um db
  1420. 49:19is the value of a which is 2.2.0
  1421. 49:23as the value of a
  1422. 49:25so these are our claimed derivatives
  1423. 49:28let's
  1424. 49:30redraw
  1425. 49:32and we see here that
  1426. 49:33a dot grad turns out to be 6 because
  1427. 49:36that is negative 2 times negative 3
  1428. 49:38and b dot grad is negative 4
  1429. 49:41times sorry is negative 2 times 2 which
  1430. 49:43is negative 4.
  1431. 49:45so those are our claims let's delete
  1432. 49:47this and let's verify them
  1433. 49:50we have
  1434. 49:52a here a dot data plus equals h
  1435. 49:57so the claim is that
  1436. 49:59a dot grad is six
  1437. 50:01let's verify
  1438. 50:03six
  1439. 50:04and we have beta data
  1440. 50:07plus equals h
  1441. 50:08so nudging b by h
  1442. 50:11and looking at what happens
  1443. 50:13we claim it's negative four
  1444. 50:15and indeed it's negative four plus minus
  1445. 50:17again float oddness
  1446. 50:20um
  1447. 50:21and uh
  1448. 50:23that's it this
  1449. 50:24that was the manual
  1450. 50:26back propagation
  1451. 50:28uh all the way from here to all the leaf
  1452. 50:30nodes and we've done it piece by piece
  1453. 50:33and really all we've done is as you saw
  1454. 50:35we iterated through all the nodes one by
  1455. 50:37one and locally applied the chain rule
  1456. 50:39we always know what is the derivative of
  1457. 50:41l with respect to this little output and
  1458. 50:44then we look at how this output was
  1459. 50:45produced this output was produced
  1460. 50:47through some operation and we have the
  1461. 50:49pointers to the children nodes of this
  1462. 50:51operation
  1463. 50:52and so in this little operation we know
  1464. 50:54what the local derivatives are and we
  1465. 50:56just multiply them onto the derivative
  1466. 50:58always
  1467. 50:59so we just go through and recursively
  1468. 51:01multiply on the local derivatives and
  1469. 51:04that's what back propagation is is just
  1470. 51:05a recursive application of chain rule
  1471. 51:08backwards through the computation graph
  1472. 51:10let's see this power in action just very
  1473. 51:12briefly what we're going to do is we're
  1474. 51:14going to
  1475. 51:15nudge our inputs to try to make l go up
  1476. 51:19so in particular what we're doing is we
  1477. 51:21want a.data we're going to change it
  1478. 51:24and if we want l to go up that means we
  1479. 51:26just have to go in the direction of the
  1480. 51:27gradient so
  1481. 51:29a
  1482. 51:30should increase in the direction of
  1483. 51:32gradient by like some small step amount
  1484. 51:34this is the step size
  1485. 51:36and we don't just want this for ba but
  1486. 51:38also for b
  1487. 51:41also for c
  1488. 51:44also for f
  1489. 51:46those are
  1490. 51:47leaf nodes which we usually have control
  1491. 51:49over
  1492. 51:50and if we nudge in direction of the
  1493. 51:52gradient we expect a positive influence
  1494. 51:54on l
  1495. 51:55so we expect l to go up
  1496. 51:58positively
  1497. 51:59so it should become less negative it
  1498. 52:01should go up to say negative you know
  1499. 52:03six or something like that
  1500. 52:05uh it's hard to tell exactly and we'd
  1501. 52:08have to rewrite the forward pass so let
  1502. 52:09me just um
  1503. 52:12do that here
  1504. 52:13um
  1505. 52:16this would be the forward pass f would
  1506. 52:18be unchanged this is effectively the
  1507. 52:20forward pass and now if we print l.data
  1508. 52:24we expect because we nudged all the
  1509. 52:27values all the inputs in the rational
  1510. 52:28gradient we expected a less negative l
  1511. 52:30we expect it to go up
  1512. 52:32so maybe it's negative six or so let's
  1513. 52:34see what happens
  1514. 52:36okay negative seven
  1515. 52:38and uh this is basically one step of an
  1516. 52:41optimization that we'll end up running
  1517. 52:43and really does gradient just give us
  1518. 52:46some power because we know how to
  1519. 52:47influence the final outcome and this
  1520. 52:49will be extremely useful for training
  1521. 52:50knowledge as well as you'll see
  1522. 52:52so now i would like to do one more uh
  1523. 52:55example of manual backpropagation using
  1524. 52:58a bit more complex and uh useful example
  1525. 53:02we are going to back propagate through a
  1526. 53:04neuron
  1527. 53:05so
  1528. 53:07we want to eventually build up neural
  1529. 53:08networks and in the simplest case these
  1530. 53:10are multilateral perceptrons as they're
  1531. 53:12called so this is a two layer neural net
  1532. 53:15and it's got these hidden layers made up
  1533. 53:17of neurons and these neurons are fully
  1534. 53:18connected to each other
  1535. 53:20now biologically neurons are very
  1536. 53:21complicated devices but we have very
  1537. 53:23simple mathematical models of them
  1538. 53:26and so this is a very simple
  1539. 53:27mathematical model of a neuron you have
  1540. 53:29some inputs axis
  1541. 53:31and then you have these synapses that
  1542. 53:33have weights on them so
  1543. 53:36the w's are weights
  1544. 53:39and then
  1545. 53:40the synapse interacts with the input to
  1546. 53:42this neuron multiplicatively so what
  1547. 53:44flows to the cell body
  1548. 53:47of this neuron is w times x
  1549. 53:49but there's multiple inputs so there's
  1550. 53:51many w times x's flowing into the cell
  1551. 53:53body
  1552. 53:54the cell body then has also like some
  1553. 53:56bias
  1554. 53:57so this is kind of like the
  1555. 53:59inert innate sort of trigger happiness
  1556. 54:02of this neuron so this bias can make it
  1557. 54:04a bit more trigger happy or a bit less
  1558. 54:06trigger happy regardless of the input
  1559. 54:08but basically we're taking all the w
  1560. 54:10times x
  1561. 54:11of all the inputs adding the bias and
  1562. 54:13then we take it through an activation
  1563. 54:15function
  1564. 54:16and this activation function is usually
  1565. 54:18some kind of a squashing function
  1566. 54:20like a sigmoid or 10h or something like
  1567. 54:22that so as an example
  1568. 54:24we're going to use the 10h in this
  1569. 54:26example
  1570. 54:28numpy has a
  1571. 54:29np.10h
  1572. 54:31so
  1573. 54:32we can call it on a range
  1574. 54:34and we can plot it
  1575. 54:36this is the 10h function and you see
  1576. 54:38that the inputs as they come in
  1577. 54:41get squashed on the y coordinate here so
  1578. 54:44um
  1579. 54:45right at zero we're going to get exactly
  1580. 54:47zero and then as you go more positive in
  1581. 54:49the input
  1582. 54:50then you'll see that the function will
  1583. 54:52only go up to one and then plateau out
  1584. 54:55and so if you pass in very positive
  1585. 54:57inputs we're gonna cap it smoothly at
  1586. 55:00one and on the negative side we're gonna
  1587. 55:02cap it smoothly to negative one
  1588. 55:04so that's 10h
  1589. 55:06and that's the squashing function or an
  1590. 55:08activation function and what comes out
  1591. 55:10of this neuron is just the activation
  1592. 55:12function applied to the dot product of
  1593. 55:14the weights and the
  1594. 55:16inputs
  1595. 55:18so let's
  1596. 55:19write one out
  1597. 55:21um
  1598. 55:22i'm going to copy paste because
  1599. 55:27i don't want to type too much
  1600. 55:28but okay so here we have the inputs
  1601. 55:31x1 x2 so this is a two-dimensional
  1602. 55:33neuron so two inputs are going to come
  1603. 55:34in
  1604. 55:35these are thought out as the weights of
  1605. 55:37this neuron
  1606. 55:38weights w1 w2 and these weights again
  1607. 55:41are the synaptic strengths for each
  1608. 55:43input
  1609. 55:45and this is the bias of the neuron
  1610. 55:47b
  1611. 55:49and now we want to do is according to
  1612. 55:51this model we need to multiply x1 times
  1613. 55:54w1
  1614. 55:55and x2 times w2
  1615. 55:57and then we need to add bias on top of
  1616. 56:00it
  1617. 56:01and it gets a little messy here but all
  1618. 56:03we are trying to do is x1 w1 plus x2 w2
  1619. 56:06plus b
  1620. 56:07and these are multiply here
  1621. 56:09except i'm doing it in small steps so
  1622. 56:12that we actually have pointers to all
  1623. 56:13these intermediate nodes so we have x1
  1624. 56:15w1 variable x times x2 w2 variable and
  1625. 56:19i'm also labeling them
  1626. 56:21so n is now
  1627. 56:23the cell body raw
  1628. 56:25raw
  1629. 56:26activation without
  1630. 56:28the activation function for now
  1631. 56:30and this should be enough to basically
  1632. 56:32plot it so draw dot of n
  1633. 56:37gives us x1 times w1 x2 times w2
  1634. 56:41being added
  1635. 56:43then the bias gets added on top of this
  1636. 56:45and this n
  1637. 56:47is this sum
  1638. 56:49so we're now going to take it through an
  1639. 56:50activation function
  1640. 56:52and let's say we use the 10h
  1641. 56:54so that we produce the output
  1642. 56:56so what we'd like to do here is we'd
  1643. 56:58like to do the output and i'll call it o
  1644. 57:01is um
  1645. 57:03n dot 10h
  1646. 57:05okay but we haven't yet written the 10h
  1647. 57:08now the reason that we need to implement
  1648. 57:09another 10h function here is that
  1649. 57:12tanh is a
  1650. 57:14hyperbolic function and we've only so
  1651. 57:16far implemented a plus and the times and
  1652. 57:18you can't make a 10h out of just pluses
  1653. 57:20and times
  1654. 57:22you also need exponentiation so 10h is
  1655. 57:25this kind of a formula here
  1656. 57:27you can use either one of these and you
  1657. 57:28see that there's exponentiation involved
  1658. 57:30which we have not implemented yet for
  1659. 57:32our low value node here so we're not
  1660. 57:34going to be able to produce 10h yet and
  1661. 57:36we have to go back up and implement
  1662. 57:37something like it
  1663. 57:39now one option here
  1664. 57:42is we could actually implement um
  1665. 57:44exponentiation
  1666. 57:46right and we could return the x of a
  1667. 57:49value instead of a 10h of a value
  1668. 57:52because if we had x then we have
  1669. 57:54everything else that we need so um
  1670. 57:56because we know how to add and we know
  1671. 57:58how to
  1672. 58:00um
  1673. 58:01we know how to add and we know how to
  1674. 58:02multiply so we'd be able to create 10h
  1675. 58:04if we knew how to x
  1676. 58:06but for the purposes of this example i
  1677. 58:08specifically wanted to
  1678. 58:10show you
  1679. 58:11that we don't necessarily need to have
  1680. 58:13the most atomic pieces
  1681. 58:15in
  1682. 58:16um
  1683. 58:16in this value object we can actually
  1684. 58:19like create functions at arbitrary
  1685. 58:23points of abstraction they can be
  1686. 58:24complicated functions but they can be
  1687. 58:26also very very simple functions like a
  1688. 58:27plus and it's totally up to us the only
  1689. 58:30thing that matters is that we know how
  1690. 58:31to differentiate through any one
  1691. 58:33function so we take some inputs and we
  1692. 58:35make an output the only thing that
  1693. 58:37matters it can be arbitrarily complex
  1694. 58:38function as long as you know how to
  1695. 58:41create the local derivative if you know
  1696. 58:43the local derivative of how the inputs
  1697. 58:44impact the output then that's all you
  1698. 58:46need so we're going to cluster up
  1699. 58:49all of this expression and we're not
  1700. 58:51going to break it down to its atomic
  1701. 58:52pieces we're just going to directly
  1702. 58:54implement tanh
  1703. 58:55so let's do that
  1704. 58:57depth nh
  1705. 58:59and then out will be a value
  1706. 59:02of
  1707. 59:03and we need this expression here so
  1708. 59:05um
  1709. 59:08let me actually
  1710. 59:10copy paste
  1711. 59:14let's grab n which is a cell.theta
  1712. 59:17and then this
  1713. 59:18i believe is the tan h
  1714. 59:21math.x of
  1715. 59:24two
  1716. 59:25no n
  1717. 59:27n minus one over
  1718. 59:28two n plus one
  1719. 59:30maybe i can call this x
  1720. 59:33just so that it matches exactly
  1721. 59:35okay and now
  1722. 59:37this will be t
  1723. 59:40and uh children of this node there's
  1724. 59:42just one child
  1725. 59:44and i'm wrapping it in a tuple so this
  1726. 59:46is a tuple of one object just self
  1727. 59:48and here the name of this operation will
  1728. 59:50be 10h
  1729. 59:52and we're going to return that
  1730. 59:56okay
  1731. 59:58so now valley should be implementing 10h
  1732. 1:00:02and now we can scroll all the way down
  1733. 1:00:03here
  1734. 1:00:04and we can actually do n.10 h and that's
  1735. 1:00:06going to return the tanhd
  1736. 1:00:09output of n
  1737. 1:00:11and now we should be able to draw it out
  1738. 1:00:12of o not of n
  1739. 1:00:14so let's see how that worked
  1740. 1:00:18there we go
  1741. 1:00:19n went through 10 h
  1742. 1:00:21to produce this output
  1743. 1:00:24so now tan h is a
  1744. 1:00:26sort of
  1745. 1:00:27our little micro grad supported node
  1746. 1:00:30here as an operation
  1747. 1:00:33and as long as we know the derivative of
  1748. 1:00:3510h
  1749. 1:00:36then we'll be able to back propagate
  1750. 1:00:37through it now let's see this 10h in
  1751. 1:00:39action currently it's not squashing too
  1752. 1:00:41much because the input to it is pretty
  1753. 1:00:43low so if the bias was increased to say
  1754. 1:00:46eight
  1755. 1:00:49then we'll see that what's flowing into
  1756. 1:00:51the 10h now is
  1757. 1:00:53two
  1758. 1:00:54and 10h is squashing it to 0.96 so we're
  1759. 1:00:57already hitting the tail of this 10h and
  1760. 1:00:59it will sort of smoothly go up to 1 and
  1761. 1:01:01then plateau out over there
  1762. 1:01:03okay so now i'm going to do something
  1763. 1:01:04slightly strange i'm going to change
  1764. 1:01:06this bias from 8 to this number
  1765. 1:01:096.88 etc
  1766. 1:01:11and i'm going to do this for specific
  1767. 1:01:13reasons because we're about to start
  1768. 1:01:15back propagation
  1769. 1:01:16and i want to make sure that our numbers
  1770. 1:01:19come out nice they're not like very
  1771. 1:01:21crazy numbers they're nice numbers that
  1772. 1:01:22we can sort of understand in our head
  1773. 1:01:24let me also add a pose label
  1774. 1:01:26o is short for output here
  1775. 1:01:30so that's zero
  1776. 1:01:31okay so
  1777. 1:01:320.88 flows into 10 h comes out 0.7 so on
  1778. 1:01:36so now we're going to do back
  1779. 1:01:37propagation and we're going to fill in
  1780. 1:01:38all the gradients
  1781. 1:01:40so what is the derivative o with respect
  1782. 1:01:43to
  1783. 1:01:44all the
  1784. 1:01:45inputs here and of course in the typical
  1785. 1:01:47neural network setting what we really
  1786. 1:01:48care about the most is the derivative of
  1787. 1:01:51these neurons on the weights
  1788. 1:01:53specifically the w2 and w1 because those
  1789. 1:01:56are the weights that we're going to be
  1790. 1:01:57changing part of the optimization
  1791. 1:01:59and the other thing that we have to
  1792. 1:02:00remember is here we have only a single
  1793. 1:02:02neuron but in the neural natives
  1794. 1:02:03typically have many neurons and they're
  1795. 1:02:04connected
  1796. 1:02:07so this is only like a one small neuron
  1797. 1:02:09a piece of a much bigger puzzle and
  1798. 1:02:10eventually there's a loss function that
  1799. 1:02:12sort of measures the accuracy of the
  1800. 1:02:13neural net and we're back propagating
  1801. 1:02:15with respect to that accuracy and trying
  1802. 1:02:16to increase it
  1803. 1:02:19so let's start off by propagation here
  1804. 1:02:21in the end
  1805. 1:02:22what is the derivative of o with respect
  1806. 1:02:24to o the base case sort of we know
  1807. 1:02:26always is that the gradient is just 1.0
  1808. 1:02:30so let me fill it in
  1809. 1:02:32and then let me
  1810. 1:02:35split out
  1811. 1:02:37the drawing function
  1812. 1:02:40here
  1813. 1:02:43and then here cell
  1814. 1:02:47clear this output here okay
  1815. 1:02:50so now when we draw o we'll see that oh
  1816. 1:02:52that grad is one
  1817. 1:02:53so now we're going to back propagate
  1818. 1:02:55through the tan h
  1819. 1:02:56so to back propagate through 10h we need
  1820. 1:02:58to know the local derivative of 10h
  1821. 1:03:01so if we have that
  1822. 1:03:03o is 10 h of
  1823. 1:03:07n
  1824. 1:03:08then what is d o by d n
  1825. 1:03:12now what you could do is you could come
  1826. 1:03:13here and you could take this expression
  1827. 1:03:15and you could
  1828. 1:03:16do your calculus derivative taking
  1829. 1:03:19um and that would work but we can also
  1830. 1:03:21just scroll down wikipedia here
  1831. 1:03:23into a section that hopefully tells us
  1832. 1:03:26that derivative uh
  1833. 1:03:28d by dx of 10 h of x is
  1834. 1:03:31any of these i like this one 1 minus 10
  1835. 1:03:33h square of x
  1836. 1:03:35so this is 1 minus 10 h
  1837. 1:03:37of x squared
  1838. 1:03:39so basically what this is saying is that
  1839. 1:03:41d o by d n
  1840. 1:03:43is
  1841. 1:03:441 minus 10 h
  1842. 1:03:47of n
  1843. 1:03:48squared
  1844. 1:03:51and we already have 10 h of n that's
  1845. 1:03:52just o
  1846. 1:03:54so it's one minus o squared
  1847. 1:03:56so o is the output here so the output is
  1848. 1:03:59this number
  1849. 1:04:02data
  1850. 1:04:04is this number
  1851. 1:04:06and then
  1852. 1:04:08what this is saying is that do by dn is
  1853. 1:04:101 minus
  1854. 1:04:11this squared so
  1855. 1:04:13one minus of that data squared
  1856. 1:04:16is 0.5 conveniently
  1857. 1:04:18so the local derivative of this 10 h
  1858. 1:04:21operation here is 0.5
  1859. 1:04:24and
  1860. 1:04:25so that would be d o by d n
  1861. 1:04:27so
  1862. 1:04:28we can fill in that in that grad
  1863. 1:04:33is 0.5 we'll just fill in
  1864. 1:04:42so this is exactly 0.5 one half
  1865. 1:04:45so now we're going to continue the back
  1866. 1:04:47propagation
  1867. 1:04:49this is 0.5 and this is a plus node
  1868. 1:04:52so how is backprop going to what is that
  1869. 1:04:55going to do here
  1870. 1:04:56and if you remember our previous example
  1871. 1:04:58a plus is just a distributor of gradient
  1872. 1:05:01so this gradient will simply flow to
  1873. 1:05:03both of these equally and that's because
  1874. 1:05:05the local derivative of this operation
  1875. 1:05:07is one for every one of its nodes so 1
  1876. 1:05:10times 0.5 is 0.5
  1877. 1:05:12so therefore we know that
  1878. 1:05:14this node here which we called this
  1879. 1:05:18its grad is just 0.5
  1880. 1:05:21and we know that b dot grad is also 0.5
  1881. 1:05:24so let's set those and let's draw
  1882. 1:05:28so 0.5
  1883. 1:05:30continuing we have another plus
  1884. 1:05:320.5 again we'll just distribute it so
  1885. 1:05:340.5 will flow to both of these
  1886. 1:05:37so we can set
  1887. 1:05:39theirs
  1888. 1:05:43x2w2 as well that grad is 0.5
  1889. 1:05:47and let's redraw pluses are my favorite
  1890. 1:05:50uh operations to back propagate through
  1891. 1:05:51because
  1892. 1:05:53it's very simple
  1893. 1:05:55so now it's flowing into these
  1894. 1:05:56expressions is 0.5 and so really again
  1895. 1:05:58keep in mind what the derivative is
  1896. 1:05:59telling us at every point in time along
  1897. 1:06:01here this is saying that
  1898. 1:06:04if we want the output of this neuron to
  1899. 1:06:06increase
  1900. 1:06:08then
  1901. 1:06:08the influence on these expressions is
  1902. 1:06:10positive on the output both of them are
  1903. 1:06:13positive
  1904. 1:06:16contribution to the output
  1905. 1:06:20so now back propagating to x2 and w2
  1906. 1:06:23first
  1907. 1:06:24this is a times node so we know that the
  1908. 1:06:26local derivative is you know the other
  1909. 1:06:28term
  1910. 1:06:28so if we want to calculate x2.grad
  1911. 1:06:32then
  1912. 1:06:33can you think through what it's going to
  1913. 1:06:34be
  1914. 1:06:40so x2.grad will be
  1915. 1:06:42w2.data
  1916. 1:06:44times this x2w2
  1917. 1:06:48by grad right
  1918. 1:06:51and
  1919. 1:06:52w2.grad will be
  1920. 1:06:55x2 that data times x2w2.grad
  1921. 1:07:01right so that's the local piece of chain
  1922. 1:07:03rule
  1923. 1:07:07let's set them and let's redraw
  1924. 1:07:09so here we see that the gradient on our
  1925. 1:07:11weight 2 is 0 because x2 data was 0
  1926. 1:07:15right but x2 will have the gradient 0.5
  1927. 1:07:18because data here was 1.
  1928. 1:07:20and so what's interesting here right is
  1929. 1:07:22because the input x2 was 0 then because
  1930. 1:07:25of the way the times works
  1931. 1:07:28of course this gradient will be zero and
  1932. 1:07:30think about intuitively why that is
  1933. 1:07:33derivative always tells us the influence
  1934. 1:07:35of
  1935. 1:07:36this on the final output if i wiggle w2
  1936. 1:07:39how is the output changing
  1937. 1:07:41it's not changing because we're
  1938. 1:07:42multiplying by zero
  1939. 1:07:44so because it's not changing there's no
  1940. 1:07:46derivative and zero is the correct
  1941. 1:07:47answer
  1942. 1:07:48because we're
  1943. 1:07:49squashing it at zero
  1944. 1:07:52and let's do it here point five should
  1945. 1:07:54come here and flow through this times
  1946. 1:07:57and so we'll have that x1.grad is
  1947. 1:08:01can you think through a little bit what
  1948. 1:08:03what
  1949. 1:08:04this should be
  1950. 1:08:07the local derivative of times
  1951. 1:08:09with respect to x1 is going to be w1
  1952. 1:08:12so w1 is data times
  1953. 1:08:15x1 w1 dot grad
  1954. 1:08:18and w1.grad will be x1.data times
  1955. 1:08:23x1 w2 w1 with graph
  1956. 1:08:27let's see what those came out to be
  1957. 1:08:29so this is 0.5 so this would be negative
  1958. 1:08:311.5 and this would be 1.
  1959. 1:08:34and we've back propagated through this
  1960. 1:08:36expression these are the actual final
  1961. 1:08:38derivatives so if we want this neuron's
  1962. 1:08:40output to increase
  1963. 1:08:43we know that what's necessary is that
  1964. 1:08:47w2 we have no gradient w2 doesn't
  1965. 1:08:49actually matter to this neuron right now
  1966. 1:08:51but this neuron this weight should uh go
  1967. 1:08:54up
  1968. 1:08:55so if this weight goes up then this
  1969. 1:08:57neuron's output would have gone up and
  1970. 1:08:59proportionally because the gradient is
  1971. 1:09:01one okay so doing the back propagation
  1972. 1:09:03manually is obviously ridiculous so we
  1973. 1:09:05are now going to put an end to this
  1974. 1:09:06suffering and we're going to see how we
  1975. 1:09:08can implement uh the backward pass a bit
  1976. 1:09:11more automatically we're not going to be
  1977. 1:09:12doing all of it manually out here
  1978. 1:09:14it's now pretty obvious to us by example
  1979. 1:09:17how these pluses and times are back
  1980. 1:09:18property ingredients so let's go up to
  1981. 1:09:20the value
  1982. 1:09:22object and we're going to start
  1983. 1:09:24codifying what we've seen
  1984. 1:09:27in the examples below
  1985. 1:09:29so we're going to do this by storing a
  1986. 1:09:31special cell dot backward
  1987. 1:09:34and underscore backward and this will be
  1988. 1:09:37a function which is going to do that
  1989. 1:09:39little piece of chain rule at each
  1990. 1:09:41little node that compute that took
  1991. 1:09:43inputs and produced output uh we're
  1992. 1:09:45going to store
  1993. 1:09:46how we are going to chain the the
  1994. 1:09:49outputs gradient into the inputs
  1995. 1:09:51gradients
  1996. 1:09:52so by default
  1997. 1:09:54this will be a function
  1998. 1:09:55that uh doesn't do anything
  1999. 1:09:58so um
  2000. 1:09:59and you can also see that here in the
  2001. 1:10:01value in micrograb
  2002. 1:10:03so
  2003. 1:10:04with this backward function by default
  2004. 1:10:06doesn't do anything
  2005. 1:10:08this is an empty function
  2006. 1:10:10and that would be sort of the case for
  2007. 1:10:11example for a leaf node for leaf node
  2008. 1:10:13there's nothing to do
  2009. 1:10:15but now if when we're creating these out
  2010. 1:10:18values these out values are an addition
  2011. 1:10:21of self and other
  2012. 1:10:24and so we will want to sell set
  2013. 1:10:27outs backward to be
  2014. 1:10:29the function that propagates the
  2015. 1:10:31gradient
  2016. 1:10:34so
  2017. 1:10:35let's define what should happen
  2018. 1:10:40and we're going to store it in a closure
  2019. 1:10:42let's define what should happen when we
  2020. 1:10:44call
  2021. 1:10:45outs grad
  2022. 1:10:47for in addition
  2023. 1:10:50our job is to take
  2024. 1:10:52outs grad and propagate it into self's
  2025. 1:10:55grad and other grad so basically we want
  2026. 1:10:57to sell self.grad to something
  2027. 1:11:00and we want to set others.grad to
  2028. 1:11:02something
  2029. 1:11:04okay
  2030. 1:11:05and the way we saw below how chain rule
  2031. 1:11:08works we want to take the local
  2032. 1:11:10derivative times
  2033. 1:11:11the
  2034. 1:11:12sort of global derivative i should call
  2035. 1:11:14it which is the derivative of the final
  2036. 1:11:16output of the expression with respect to
  2037. 1:11:18out's data
  2038. 1:11:21with respect to out
  2039. 1:11:22so
  2040. 1:11:24the local derivative of self in an
  2041. 1:11:27addition is 1.0
  2042. 1:11:29so it's just 1.0 times
  2043. 1:11:31outs grad
  2044. 1:11:34that's the chain rule
  2045. 1:11:35and others.grad will be 1.0 times
  2046. 1:11:38outgrad
  2047. 1:11:39and what you basically what you're
  2048. 1:11:40seeing here is that outscrad
  2049. 1:11:42will simply be copied onto selfs grad
  2050. 1:11:45and others grad as we saw happens for an
  2051. 1:11:48addition operation
  2052. 1:11:49so we're going to later call this
  2053. 1:11:51function to propagate the gradient
  2054. 1:11:53having done an addition
  2055. 1:11:55let's now do multiplication we're going
  2056. 1:11:57to also define that backward
  2057. 1:12:02and we're going to set its backward to
  2058. 1:12:04be backward
  2059. 1:12:07and we want to chain outgrad into
  2060. 1:12:11self.grad
  2061. 1:12:14and others.grad
  2062. 1:12:17and this will be a little piece of chain
  2063. 1:12:18rule for multiplication
  2064. 1:12:20so we'll have
  2065. 1:12:21so what should this be
  2066. 1:12:23can you think through
  2067. 1:12:28so what is the local derivative
  2068. 1:12:30here the local derivative was
  2069. 1:12:32others.data
  2070. 1:12:35and then
  2071. 1:12:36oops others.data and the times of that
  2072. 1:12:39grad that's channel
  2073. 1:12:42and here we have self.data times of that
  2074. 1:12:44grad
  2075. 1:12:45that's what we've been doing
  2076. 1:12:49and finally here for 10 h
  2077. 1:12:51left backward
  2078. 1:12:54and then we want to set out backwards to
  2079. 1:12:57be just backward
  2080. 1:13:00and here we need to
  2081. 1:13:02back propagate we have out that grad and
  2082. 1:13:04we want to chain it into self.grad
  2083. 1:13:09and salt.grad will be
  2084. 1:13:11the local derivative of this operation
  2085. 1:13:13that we've done here which is 10h
  2086. 1:13:16and so we saw that the local the
  2087. 1:13:17gradient is 1 minus the tan h of x
  2088. 1:13:20squared which here is t
  2089. 1:13:23that's the local derivative because
  2090. 1:13:25that's t is the output of this 10 h so 1
  2091. 1:13:27minus t squared is the local derivative
  2092. 1:13:30and then gradient um
  2093. 1:13:32has to be multiplied because of the
  2094. 1:13:33chain rule
  2095. 1:13:34so outgrad is chained through the local
  2096. 1:13:36gradient into salt.grad
  2097. 1:13:39and that should be basically it so we're
  2098. 1:13:41going to redefine our value node
  2099. 1:13:44we're going to swing all the way down
  2100. 1:13:46here
  2101. 1:13:48and we're going to
  2102. 1:13:49redefine
  2103. 1:13:51our expression
  2104. 1:13:52make sure that all the grads are zero
  2105. 1:13:55okay
  2106. 1:13:56but now we don't have to do this
  2107. 1:13:57manually anymore
  2108. 1:13:59we are going to basically be calling the
  2109. 1:14:01dot backward in the right order
  2110. 1:14:04so
  2111. 1:14:05first we want to call os
  2112. 1:14:07dot backwards
  2113. 1:14:14so o was the outcome of 10h
  2114. 1:14:17right so calling all that those who's
  2115. 1:14:20backward
  2116. 1:14:22will be
  2117. 1:14:23this function this is what it will do
  2118. 1:14:26now we have to be careful because
  2119. 1:14:29there's a times out.grad
  2120. 1:14:31and out.grad remember is initialized to
  2121. 1:14:34zero
  2122. 1:14:38so here we see grad zero so as a base
  2123. 1:14:41case we need to set both.grad to 1.0
  2124. 1:14:46to initialize this with 1
  2125. 1:14:53and then once this is 1 we can call oda
  2126. 1:14:56backward
  2127. 1:14:57and what that should do is it should
  2128. 1:14:58propagate this grad through 10h
  2129. 1:15:02so the local derivative times
  2130. 1:15:04the global derivative which is
  2131. 1:15:05initialized at one so
  2132. 1:15:08this should
  2133. 1:15:11um
  2134. 1:15:15a dope
  2135. 1:15:17so i thought about redoing it but i
  2136. 1:15:19figured i should just leave the error in
  2137. 1:15:20here because it's pretty funny why is
  2138. 1:15:22anti-object not callable
  2139. 1:15:24uh it's because
  2140. 1:15:27i screwed up we're trying to save these
  2141. 1:15:29functions so this is correct
  2142. 1:15:31this here
  2143. 1:15:33we don't want to call the function
  2144. 1:15:34because that returns none these
  2145. 1:15:36functions return none we just want to
  2146. 1:15:38store the function
  2147. 1:15:39so let me redefine the value object
  2148. 1:15:42and then we're going to come back in
  2149. 1:15:43redefine the expression draw a dot
  2150. 1:15:46everything is great o dot grad is one
  2151. 1:15:50o dot grad is one and now
  2152. 1:15:53now this should work of course
  2153. 1:15:55okay so all that backward should
  2154. 1:15:58this grant should now be 0.5 if we
  2155. 1:16:00redraw and if everything went correctly
  2156. 1:16:030.5 yay
  2157. 1:16:05okay so now we need to call ns.grad
  2158. 1:16:10and it's not awkward sorry
  2159. 1:16:13ends backward
  2160. 1:16:14so that seems to have worked
  2161. 1:16:17so instead backward routed the gradient
  2162. 1:16:21to both of these so this is looking
  2163. 1:16:22great
  2164. 1:16:24now we could of course called uh called
  2165. 1:16:26b grad
  2166. 1:16:27beat up backwards sorry
  2167. 1:16:30what's gonna happen
  2168. 1:16:32well b doesn't have it backward b is
  2169. 1:16:34backward
  2170. 1:16:35because b is a leaf node
  2171. 1:16:37b's backward is by initialization the
  2172. 1:16:40empty function
  2173. 1:16:41so nothing would happen but we can call
  2174. 1:16:44call it on it
  2175. 1:16:45but when we call
  2176. 1:16:48this one
  2177. 1:16:50it's backward
  2178. 1:16:53then we expect this 0.5 to get further
  2179. 1:16:56routed
  2180. 1:16:57right so there we go 0.5.5
  2181. 1:17:00and then finally
  2182. 1:17:02we want to call
  2183. 1:17:05it here on x2 w2
  2184. 1:17:10and on x1 w1
  2185. 1:17:16do both of those
  2186. 1:17:17and there we go
  2187. 1:17:19so we get 0 0.5 negative 1.5 and 1
  2188. 1:17:23exactly as we did before but now
  2189. 1:17:26we've done it through
  2190. 1:17:28calling that backward um
  2191. 1:17:30sort of manually
  2192. 1:17:32so we have the lamp one last piece to
  2193. 1:17:34get rid of which is us calling
  2194. 1:17:36underscore backward manually so let's
  2195. 1:17:38think through what we are actually doing
  2196. 1:17:40um
  2197. 1:17:41we've laid out a mathematical expression
  2198. 1:17:43and now we're trying to go backwards
  2199. 1:17:44through that expression
  2200. 1:17:46um so going backwards through the
  2201. 1:17:48expression just means that we never want
  2202. 1:17:50to call a dot backward for any node
  2203. 1:17:54before
  2204. 1:17:55we've done a sort of um everything after
  2205. 1:17:58it
  2206. 1:17:59so we have to do everything after it
  2207. 1:18:01before we're ever going to call that
  2208. 1:18:02backward on any one node we have to get
  2209. 1:18:04all of its full dependencies everything
  2210. 1:18:06that it depends on has to
  2211. 1:18:08propagate to it before we can continue
  2212. 1:18:10back propagation so this ordering of
  2213. 1:18:14graphs can be achieved using something
  2214. 1:18:16called topological sort
  2215. 1:18:17so topological sort
  2216. 1:18:20is basically a laying out of a graph
  2217. 1:18:23such that all the edges go only from
  2218. 1:18:24left to right basically
  2219. 1:18:26so here we have a graph it's a directory
  2220. 1:18:29a cyclic graph a dag
  2221. 1:18:31and this is two different topological
  2222. 1:18:34orders of it i believe where basically
  2223. 1:18:36you'll see that it's laying out of the
  2224. 1:18:37notes such that all the edges go only
  2225. 1:18:39one way from left to right
  2226. 1:18:41and implementing topological sort you
  2227. 1:18:44can look in wikipedia and so on i'm not
  2228. 1:18:46going to go through it in detail
  2229. 1:18:48but basically this is what builds a
  2230. 1:18:51topological graph
  2231. 1:18:54we maintain a set of visited nodes and
  2232. 1:18:56then we are
  2233. 1:18:59going through starting at some root node
  2234. 1:19:02which for us is o that's where we want
  2235. 1:19:03to start the topological sort
  2236. 1:19:05and starting at o we go through all of
  2237. 1:19:08its children and we need to lay them out
  2238. 1:19:10from left to right
  2239. 1:19:12and basically this starts at o
  2240. 1:19:14if it's not visited then it marks it as
  2241. 1:19:17visited and then it iterates through all
  2242. 1:19:19of its children
  2243. 1:19:20and calls build topological on them
  2244. 1:19:24and then uh after it's gone through all
  2245. 1:19:26the children it adds itself
  2246. 1:19:28so basically
  2247. 1:19:29this node that we're going to call it on
  2248. 1:19:31like say o is only going to add itself
  2249. 1:19:34to the topo list after all of the
  2250. 1:19:37children have been processed and that's
  2251. 1:19:39how this function is guaranteeing
  2252. 1:19:41that you're only going to be in the list
  2253. 1:19:43once all your children are in the list
  2254. 1:19:45and that's the invariant that is being
  2255. 1:19:46maintained so if we built upon o and
  2256. 1:19:49then inspect this list
  2257. 1:19:52we're going to see that it ordered our
  2258. 1:19:54value objects
  2259. 1:19:56and the last one
  2260. 1:19:58is the value of 0.707 which is the
  2261. 1:20:00output
  2262. 1:20:01so this is o and then this is n
  2263. 1:20:04and then all the other nodes get laid
  2264. 1:20:07out before it
  2265. 1:20:09so that builds the topological graph and
  2266. 1:20:12really what we're doing now is we're
  2267. 1:20:13just calling dot underscore backward on
  2268. 1:20:16all of the nodes in a topological order
  2269. 1:20:19so if we just reset the gradients
  2270. 1:20:22they're all zero
  2271. 1:20:23what did we do
  2272. 1:20:24we started by
  2273. 1:20:27setting o dot grad
  2274. 1:20:29to b1
  2275. 1:20:31that's the base case
  2276. 1:20:33then we built the topological order
  2277. 1:20:38and then we went for node
  2278. 1:20:41in
  2279. 1:20:42reversed
  2280. 1:20:44of topo
  2281. 1:20:46now
  2282. 1:20:47in in the reverse order because this
  2283. 1:20:49list goes from
  2284. 1:20:50you know we need to go through it in
  2285. 1:20:52reversed order
  2286. 1:20:53so starting at o
  2287. 1:20:56note that backward
  2288. 1:20:58and this should be
  2289. 1:21:01it
  2290. 1:21:03there we go
  2291. 1:21:05those are the correct derivatives
  2292. 1:21:07finally we are going to hide this
  2293. 1:21:08functionality
  2294. 1:21:10so i'm going to
  2295. 1:21:11copy this and we're going to hide it
  2296. 1:21:13inside the valley class because we don't
  2297. 1:21:15want to have all that code lying around
  2298. 1:21:18so instead of an underscore backward
  2299. 1:21:19we're now going to define an actual
  2300. 1:21:21backward so that's backward without the
  2301. 1:21:23underscore
  2302. 1:21:26and that's going to do all the stuff
  2303. 1:21:27that we just arrived
  2304. 1:21:29so let me just clean this up a little
  2305. 1:21:30bit so
  2306. 1:21:32we're first going to
  2307. 1:21:37build a topological graph
  2308. 1:21:38starting at self
  2309. 1:21:41so build topo of self
  2310. 1:21:44will populate the topological order into
  2311. 1:21:46the topo list which is a local variable
  2312. 1:21:49then we set self.grad to be one
  2313. 1:21:52and then for each node in the reversed
  2314. 1:21:55list so starting at us and going to all
  2315. 1:21:57the children
  2316. 1:22:00underscore backward
  2317. 1:22:02and
  2318. 1:22:03that should be it so
  2319. 1:22:06save
  2320. 1:22:08come down here
  2321. 1:22:09redefine
  2322. 1:22:09[Music]
  2323. 1:22:11okay all the grands are zero
  2324. 1:22:13and now what we can do is oh that
  2325. 1:22:15backward without the underscore
  2326. 1:22:17and
  2327. 1:22:21there we go
  2328. 1:22:22and that's uh that's back propagation
  2329. 1:22:26place for one neuron
  2330. 1:22:28now we shouldn't be too happy with
  2331. 1:22:29ourselves actually because we have a bad
  2332. 1:22:32bug um and we have not surfaced the bug
  2333. 1:22:35because of some specific conditions that
  2334. 1:22:36we are we have to think about right now
  2335. 1:22:39so here's the simplest case that shows
  2336. 1:22:42the bug
  2337. 1:22:43say i create a single node a
  2338. 1:22:48and then i create a b that is a plus a
  2339. 1:22:51and then i called backward
  2340. 1:22:54so what's going to happen is a is 3
  2341. 1:22:57and then a b is a plus a so there's two
  2342. 1:23:00arrows on top of each other here
  2343. 1:23:03then we can see that b is of course the
  2344. 1:23:05forward pass works
  2345. 1:23:06b is just
  2346. 1:23:08a plus a which is six
  2347. 1:23:10but the gradient here is not actually
  2348. 1:23:11correct
  2349. 1:23:12that we calculate it automatically
  2350. 1:23:15and that's because
  2351. 1:23:17um
  2352. 1:23:19of course uh
  2353. 1:23:20just doing calculus in your head the
  2354. 1:23:22derivative of b with respect to a
  2355. 1:23:24should be uh two
  2356. 1:23:27one plus one
  2357. 1:23:28it's not one
  2358. 1:23:30intuitively what's happening here right
  2359. 1:23:32so b is the result of a plus a and then
  2360. 1:23:34we call backward on it
  2361. 1:23:36so let's go up and see what that does
  2362. 1:23:42um
  2363. 1:23:43b is a result of addition
  2364. 1:23:45so out as
  2365. 1:23:46b and then when we called backward what
  2366. 1:23:49happened is
  2367. 1:23:50self.grad was set
  2368. 1:23:53to one
  2369. 1:23:54and then other that grad was set to one
  2370. 1:23:57but because we're doing a plus a
  2371. 1:23:59self and other are actually the exact
  2372. 1:24:02same object
  2373. 1:24:03so we are overriding the gradient we are
  2374. 1:24:06setting it to one and then we are
  2375. 1:24:07setting it again to one and that's why
  2376. 1:24:10it stays
  2377. 1:24:11at one
  2378. 1:24:13so that's a problem
  2379. 1:24:14there's another way to see this in a
  2380. 1:24:16little bit more complicated expression
  2381. 1:24:21so here we have
  2382. 1:24:23a and b
  2383. 1:24:25and then uh d will be the multiplication
  2384. 1:24:28of the two and e will be the addition of
  2385. 1:24:30the two
  2386. 1:24:32and
  2387. 1:24:33then we multiply e times d to get f and
  2388. 1:24:35then we called fda backward
  2389. 1:24:37and these gradients if you check will be
  2390. 1:24:39incorrect
  2391. 1:24:40so fundamentally what's happening here
  2392. 1:24:42again is
  2393. 1:24:45basically we're going to see an issue
  2394. 1:24:46anytime we use a variable more than once
  2395. 1:24:49until now in these expressions above
  2396. 1:24:51every variable is used exactly once so
  2397. 1:24:53we didn't see the issue
  2398. 1:24:54but here if a variable is used more than
  2399. 1:24:56once what's going to happen during
  2400. 1:24:57backward pass we're backpropagating from
  2401. 1:25:00f to e to d so far so good but now
  2402. 1:25:03equals it backward and it deposits its
  2403. 1:25:05gradients to a and b but then we come
  2404. 1:25:08back to d
  2405. 1:25:09and call backward and it overwrites
  2406. 1:25:11those gradients at a and b
  2407. 1:25:14so that's obviously a problem
  2408. 1:25:17and the solution here if you look at
  2409. 1:25:19the multivariate case of the chain rule
  2410. 1:25:22and its generalization there
  2411. 1:25:23the solution there is basically that we
  2412. 1:25:26have to accumulate these gradients these
  2413. 1:25:28gradients add
  2414. 1:25:30and so instead of setting those
  2415. 1:25:32gradients
  2416. 1:25:34we can simply do plus equals we need to
  2417. 1:25:37accumulate those gradients
  2418. 1:25:39plus equals plus equals
  2419. 1:25:41plus equals
  2420. 1:25:44plus equals
  2421. 1:25:46and this will be okay remember because
  2422. 1:25:48we are initializing them at zero so they
  2423. 1:25:50start at zero
  2424. 1:25:51and then any
  2425. 1:25:53contribution
  2426. 1:25:54that flows backwards
  2427. 1:25:57will simply add
  2428. 1:25:58so now if we redefine
  2429. 1:26:01this one
  2430. 1:26:03because the plus equals this now works
  2431. 1:26:06because a.grad started at zero and we
  2432. 1:26:08called beta backward we deposit one and
  2433. 1:26:11then we deposit one again and now this
  2434. 1:26:13is two which is correct
  2435. 1:26:14and here this will also work and we'll
  2436. 1:26:16get correct gradients
  2437. 1:26:18because when we call eta backward we
  2438. 1:26:20will deposit the gradients from this
  2439. 1:26:21branch and then we get to back into
  2440. 1:26:23detail backward it will deposit its own
  2441. 1:26:26gradients and then those gradients
  2442. 1:26:28simply add on top of each other and so
  2443. 1:26:30we just accumulate those gradients and
  2444. 1:26:31that fixes the issue okay now before we
  2445. 1:26:34move on let me actually do a bit of
  2446. 1:26:35cleanup here and delete some of these
  2447. 1:26:38some of this intermediate work so
  2448. 1:26:41we're not gonna need any of this now
  2449. 1:26:42that we've derived all of it
  2450. 1:26:44um
  2451. 1:26:45we are going to keep this because i want
  2452. 1:26:48to come back to it
  2453. 1:26:49delete the 10h
  2454. 1:26:51delete our morning example
  2455. 1:26:53delete the step
  2456. 1:26:55delete this keep the code that draws
  2457. 1:26:59and then delete this example
  2458. 1:27:02and leave behind only the definition of
  2459. 1:27:03value
  2460. 1:27:05and now let's come back to this
  2461. 1:27:06non-linearity here that we implemented
  2462. 1:27:08the tanh now i told you that we could
  2463. 1:27:10have broken down 10h into its explicit
  2464. 1:27:13atoms in terms of other expressions if
  2465. 1:27:16we had the x function so if you remember
  2466. 1:27:18tan h is defined like this and we chose
  2467. 1:27:20to develop tan h as a single function
  2468. 1:27:22and we can do that because we know its
  2469. 1:27:24derivative and we can back propagate
  2470. 1:27:26through it
  2471. 1:27:26but we can also break down tan h into
  2472. 1:27:29and express it as a function of x and i
  2473. 1:27:31would like to do that now because i want
  2474. 1:27:33to prove to you that you get all the
  2475. 1:27:34same results and all those ingredients
  2476. 1:27:36but also because it forces us to
  2477. 1:27:38implement a few more expressions it
  2478. 1:27:40forces us to do exponentiation addition
  2479. 1:27:42subtraction division and things like
  2480. 1:27:44that and i think it's a good exercise to
  2481. 1:27:46go through a few more of these
  2482. 1:27:48okay so let's scroll up
  2483. 1:27:50to the definition of value
  2484. 1:27:52and here one thing that we currently
  2485. 1:27:53can't do is we can do like a value of
  2486. 1:27:56say 2.0
  2487. 1:27:58but we can't do you know here for
  2488. 1:28:00example we want to add constant one and
  2489. 1:28:02we can't do something like this
  2490. 1:28:05and we can't do it because it says
  2491. 1:28:06object has no attribute data that's
  2492. 1:28:08because a plus one comes right here to
  2493. 1:28:11add
  2494. 1:28:12and then other is the integer one and
  2495. 1:28:14then here python is trying to access
  2496. 1:28:16one.data and that's not a thing and
  2497. 1:28:18that's because basically one is not a
  2498. 1:28:20value object and we only have addition
  2499. 1:28:22for value objects so as a matter of
  2500. 1:28:24convenience so that we can create
  2501. 1:28:26expressions like this and make them make
  2502. 1:28:28sense
  2503. 1:28:29we can simply do something like this
  2504. 1:28:32basically
  2505. 1:28:33we let other alone if other is an
  2506. 1:28:35instance of value but if it's not an
  2507. 1:28:37instance of value we're going to assume
  2508. 1:28:39that it's a number like an integer float
  2509. 1:28:40and we're going to simply wrap it in in
  2510. 1:28:43value and then other will just become
  2511. 1:28:45value of other and then other will have
  2512. 1:28:46a data attribute and this should work so
  2513. 1:28:49if i just say this predefined value then
  2514. 1:28:51this should work
  2515. 1:28:53there we go okay now let's do the exact
  2516. 1:28:55same thing for multiply because we can't
  2517. 1:28:57do something like this
  2518. 1:28:58again
  2519. 1:28:59for the exact same reason so we just
  2520. 1:29:01have to go to mole and if other is
  2521. 1:29:04not a value then let's wrap it in value
  2522. 1:29:07let's redefine value and now this works
  2523. 1:29:10now here's a kind of unfortunate and not
  2524. 1:29:12obvious part a times two works we saw
  2525. 1:29:15that but two times a is that gonna work
  2526. 1:29:19you'd expect it to right but actually it
  2527. 1:29:21will not
  2528. 1:29:22and the reason it won't is because
  2529. 1:29:24python doesn't know
  2530. 1:29:26like when when you do a times two
  2531. 1:29:28basically um so a times two python will
  2532. 1:29:31go and it will basically do something
  2533. 1:29:32like a dot mul
  2534. 1:29:34of two that's basically what it will
  2535. 1:29:36call but to it 2 times a is the same as
  2536. 1:29:392 dot mol of a
  2537. 1:29:41and it doesn't 2 can't multiply
  2538. 1:29:44value and so it's really confused about
  2539. 1:29:46that
  2540. 1:29:47so instead what happens is in python the
  2541. 1:29:49way this works is you are free to define
  2542. 1:29:51something called the r mold
  2543. 1:29:54and our mole
  2544. 1:29:55is kind of like a fallback so if python
  2545. 1:29:58can't do 2 times a it will check if um
  2546. 1:30:02if by any chance a knows how to multiply
  2547. 1:30:05two and that will be called into our
  2548. 1:30:07mole
  2549. 1:30:08so because python can't do two times a
  2550. 1:30:11it will check is there an our mole in
  2551. 1:30:12value and because there is it will now
  2552. 1:30:15call that
  2553. 1:30:16and what we'll do here is we will swap
  2554. 1:30:18the order of the operands so basically
  2555. 1:30:21two times a will redirect to armel and
  2556. 1:30:23our mole will basically call a times two
  2557. 1:30:26and that's how that will work
  2558. 1:30:28so
  2559. 1:30:29redefining now with armor two times a
  2560. 1:30:31becomes four okay now looking at the
  2561. 1:30:33other elements that we still need we
  2562. 1:30:34need to know how to exponentiate and how
  2563. 1:30:36to divide so let's first the explanation
  2564. 1:30:38to the exponentiation part we're going
  2565. 1:30:40to introduce
  2566. 1:30:41a single
  2567. 1:30:42function x here
  2568. 1:30:45and x is going to mirror 10h in the
  2569. 1:30:47sense that it's a simple single function
  2570. 1:30:49that transforms a single scalar value
  2571. 1:30:51and outputs a single scalar value
  2572. 1:30:53so we pop out the python number we use
  2573. 1:30:56math.x to exponentiate it create a new
  2574. 1:30:58value object
  2575. 1:30:59everything that we've seen before the
  2576. 1:31:00tricky part of course is how do you
  2577. 1:31:02propagate through e to the x
  2578. 1:31:04and
  2579. 1:31:05so here you can potentially pause the
  2580. 1:31:07video and think about what should go
  2581. 1:31:09here
  2582. 1:31:13okay so basically we need to know what
  2583. 1:31:15is the local derivative of e to the x so
  2584. 1:31:18d by d x of e to the x is famously just
  2585. 1:31:21e to the x and we've already just
  2586. 1:31:23calculated e to the x and it's inside
  2587. 1:31:25out that data so we can do up that data
  2588. 1:31:27times
  2589. 1:31:28and
  2590. 1:31:29out that grad that's the chain rule
  2591. 1:31:32so we're just chaining on to the current
  2592. 1:31:33running grad
  2593. 1:31:35and this is what the expression looks
  2594. 1:31:36like it looks a little confusing but
  2595. 1:31:38this is what it is and that's the
  2596. 1:31:40exponentiation
  2597. 1:31:41so redefining we should now be able to
  2598. 1:31:43call a.x
  2599. 1:31:45and
  2600. 1:31:46hopefully the backward pass works as
  2601. 1:31:47well okay and the last thing we'd like
  2602. 1:31:49to do of course is we'd like to be able
  2603. 1:31:50to divide
  2604. 1:31:52now
  2605. 1:31:53i actually will implement something
  2606. 1:31:54slightly more powerful than division
  2607. 1:31:56because division is just a special case
  2608. 1:31:57of something a bit more powerful
  2609. 1:31:59so in particular just by rearranging
  2610. 1:32:02if we have some kind of a b equals
  2611. 1:32:04value of 4.0 here we'd like to basically
  2612. 1:32:07be able to do a divide b and we'd like
  2613. 1:32:09this to be able to give us 0.5
  2614. 1:32:11now division actually can be reshuffled
  2615. 1:32:14as follows if we have a divide b that's
  2616. 1:32:17actually the same as a multiplying one
  2617. 1:32:18over b
  2618. 1:32:19and that's the same as a multiplying b
  2619. 1:32:21to the power of negative one
  2620. 1:32:24and so what i'd like to do instead is i
  2621. 1:32:25basically like to implement the
  2622. 1:32:27operation of x to the k for some
  2623. 1:32:29constant uh k so it's an integer or a
  2624. 1:32:32float um and we would like to be able to
  2625. 1:32:35differentiate this and then as a special
  2626. 1:32:36case uh negative one will be division
  2627. 1:32:40and so i'm doing that just because uh
  2628. 1:32:42it's more general and um yeah you might
  2629. 1:32:45as well do it that way so basically what
  2630. 1:32:46i'm saying is we can redefine
  2631. 1:32:49uh division
  2632. 1:32:51which we will put here somewhere
  2633. 1:32:54yeah we can put it here somewhere what
  2634. 1:32:56i'm saying is that we can redefine
  2635. 1:32:58division so self-divide other
  2636. 1:33:00can actually be rewritten as self times
  2637. 1:33:03other to the power of negative one
  2638. 1:33:05and now
  2639. 1:33:07a value raised to the power of negative
  2640. 1:33:09one we have now defined that
  2641. 1:33:11so
  2642. 1:33:12here's
  2643. 1:33:13so we need to implement the pow function
  2644. 1:33:15where am i going to put the power
  2645. 1:33:17function maybe here somewhere
  2646. 1:33:20this is the skeleton for it
  2647. 1:33:22so this function will be called when we
  2648. 1:33:24try to raise a value to some power and
  2649. 1:33:26other will be that power
  2650. 1:33:28now i'd like to make sure that other is
  2651. 1:33:30only an int or a float usually other is
  2652. 1:33:33some kind of a different value object
  2653. 1:33:35but here other will be forced to be an
  2654. 1:33:37end or a float otherwise the math
  2655. 1:33:40won't work for
  2656. 1:33:42for or try to achieve in the specific
  2657. 1:33:43case that would be a different
  2658. 1:33:45derivative expression if we wanted other
  2659. 1:33:47to be a value
  2660. 1:33:49so here we create the output value which
  2661. 1:33:51is just uh you know this data raised to
  2662. 1:33:53the power of other and other here could
  2663. 1:33:55be for example negative one that's what
  2664. 1:33:56we are hoping to achieve
  2665. 1:33:59and then uh this is the backwards stub
  2666. 1:34:01and this is the fun part which is what
  2667. 1:34:03is the uh chain rule expression here for
  2668. 1:34:07back for um
  2669. 1:34:09back propagating through the power
  2670. 1:34:11function where the power is to the power
  2671. 1:34:13of some kind of a constant
  2672. 1:34:15so this is the exercise and maybe pause
  2673. 1:34:17the video here and see if you can figure
  2674. 1:34:18it out yourself as to what we should put
  2675. 1:34:20here
  2676. 1:34:26okay so
  2677. 1:34:29you can actually go here and look at
  2678. 1:34:30derivative rules as an example and we
  2679. 1:34:32see lots of derivatives that you can
  2680. 1:34:34hopefully know from calculus in
  2681. 1:34:36particular what we're looking for is the
  2682. 1:34:37power rule
  2683. 1:34:39because that's telling us that if we're
  2684. 1:34:40trying to take d by dx of x to the n
  2685. 1:34:42which is what we're doing here
  2686. 1:34:44then that is just n times x to the n
  2687. 1:34:46minus 1
  2688. 1:34:48right
  2689. 1:34:49okay
  2690. 1:34:50so
  2691. 1:34:51that's telling us about the local
  2692. 1:34:53derivative of this power operation
  2693. 1:34:55so all we want here
  2694. 1:34:58basically n is now other
  2695. 1:35:00and self.data is x
  2696. 1:35:03and so this now becomes
  2697. 1:35:06other which is n times
  2698. 1:35:08self.data
  2699. 1:35:10which is now a python in torah float
  2700. 1:35:13it's not a valley object we're accessing
  2701. 1:35:14the data attribute
  2702. 1:35:16raised
  2703. 1:35:17to the power of other minus one or n
  2704. 1:35:19minus one
  2705. 1:35:21i can put brackets around this but this
  2706. 1:35:22doesn't matter because
  2707. 1:35:25power takes precedence over multiply and
  2708. 1:35:27python so that would have been okay
  2709. 1:35:29and that's the local derivative only but
  2710. 1:35:31now we have to chain it and we change
  2711. 1:35:33just simply by multiplying by output
  2712. 1:35:34grad that's chain rule
  2713. 1:35:36and this should technically work
  2714. 1:35:40and we're going to find out soon but now
  2715. 1:35:42if we
  2716. 1:35:43do this this should now work
  2717. 1:35:46and we get 0.5 so the forward pass works
  2718. 1:35:49but does the backward pass work and i
  2719. 1:35:51realize that we actually also have to
  2720. 1:35:52know how to subtract so
  2721. 1:35:54right now a minus b will not work
  2722. 1:35:57to make it work we need one more
  2723. 1:36:00piece of code here
  2724. 1:36:01and
  2725. 1:36:02basically this is the
  2726. 1:36:05subtraction and the way we're going to
  2727. 1:36:06implement subtraction is we're going to
  2728. 1:36:08implement it by addition of a negation
  2729. 1:36:10and then to implement negation we're
  2730. 1:36:12gonna multiply by negative one so just
  2731. 1:36:14again using the stuff we've already
  2732. 1:36:15built and just um expressing it in terms
  2733. 1:36:17of what we have and a minus b is now
  2734. 1:36:20working okay so now let's scroll again
  2735. 1:36:22to this expression here for this neuron
  2736. 1:36:25and let's just
  2737. 1:36:26compute the backward pass here once
  2738. 1:36:28we've defined o
  2739. 1:36:30and let's draw it
  2740. 1:36:32so here's the gradients for all these
  2741. 1:36:33leaf nodes for this two-dimensional
  2742. 1:36:35neuron that has a 10h that we've seen
  2743. 1:36:37before so now what i'd like to do is i'd
  2744. 1:36:39like to break up this 10h
  2745. 1:36:41into this expression here
  2746. 1:36:44so let me copy paste this
  2747. 1:36:46here
  2748. 1:36:47and now instead of we'll preserve the
  2749. 1:36:49label
  2750. 1:36:50and we will change how we define o
  2751. 1:36:53so in particular we're going to
  2752. 1:36:55implement this formula here
  2753. 1:36:56so we need e to the 2x
  2754. 1:36:58minus 1 over e to the x plus 1. so e to
  2755. 1:37:01the 2x we need to take 2 times n and we
  2756. 1:37:04need to exponentiate it that's e to the
  2757. 1:37:07two x and then because we're using it
  2758. 1:37:08twice let's create an intermediate
  2759. 1:37:10variable e
  2760. 1:37:12and then define o as
  2761. 1:37:14e plus one over
  2762. 1:37:16e minus one over e plus one
  2763. 1:37:19e minus one over e plus one
  2764. 1:37:22and that should be it and then we should
  2765. 1:37:24be able to draw that of o
  2766. 1:37:26so now before i run this what do we
  2767. 1:37:29expect to see
  2768. 1:37:30number one we're expecting to see a much
  2769. 1:37:32longer
  2770. 1:37:33graph here because we've broken up 10h
  2771. 1:37:35into a bunch of other operations
  2772. 1:37:37but those operations are mathematically
  2773. 1:37:39equivalent and so what we're expecting
  2774. 1:37:41to see is number one the same
  2775. 1:37:43result here so the forward pass works
  2776. 1:37:45and number two because of that
  2777. 1:37:47mathematical equivalence we expect to
  2778. 1:37:49see the same backward pass and the same
  2779. 1:37:51gradients on these leaf nodes so these
  2780. 1:37:53gradients should be identical
  2781. 1:37:55so let's run this
  2782. 1:37:58so number one let's verify that instead
  2783. 1:38:00of a single 10h node we have now x and
  2784. 1:38:03we have plus we have times negative one
  2785. 1:38:06uh this is the division
  2786. 1:38:08and we end up with the same forward pass
  2787. 1:38:10here
  2788. 1:38:11and then the gradients we have to be
  2789. 1:38:13careful because they're in slightly
  2790. 1:38:14different order potentially the
  2791. 1:38:16gradients for w2x2 should be 0 and 0.5
  2792. 1:38:19w2 and x2 are 0 and 0.5
  2793. 1:38:22and w1 x1 are 1 and negative 1.5
  2794. 1:38:251 and negative 1.5
  2795. 1:38:27so that means that both our forward
  2796. 1:38:28passes and backward passes were correct
  2797. 1:38:31because this turned out to be equivalent
  2798. 1:38:33to
  2799. 1:38:3410h before
  2800. 1:38:35and so the reason i wanted to go through
  2801. 1:38:37this exercise is number one we got to
  2802. 1:38:39practice a few more operations and uh
  2803. 1:38:41writing more backwards passes and number
  2804. 1:38:43two i wanted to illustrate the point
  2805. 1:38:45that
  2806. 1:38:46the um
  2807. 1:38:47the level at which you implement your
  2808. 1:38:49operations is totally up to you you can
  2809. 1:38:51implement backward passes for tiny
  2810. 1:38:53expressions like a single individual
  2811. 1:38:54plus or a single times
  2812. 1:38:56or you can implement them for say
  2813. 1:38:5810h
  2814. 1:39:00which is a kind of a potentially you can
  2815. 1:39:01see it as a composite operation because
  2816. 1:39:03it's made up of all these more atomic
  2817. 1:39:05operations but really all of this is
  2818. 1:39:07kind of like a fake concept all that
  2819. 1:39:08matters is we have some kind of inputs
  2820. 1:39:10and some kind of an output and this
  2821. 1:39:11output is a function of the inputs in
  2822. 1:39:13some way and as long as you can do
  2823. 1:39:14forward pass and the backward pass of
  2824. 1:39:16that little operation it doesn't matter
  2825. 1:39:19what that operation is
  2826. 1:39:21and how composite it is
  2827. 1:39:23if you can write the local gradients you
  2828. 1:39:24can chain the gradient and you can
  2829. 1:39:26continue back propagation so the design
  2830. 1:39:28of what those functions are is
  2831. 1:39:30completely up to you
  2832. 1:39:31so now i would like to show you how you
  2833. 1:39:33can do the exact same thing by using a
  2834. 1:39:35modern deep neural network library like
  2835. 1:39:37for example pytorch which i've roughly
  2836. 1:39:40modeled micrograd
  2837. 1:39:41by
  2838. 1:39:42and so
  2839. 1:39:43pytorch is something you would use in
  2840. 1:39:44production and i'll show you how you can
  2841. 1:39:46do the exact same thing but in pytorch
  2842. 1:39:48api so i'm just going to copy paste it
  2843. 1:39:50in and walk you through it a little bit
  2844. 1:39:52this is what it looks like
  2845. 1:39:54so we're going to import pi torch and
  2846. 1:39:56then we need to define these
  2847. 1:39:59value objects like we have here
  2848. 1:40:01now micrograd is a scalar valued
  2849. 1:40:04engine so we only have scalar values
  2850. 1:40:07like 2.0 but in pi torch everything is
  2851. 1:40:10based around tensors and like i
  2852. 1:40:11mentioned tensors are just n-dimensional
  2853. 1:40:13arrays of scalars
  2854. 1:40:15so that's why things get a little bit
  2855. 1:40:17more complicated here i just need a
  2856. 1:40:19scalar value to tensor a tensor with
  2857. 1:40:21just a single element
  2858. 1:40:23but by default when you work with
  2859. 1:40:25pytorch you would use um
  2860. 1:40:28more complicated tensors like this so if
  2861. 1:40:30i import pytorch
  2862. 1:40:33then i can create tensors like this and
  2863. 1:40:36this tensor for example is a two by
  2864. 1:40:38three array
  2865. 1:40:39of scalar
  2866. 1:40:41scalars
  2867. 1:40:42in a single compact representation so we
  2868. 1:40:45can check its shape we see that it's a
  2869. 1:40:46two by three array
  2870. 1:40:48and so on
  2871. 1:40:49so this is usually what you would work
  2872. 1:40:50with um in the actual libraries so here
  2873. 1:40:54i'm creating
  2874. 1:40:55a tensor that has only a single element
  2875. 1:40:582.0
  2876. 1:41:00and then i'm casting it to be double
  2877. 1:41:03because python is by default using
  2878. 1:41:05double precision for its floating point
  2879. 1:41:07numbers so i'd like everything to be
  2880. 1:41:08identical by default the data type of
  2881. 1:41:12these tensors will be float32 so it's
  2882. 1:41:14only using a single precision float so
  2883. 1:41:16i'm casting it to double
  2884. 1:41:18so that we have float64 just like in
  2885. 1:41:21python
  2886. 1:41:22so i'm casting to double and then we get
  2887. 1:41:24something similar to value of two the
  2888. 1:41:28next thing i have to do is because these
  2889. 1:41:29are leaf nodes by default pytorch
  2890. 1:41:31assumes that they do not require
  2891. 1:41:32gradients so i need to explicitly say
  2892. 1:41:35that all of these nodes require
  2893. 1:41:36gradients
  2894. 1:41:37okay so this is going to construct
  2895. 1:41:39scalar valued one element tensors
  2896. 1:41:43make sure that fighters knows that they
  2897. 1:41:44require gradients now by default these
  2898. 1:41:47are set to false by the way because of
  2899. 1:41:48efficiency reasons because usually you
  2900. 1:41:50would not want gradients for leaf nodes
  2901. 1:41:53like the inputs to the network and this
  2902. 1:41:55is just trying to be efficient in the
  2903. 1:41:57most common cases
  2904. 1:41:59so once we've defined all of our values
  2905. 1:42:01in python we can perform arithmetic just
  2906. 1:42:03like we can here in microgradlend so
  2907. 1:42:06this will just work and then there's a
  2908. 1:42:07torch.10h also
  2909. 1:42:09and when we get back is a tensor again
  2910. 1:42:12and we can
  2911. 1:42:13just like in micrograd it's got a data
  2912. 1:42:15attribute and it's got grant attributes
  2913. 1:42:18so these tensor objects just like in
  2914. 1:42:19micrograd have a dot data and a dot grad
  2915. 1:42:22and
  2916. 1:42:23the only difference here is that we need
  2917. 1:42:25to call it that item because otherwise
  2918. 1:42:28um pi torch
  2919. 1:42:30that item basically takes
  2920. 1:42:32a single tensor of one element and it
  2921. 1:42:34just returns that element stripping out
  2922. 1:42:36the tensor
  2923. 1:42:37so let me just run this and hopefully we
  2924. 1:42:39are going to get this is going to print
  2925. 1:42:41the forward pass
  2926. 1:42:42which is 0.707
  2927. 1:42:44and this will be the gradients which
  2928. 1:42:46hopefully are
  2929. 1:42:480.5 0 negative 1.5 and 1.
  2930. 1:42:51so if we just run this
  2931. 1:42:53there we go
  2932. 1:42:540.7 so the forward pass agrees and then
  2933. 1:42:57point five zero negative one point five
  2934. 1:42:59and one
  2935. 1:43:00so pi torch agrees with us
  2936. 1:43:02and just to show you here basically o
  2937. 1:43:05here's a tensor with a single element
  2938. 1:43:08and it's a double
  2939. 1:43:09and we can call that item on it to just
  2940. 1:43:12get the single number out
  2941. 1:43:14so that's what item does and o is a
  2942. 1:43:16tensor object like i mentioned and it's
  2943. 1:43:18got a backward function just like we've
  2944. 1:43:20implemented
  2945. 1:43:22and then all of these also have a dot
  2946. 1:43:23graph so like x2 for example in the grad
  2947. 1:43:26and it's a tensor and we can pop out the
  2948. 1:43:28individual number with that actin
  2949. 1:43:31so basically
  2950. 1:43:32torches torch can do what we did in
  2951. 1:43:35micrograph is a special case when your
  2952. 1:43:37tensors are all single element tensors
  2953. 1:43:40but the big deal with pytorch is that
  2954. 1:43:42everything is significantly more
  2955. 1:43:43efficient because we are working with
  2956. 1:43:45these tensor objects and we can do lots
  2957. 1:43:47of operations in parallel on all of
  2958. 1:43:49these tensors
  2959. 1:43:51but otherwise what we've built very much
  2960. 1:43:53agrees with the api of pytorch
  2961. 1:43:55okay so now that we have some machinery
  2962. 1:43:57to build out pretty complicated
  2963. 1:43:58mathematical expressions we can also
  2964. 1:44:00start building out neural nets and as i
  2965. 1:44:02mentioned neural nets are just a
  2966. 1:44:03specific class of mathematical
  2967. 1:44:05expressions
  2968. 1:44:07so we're going to start building out a
  2969. 1:44:08neural net piece by piece and eventually
  2970. 1:44:09we'll build out a two-layer multi-layer
  2971. 1:44:12layer perceptron as it's called and i'll
  2972. 1:44:14show you exactly what that means
  2973. 1:44:15let's start with a single individual
  2974. 1:44:17neuron we've implemented one here but
  2975. 1:44:19here i'm going to implement one that
  2976. 1:44:21also subscribes to the pytorch api in
  2977. 1:44:24how it designs its neural network
  2978. 1:44:26modules
  2979. 1:44:27so just like we saw that we can like
  2980. 1:44:28match the api of pytorch
  2981. 1:44:31on the auto grad side we're going to try
  2982. 1:44:33to do that on the neural network modules
  2983. 1:44:35so here's class neuron
  2984. 1:44:38and just for the sake of efficiency i'm
  2985. 1:44:40going to copy paste some sections that
  2986. 1:44:42are relatively straightforward
  2987. 1:44:45so the constructor will take
  2988. 1:44:47number of inputs to this neuron which is
  2989. 1:44:49how many inputs come to a neuron so this
  2990. 1:44:52one for example has three inputs
  2991. 1:44:55and then it's going to create a weight
  2992. 1:44:57there is some random number between
  2993. 1:44:58negative one and one for every one of
  2994. 1:45:00those inputs
  2995. 1:45:01and a bias that controls the overall
  2996. 1:45:03trigger happiness of this neuron
  2997. 1:45:06and then we're going to implement a def
  2998. 1:45:08underscore underscore call
  2999. 1:45:11of self and x some input x
  3000. 1:45:14and really what we don't do here is w
  3001. 1:45:15times x plus b
  3002. 1:45:17where w times x here is a dot product
  3003. 1:45:19specifically
  3004. 1:45:21now if you haven't seen
  3005. 1:45:22call
  3006. 1:45:24let me just return 0.0 here for now the
  3007. 1:45:26way this works now is we can have an x
  3008. 1:45:28which is say like 2.0 3.0 then we can
  3009. 1:45:31initialize a neuron that is
  3010. 1:45:32two-dimensional
  3011. 1:45:33because these are two numbers and then
  3012. 1:45:35we can feed those two numbers into that
  3013. 1:45:37neuron to get an output
  3014. 1:45:39and so when you use this notation n of x
  3015. 1:45:42python will use call
  3016. 1:45:45so currently call just return 0.0
  3017. 1:45:50now we'd like to actually do the forward
  3018. 1:45:52pass of this neuron instead
  3019. 1:45:54so we're going to do here first is we
  3020. 1:45:57need to basically multiply all of the
  3021. 1:45:58elements of w with all of the elements
  3022. 1:46:01of x pairwise we need to multiply them
  3023. 1:46:04so the first thing we're going to do is
  3024. 1:46:05we're going to zip up
  3025. 1:46:07celta w and x
  3026. 1:46:09and in python zip takes two iterators
  3027. 1:46:12and it creates a new iterator that
  3028. 1:46:14iterates over the tuples of the
  3029. 1:46:16corresponding entries
  3030. 1:46:17so for example just to show you we can
  3031. 1:46:20print this list
  3032. 1:46:22and still return 0.0 here
  3033. 1:46:30sorry
  3034. 1:46:34so we see that these w's are paired up
  3035. 1:46:36with the x's w with x
  3036. 1:46:41and now what we want to do is
  3037. 1:46:47for w i x i in
  3038. 1:46:50we want to multiply w times
  3039. 1:46:52w wi times x i
  3040. 1:46:54and then we want to sum all of that
  3041. 1:46:56together
  3042. 1:46:57to come up with an activation
  3043. 1:46:59and add also subnet b on top
  3044. 1:47:02so that's the raw activation and then of
  3045. 1:47:04course we need to pass that through a
  3046. 1:47:05non-linearity so what we're going to be
  3047. 1:47:07returning is act.10h
  3048. 1:47:09and here's out
  3049. 1:47:12so
  3050. 1:47:13now we see that we are getting some
  3051. 1:47:14outputs and we get a different output
  3052. 1:47:16from a neuron each time because we are
  3053. 1:47:17initializing different weights and by
  3054. 1:47:19and biases
  3055. 1:47:21and then to be a bit more efficient here
  3056. 1:47:22actually sum by the way takes a second
  3057. 1:47:25optional parameter which is the start
  3058. 1:47:28and by default the start is zero so
  3059. 1:47:31these elements of this sum will be added
  3060. 1:47:34on top of zero to begin with but
  3061. 1:47:35actually we can just start with cell dot
  3062. 1:47:37b
  3063. 1:47:38and then we just have an expression like
  3064. 1:47:39this
  3065. 1:47:45and then the generator expression here
  3066. 1:47:47must be parenthesized in python
  3067. 1:47:49there we go
  3068. 1:47:53yep so now we can forward a single
  3069. 1:47:55neuron next up we're going to define a
  3070. 1:47:57layer of neurons so here we have a
  3071. 1:47:59schematic for a mlb
  3072. 1:48:02so we see that these mlps each layer
  3073. 1:48:05this is one layer has actually a number
  3074. 1:48:07of neurons and they're not connected to
  3075. 1:48:08each other but all of them are fully
  3076. 1:48:09connected to the input
  3077. 1:48:11so what is a layer of neurons it's just
  3078. 1:48:13it's just a set of neurons evaluated
  3079. 1:48:15independently
  3080. 1:48:16so
  3081. 1:48:17in the interest of time i'm going to do
  3082. 1:48:20something fairly straightforward here
  3083. 1:48:23it's um
  3084. 1:48:25literally a layer is just a list of
  3085. 1:48:27neurons
  3086. 1:48:28and then how many neurons do we have we
  3087. 1:48:30take that as an input argument here how
  3088. 1:48:32many neurons do you want in your layer
  3089. 1:48:34number of outputs in this layer
  3090. 1:48:36and so we just initialize completely
  3091. 1:48:38independent neurons with this given
  3092. 1:48:40dimensionality and when we call on it we
  3093. 1:48:43just independently
  3094. 1:48:44evaluate them so now instead of a neuron
  3095. 1:48:47we can make a layer of neurons they are
  3096. 1:48:49two-dimensional neurons and let's have
  3097. 1:48:51three of them
  3098. 1:48:52and now we see that we have three
  3099. 1:48:53independent evaluations of three
  3100. 1:48:55different neurons
  3101. 1:48:57right
  3102. 1:48:58okay finally let's complete this picture
  3103. 1:49:00and define an entire multi-layer
  3104. 1:49:02perceptron or mlp
  3105. 1:49:04and as we can see here in an mlp these
  3106. 1:49:06layers just feed into each other
  3107. 1:49:07sequentially
  3108. 1:49:09so let's come here and i'm just going to
  3109. 1:49:11copy the code here in interest of time
  3110. 1:49:14so an mlp is very similar
  3111. 1:49:16we're taking the number of inputs
  3112. 1:49:18as before but now instead of taking a
  3113. 1:49:20single n out which is number of neurons
  3114. 1:49:22in a single layer we're going to take a
  3115. 1:49:24list of an outs and this list defines
  3116. 1:49:26the sizes of all the layers that we want
  3117. 1:49:28in our mlp
  3118. 1:49:30so here we just put them all together
  3119. 1:49:31and then iterate over consecutive pairs
  3120. 1:49:34of these sizes and create layer objects
  3121. 1:49:36for them
  3122. 1:49:37and then in the call function we are
  3123. 1:49:39just calling them sequentially so that's
  3124. 1:49:41an mlp really
  3125. 1:49:42and let's actually re-implement this
  3126. 1:49:44picture so we want three input neurons
  3127. 1:49:46and then two layers of four and an
  3128. 1:49:48output unit
  3129. 1:49:49so
  3130. 1:49:50we want
  3131. 1:49:52a three-dimensional input say this is an
  3132. 1:49:54example input we want three inputs into
  3133. 1:49:57two layers of four and one output
  3134. 1:50:00and this of course is an mlp
  3135. 1:50:03and there we go that's a forward pass of
  3136. 1:50:05an mlp
  3137. 1:50:06to make this a little bit nicer you see
  3138. 1:50:08how we have just a single element but
  3139. 1:50:09it's wrapped in a list because layer
  3140. 1:50:11always returns lists
  3141. 1:50:13circle for convenience
  3142. 1:50:15return outs at zero if len out is
  3143. 1:50:18exactly a single element
  3144. 1:50:20else return fullest
  3145. 1:50:22and this will allow us to just get a
  3146. 1:50:23single value out at the last layer that
  3147. 1:50:25only has a single neuron
  3148. 1:50:28and finally we should be able to draw
  3149. 1:50:29dot of n of x
  3150. 1:50:31and
  3151. 1:50:32as you might imagine
  3152. 1:50:34these expressions are now getting
  3153. 1:50:36relatively involved
  3154. 1:50:38so this is an entire mlp that we're
  3155. 1:50:40defining now
  3156. 1:50:45all the way until a single output
  3157. 1:50:48okay
  3158. 1:50:49and so obviously you would never
  3159. 1:50:50differentiate on pen and paper these
  3160. 1:50:52expressions but with micrograd we will
  3161. 1:50:55be able to back propagate all the way
  3162. 1:50:56through this
  3163. 1:50:58and back propagate
  3164. 1:50:59into
  3165. 1:51:00these weights of all these neurons so
  3166. 1:51:02let's see how that works okay so let's
  3167. 1:51:04create ourselves a very simple
  3168. 1:51:06example data set here
  3169. 1:51:08so this data set has four examples
  3170. 1:51:11and so we have four possible
  3171. 1:51:13inputs into the neural net
  3172. 1:51:15and we have four desired targets so we'd
  3173. 1:51:17like the neural net to assign
  3174. 1:51:21or output 1.0 when it's fed this example
  3175. 1:51:24negative one when it's fed these
  3176. 1:51:25examples and one when it's fed this
  3177. 1:51:26example so it's a very simple binary
  3178. 1:51:28classifier neural net basically that we
  3179. 1:51:30would like here
  3180. 1:51:32now let's think what the neural net
  3181. 1:51:33currently thinks about these four
  3182. 1:51:34examples we can just get their
  3183. 1:51:36predictions
  3184. 1:51:37um basically we can just call n of x for
  3185. 1:51:40x in axis
  3186. 1:51:42and then we can
  3187. 1:51:43print
  3188. 1:51:45so these are the outputs of the neural
  3189. 1:51:46net on those four examples
  3190. 1:51:48so
  3191. 1:51:50the first one is 0.91 but we'd like it
  3192. 1:51:52to be one so we should push this one
  3193. 1:51:55higher this one we want to be higher
  3194. 1:51:58this one says 0.88 and we want this to
  3195. 1:52:00be negative one
  3196. 1:52:02this is 0.8 we want it to be negative
  3197. 1:52:04one
  3198. 1:52:05and this one is 0.8 we want it to be one
  3199. 1:52:08so how do we make the neural net and how
  3200. 1:52:10do we tune the weights
  3201. 1:52:12to
  3202. 1:52:12better predict the desired targets
  3203. 1:52:16and the trick used in deep learning to
  3204. 1:52:18achieve this is to
  3205. 1:52:20calculate a single number that somehow
  3206. 1:52:22measures the total performance of your
  3207. 1:52:24neural net and we call this single
  3208. 1:52:25number the loss
  3209. 1:52:28so the loss
  3210. 1:52:29first
  3211. 1:52:31is is a single number that we're going
  3212. 1:52:32to define that basically measures how
  3213. 1:52:34well the neural net is performing right
  3214. 1:52:36now we have the intuitive sense that
  3215. 1:52:37it's not performing very well because
  3216. 1:52:38we're not very much close to this
  3217. 1:52:40so the loss will be high and we'll want
  3218. 1:52:43to minimize the loss
  3219. 1:52:44so in particular in this case what we're
  3220. 1:52:46going to do is we're going to implement
  3221. 1:52:47the mean squared error loss
  3222. 1:52:49so this is doing is we're going to
  3223. 1:52:51basically iterate um
  3224. 1:52:54for y ground truth
  3225. 1:52:56and y output in zip of um
  3226. 1:52:59wise and white red so we're going to
  3227. 1:53:01pair up the
  3228. 1:53:03ground truths with the predictions
  3229. 1:53:06and this zip iterates over tuples of
  3230. 1:53:07them
  3231. 1:53:08and for each
  3232. 1:53:11y ground truth and y output we're going
  3233. 1:53:13to subtract them
  3234. 1:53:16and square them
  3235. 1:53:18so let's first see what these losses are
  3236. 1:53:20these are individual loss components
  3237. 1:53:22and so basically for each
  3238. 1:53:25one of the four
  3239. 1:53:26we are taking the prediction and the
  3240. 1:53:28ground truth we are subtracting them and
  3241. 1:53:30squaring them
  3242. 1:53:32so because
  3243. 1:53:33this one is so close to its target 0.91
  3244. 1:53:36is almost one
  3245. 1:53:38subtracting them gives a very small
  3246. 1:53:40number
  3247. 1:53:41so here we would get like a negative
  3248. 1:53:43point one and then squaring it
  3249. 1:53:45just makes sure
  3250. 1:53:47that regardless of whether we are more
  3251. 1:53:49negative or more positive we always get
  3252. 1:53:51a positive
  3253. 1:53:52number instead of squaring we should we
  3254. 1:53:55could also take for example the absolute
  3255. 1:53:56value we need to discard the sign
  3256. 1:53:59and so you see that the expression is
  3257. 1:54:00ranged so that you only get zero exactly
  3258. 1:54:03when y out is equal to y ground truth
  3259. 1:54:06when those two are equal so your
  3260. 1:54:07prediction is exactly the target you are
  3261. 1:54:09going to get zero
  3262. 1:54:10and if your prediction is not the target
  3263. 1:54:12you are going to get some other number
  3264. 1:54:15so here for example we are way off and
  3265. 1:54:17so that's why the loss is quite high
  3266. 1:54:19and the more off we are the greater the
  3267. 1:54:22loss will be
  3268. 1:54:24so we don't want high loss we want low
  3269. 1:54:26loss
  3270. 1:54:27and so the final loss here will be just
  3271. 1:54:30the sum
  3272. 1:54:32of all of these
  3273. 1:54:33numbers
  3274. 1:54:34so you see that this should be zero
  3275. 1:54:36roughly plus zero roughly
  3276. 1:54:38but plus
  3277. 1:54:39seven
  3278. 1:54:40so loss should be about seven
  3279. 1:54:43here
  3280. 1:54:44and now we want to minimize the loss we
  3281. 1:54:47want the loss to be low
  3282. 1:54:49because if loss is low
  3283. 1:54:51then every one of the predictions is
  3284. 1:54:54equal to its target
  3285. 1:54:56so the loss the lowest it can be is zero
  3286. 1:54:58and the greater it is the worse off the
  3287. 1:55:01neural net is predicting
  3288. 1:55:04so now of course if we do lost that
  3289. 1:55:05backward
  3290. 1:55:07something magical happened when i hit
  3291. 1:55:09enter
  3292. 1:55:10and the magical thing of course that
  3293. 1:55:12happened is that we can look at
  3294. 1:55:14end.layers.neuron and that layers at say
  3295. 1:55:16like the the first layer
  3296. 1:55:18that neurons at zero
  3297. 1:55:22because remember that mlp has the layers
  3298. 1:55:24which is a list
  3299. 1:55:26and each layer has a neurons which is a
  3300. 1:55:28list and that gives us an individual
  3301. 1:55:29neuron
  3302. 1:55:30and then it's got some weights
  3303. 1:55:32and so we can for example look at the
  3304. 1:55:34weights at zero
  3305. 1:55:38um
  3306. 1:55:40oops it's not called weights it's called
  3307. 1:55:42w
  3308. 1:55:44and that's a value but now this value
  3309. 1:55:46also has a groud because of the backward
  3310. 1:55:48pass
  3311. 1:55:50and so we see that because this gradient
  3312. 1:55:52here on this particular weight of this
  3313. 1:55:54particular neuron of this particular
  3314. 1:55:56layer is negative
  3315. 1:55:57we see that its influence on the loss is
  3316. 1:56:00also negative so slightly increasing
  3317. 1:56:02this particular weight of this neuron of
  3318. 1:56:04this layer would make the loss go down
  3319. 1:56:08and we actually have this information
  3320. 1:56:10for every single one of our neurons and
  3321. 1:56:12all their parameters actually it's worth
  3322. 1:56:13looking at also the draw dot loss by the
  3323. 1:56:16way
  3324. 1:56:17so previously we looked at the draw dot
  3325. 1:56:19of a single neural neuron forward pass
  3326. 1:56:21and that was already a large expression
  3327. 1:56:23but what is this expression we actually
  3328. 1:56:25forwarded
  3329. 1:56:27every one of those four examples and
  3330. 1:56:29then we have the loss on top of them
  3331. 1:56:30with the mean squared error
  3332. 1:56:32and so this is a really massive graph
  3333. 1:56:36because this graph that we've built up
  3334. 1:56:38now
  3335. 1:56:39oh my gosh this graph that we've built
  3336. 1:56:41up now
  3337. 1:56:42which is kind of excessive it's
  3338. 1:56:44excessive because it has four forward
  3339. 1:56:46passes of a neural net for every one of
  3340. 1:56:48the examples and then it has the loss on
  3341. 1:56:50top
  3342. 1:56:51and it ends with the value of the loss
  3343. 1:56:53which was 7.12
  3344. 1:56:55and this loss will now back propagate
  3345. 1:56:56through all the four forward passes all
  3346. 1:56:58the way through just every single
  3347. 1:57:00intermediate value of the neural net
  3348. 1:57:03all the way back to of course the
  3349. 1:57:05parameters of the weights which are the
  3350. 1:57:06input
  3351. 1:57:07so these weight parameters here are
  3352. 1:57:10inputs to this neural net
  3353. 1:57:12and
  3354. 1:57:13these numbers here these scalars are
  3355. 1:57:15inputs to the neural net
  3356. 1:57:16so if we went around here
  3357. 1:57:18we'll probably find
  3358. 1:57:20some of these examples this 1.0
  3359. 1:57:22potentially maybe this 1.0 or you know
  3360. 1:57:25some of the others and you'll see that
  3361. 1:57:26they all have gradients as well
  3362. 1:57:28the thing is these gradients on the
  3363. 1:57:30input data are not that useful to us
  3364. 1:57:33and that's because the input data seems
  3365. 1:57:36to be not changeable it's it's a given
  3366. 1:57:38to the problem and so it's a fixed input
  3367. 1:57:40we're not going to be changing it or
  3368. 1:57:42messing with it even though we do have
  3369. 1:57:43gradients for it
  3370. 1:57:46but some of these gradients here
  3371. 1:57:49will be for the neural network
  3372. 1:57:50parameters the ws and the bs and those
  3373. 1:57:53we of course we want to change
  3374. 1:57:55okay so now we're going to want some
  3375. 1:57:58convenience code to gather up all of the
  3376. 1:57:59parameters of the neural net so that we
  3377. 1:58:01can operate on all of them
  3378. 1:58:03simultaneously and every one of them we
  3379. 1:58:05will nudge a tiny amount
  3380. 1:58:08based on the gradient information
  3381. 1:58:10so let's collect the parameters of the
  3382. 1:58:11neural net all in one array
  3383. 1:58:14so let's create a parameters of self
  3384. 1:58:17that just
  3385. 1:58:18returns celta w which is a list
  3386. 1:58:22concatenated with
  3387. 1:58:24a list of self.b
  3388. 1:58:27so this will just return a list
  3389. 1:58:29list plus list just you know gives you a
  3390. 1:58:31list
  3391. 1:58:32so that's parameters of neuron and i'm
  3392. 1:58:35calling it this way because also pi
  3393. 1:58:36torch has a parameters on every single
  3394. 1:58:38and in module
  3395. 1:58:40and uh it does exactly what we're doing
  3396. 1:58:42here it just returns the
  3397. 1:58:44parameter tensors for us as the
  3398. 1:58:46parameter scalars
  3399. 1:58:48now layer is also a module so it will
  3400. 1:58:50have parameters
  3401. 1:58:52itself
  3402. 1:58:54and basically what we want to do here is
  3403. 1:58:56something like this like
  3404. 1:59:00params is here and then for
  3405. 1:59:03neuron in salt out neurons
  3406. 1:59:07we want to get neuron.parameters
  3407. 1:59:10and we want to params.extend
  3408. 1:59:14right so these are the parameters of
  3409. 1:59:16this neuron and then we want to put them
  3410. 1:59:17on top of params so params dot extend
  3411. 1:59:21of peace
  3412. 1:59:22and then we want to return brands
  3413. 1:59:25so this is way too much code so actually
  3414. 1:59:28there's a way to simplify this which is
  3415. 1:59:31return
  3416. 1:59:33p
  3417. 1:59:35for neuron in self
  3418. 1:59:38neurons
  3419. 1:59:39for
  3420. 1:59:41p in neuron dot parameters
  3421. 1:59:45so it's a single list comprehension in
  3422. 1:59:47python you can sort of nest them like
  3423. 1:59:49this and you can um
  3424. 1:59:51then create
  3425. 1:59:52uh the desired
  3426. 1:59:54array so this is these are identical
  3427. 1:59:57we can take this out
  3428. 2:00:00and then let's do the same here
  3429. 2:00:04def parameters
  3430. 2:00:06self
  3431. 2:00:07and return
  3432. 2:00:09a parameter for layer in self dot layers
  3433. 2:00:13for
  3434. 2:00:15p in layer dot parameters
  3435. 2:00:20and that should be good
  3436. 2:00:23now let me pop out this so
  3437. 2:00:26we don't re-initialize our network
  3438. 2:00:28because we need to re-initialize
  3439. 2:00:31our
  3440. 2:00:35okay so unfortunately we will have to
  3441. 2:00:37probably re-initialize the network
  3442. 2:00:38because we just add functionality
  3443. 2:00:41because this class of course we i want
  3444. 2:00:43to get all the and that parameters but
  3445. 2:00:45that's not going to work because this is
  3446. 2:00:47the old class
  3447. 2:00:49okay
  3448. 2:00:50so unfortunately we do have to
  3449. 2:00:52reinitialize the network which will
  3450. 2:00:53change some of the numbers
  3451. 2:00:55but let me do that so that we pick up
  3452. 2:00:57the new api we can now do in the
  3453. 2:00:58parameters
  3454. 2:01:00and these are all the weights and biases
  3455. 2:01:02inside the entire neural net
  3456. 2:01:05so in total this mlp has 41 parameters
  3457. 2:01:11and
  3458. 2:01:12now we'll be able to change them
  3459. 2:01:15if we recalculate the loss here we see
  3460. 2:01:18that unfortunately we have slightly
  3461. 2:01:19different
  3462. 2:01:22predictions and slightly different laws
  3463. 2:01:26but that's okay
  3464. 2:01:28okay so we see that this neurons
  3465. 2:01:31gradient is slightly negative we can
  3466. 2:01:33also look at its data right now
  3467. 2:01:36which is 0.85 so this is the current
  3468. 2:01:38value of this neuron and this is its
  3469. 2:01:40gradient on the loss
  3470. 2:01:43so what we want to do now is we want to
  3471. 2:01:45iterate for every p in
  3472. 2:01:47n dot parameters so for all the 41
  3473. 2:01:49parameters in this neural net
  3474. 2:01:51we actually want to change p data
  3475. 2:01:55slightly
  3476. 2:01:56according to the gradient information
  3477. 2:01:59okay so
  3478. 2:02:00dot dot to do here
  3479. 2:02:02but this will be basically a tiny update
  3480. 2:02:05in this gradient descent scheme in
  3481. 2:02:08gradient descent we are thinking of the
  3482. 2:02:10gradient as a vector pointing in the
  3483. 2:02:13direction
  3484. 2:02:14of
  3485. 2:02:15increased
  3486. 2:02:16loss
  3487. 2:02:19and so
  3488. 2:02:20in gradient descent we are modifying
  3489. 2:02:22p data
  3490. 2:02:24by a small step size in the direction of
  3491. 2:02:26the gradient so the step size as an
  3492. 2:02:28example could be like a very small
  3493. 2:02:29number like 0.01 is the step size times
  3494. 2:02:32p dot grad
  3495. 2:02:35right
  3496. 2:02:36but we have to think through some of the
  3497. 2:02:37signs here
  3498. 2:02:38so uh
  3499. 2:02:40in particular working with this specific
  3500. 2:02:43example here
  3501. 2:02:44we see that if we just left it like this
  3502. 2:02:47then this neuron's value
  3503. 2:02:49would be currently increased by a tiny
  3504. 2:02:51amount of the gradient
  3505. 2:02:53the grain is negative so this value of
  3506. 2:02:56this neuron would go slightly down it
  3507. 2:02:58would become like 0.8 you know four or
  3508. 2:03:00something like that
  3509. 2:03:02but if this neuron's value goes lower
  3510. 2:03:06that would actually
  3511. 2:03:08increase the loss
  3512. 2:03:10that's because
  3513. 2:03:12the derivative of this neuron is
  3514. 2:03:14negative so increasing
  3515. 2:03:16this makes the loss go down so
  3516. 2:03:19increasing it is what we want to do
  3517. 2:03:21instead of decreasing it so basically
  3518. 2:03:23what we're missing here is we're
  3519. 2:03:24actually missing a negative sign
  3520. 2:03:26and again this other interpretation
  3521. 2:03:29and that's because we want to minimize
  3522. 2:03:30the loss we don't want to maximize the
  3523. 2:03:31loss we want to decrease it
  3524. 2:03:33and the other interpretation as i
  3525. 2:03:34mentioned is you can think of the
  3526. 2:03:36gradient vector
  3527. 2:03:37so basically just the vector of all the
  3528. 2:03:39gradients
  3529. 2:03:40as pointing in the direction of
  3530. 2:03:42increasing
  3531. 2:03:44the loss but then we want to decrease it
  3532. 2:03:46so we actually want to go in the
  3533. 2:03:47opposite direction
  3534. 2:03:49and so you can convince yourself that
  3535. 2:03:50this sort of plug does the right thing
  3536. 2:03:51here with the negative because we want
  3537. 2:03:53to minimize the loss
  3538. 2:03:55so if we nudge all the parameters by
  3539. 2:03:57tiny amount
  3540. 2:04:00then we'll see that
  3541. 2:04:02this data will have changed a little bit
  3542. 2:04:04so now this neuron
  3543. 2:04:06is a tiny amount greater
  3544. 2:04:08value so 0.854 went to 0.857
  3545. 2:04:13and that's a good thing because slightly
  3546. 2:04:16increasing this neuron
  3547. 2:04:18uh
  3548. 2:04:18data makes the loss go down according to
  3549. 2:04:21the gradient and so the correct thing
  3550. 2:04:23has happened sign wise
  3551. 2:04:26and so now what we would expect of
  3552. 2:04:27course is that
  3553. 2:04:29because we've changed all these
  3554. 2:04:30parameters we expect that the loss
  3555. 2:04:32should have gone down a bit
  3556. 2:04:35so we want to re-evaluate the loss let
  3557. 2:04:37me basically
  3558. 2:04:39this is just a data definition that
  3559. 2:04:41hasn't changed but the forward pass here
  3560. 2:04:44of the network we can recalculate
  3561. 2:04:49and actually let me do it outside here
  3562. 2:04:51so that we can compare the two loss
  3563. 2:04:52values
  3564. 2:04:54so here if i recalculate the loss
  3565. 2:04:57we'd expect the new loss now to be
  3566. 2:04:59slightly lower than this number so
  3567. 2:05:01hopefully what we're getting now is a
  3568. 2:05:03tiny bit lower than 4.84
  3569. 2:05:064.36
  3570. 2:05:08okay and remember the way we've arranged
  3571. 2:05:10this is that low loss means that our
  3572. 2:05:12predictions are matching the targets so
  3573. 2:05:15our predictions now are probably
  3574. 2:05:16slightly closer to the
  3575. 2:05:18targets and now all we have to do is we
  3576. 2:05:22have to iterate this process
  3577. 2:05:24so again um we've done the forward pass
  3578. 2:05:26and this is the loss
  3579. 2:05:28now we can lost that backward
  3580. 2:05:30let me take these out and we can do a
  3581. 2:05:32step size
  3582. 2:05:34and now we should have a slightly lower
  3583. 2:05:35loss 4.36 goes to 3.9
  3584. 2:05:39and okay so
  3585. 2:05:41we've done the forward pass here's the
  3586. 2:05:43backward pass
  3587. 2:05:44nudge
  3588. 2:05:45and now the loss is 3.66
  3589. 2:05:503.47
  3590. 2:05:52and you get the idea we just continue
  3591. 2:05:54doing this and this is uh gradient
  3592. 2:05:56descent we're just iteratively doing
  3593. 2:05:58forward pass backward pass update
  3594. 2:06:01forward pass backward pass update and
  3595. 2:06:02the neural net is improving its
  3596. 2:06:04predictions
  3597. 2:06:05so here if we look at why pred now
  3598. 2:06:09like red
  3599. 2:06:12we see that um
  3600. 2:06:14this value should be getting closer to
  3601. 2:06:16one
  3602. 2:06:16so this value should be getting more
  3603. 2:06:17positive these should be getting more
  3604. 2:06:19negative and this one should be also
  3605. 2:06:20getting more positive so if we just
  3606. 2:06:22iterate this
  3607. 2:06:23a few more times
  3608. 2:06:26actually we may be able to afford go to
  3609. 2:06:28go a bit faster let's try a slightly
  3610. 2:06:30higher learning rate
  3611. 2:06:34oops okay there we go so now we're at
  3612. 2:06:350.31
  3613. 2:06:39if you go too fast by the way if you try
  3614. 2:06:41to make it too big of a step you may
  3615. 2:06:43actually overstep
  3616. 2:06:47it's overconfidence because again
  3617. 2:06:48remember we don't actually know exactly
  3618. 2:06:50about the loss function the loss
  3619. 2:06:51function has all kinds of structure and
  3620. 2:06:53we only know about the very local
  3621. 2:06:55dependence of all these parameters on
  3622. 2:06:57the loss but if we step too far
  3623. 2:06:59we may step into you know a part of the
  3624. 2:07:01loss that is completely different
  3625. 2:07:03and that can destabilize training and
  3626. 2:07:04make your loss actually blow up even
  3627. 2:07:08so the loss is now 0.04 so actually the
  3628. 2:07:11predictions should be really quite close
  3629. 2:07:13let's take a look
  3630. 2:07:15so you see how this is almost one
  3631. 2:07:17almost negative one almost one we can
  3632. 2:07:19continue going
  3633. 2:07:21uh so
  3634. 2:07:22yep backward
  3635. 2:07:24update
  3636. 2:07:25oops there we go so we went way too fast
  3637. 2:07:28and um
  3638. 2:07:29we actually overstepped
  3639. 2:07:31so we got two uh too eager where are we
  3640. 2:07:34now oops
  3641. 2:07:36okay
  3642. 2:07:37seven e negative nine so this is very
  3643. 2:07:39very low loss
  3644. 2:07:41and the predictions
  3645. 2:07:43are basically perfect
  3646. 2:07:45so somehow we
  3647. 2:07:47basically we were doing way too big
  3648. 2:07:48updates and we briefly exploded but then
  3649. 2:07:50somehow we ended up getting into a
  3650. 2:07:51really good spot so usually this
  3651. 2:07:54learning rate and the tuning of it is a
  3652. 2:07:56subtle art you want to set your learning
  3653. 2:07:58rate if it's too low you're going to
  3654. 2:08:00take way too long to converge but if
  3655. 2:08:02it's too high the whole thing gets
  3656. 2:08:03unstable and you might actually even
  3657. 2:08:05explode the loss
  3658. 2:08:07depending on your loss function
  3659. 2:08:08so finding the step size to be just
  3660. 2:08:10right it's it's a pretty subtle art
  3661. 2:08:12sometimes when you're using sort of
  3662. 2:08:14vanilla gradient descent
  3663. 2:08:15but we happen to get into a good spot we
  3664. 2:08:17can look at
  3665. 2:08:19n-dot parameters
  3666. 2:08:22so this is the setting of weights and
  3667. 2:08:25biases
  3668. 2:08:26that makes our network
  3669. 2:08:29predict
  3670. 2:08:30the desired targets
  3671. 2:08:31very very close
  3672. 2:08:33and
  3673. 2:08:35basically we've successfully trained
  3674. 2:08:37neural net
  3675. 2:08:38okay let's make this a tiny bit more
  3676. 2:08:40respectable and implement an actual
  3677. 2:08:41training loop and what that looks like
  3678. 2:08:43so this is the data definition that
  3679. 2:08:45stays this is the forward pass
  3680. 2:08:47um so
  3681. 2:08:49for uh k in range you know we're going
  3682. 2:08:52to
  3683. 2:08:53take a bunch of steps
  3684. 2:08:57first you do the forward pass
  3685. 2:09:00we validate the loss
  3686. 2:09:03let's re-initialize the neural net from
  3687. 2:09:05scratch
  3688. 2:09:06and here's the data
  3689. 2:09:08and we first do before pass then we do
  3690. 2:09:11the backward pass
  3691. 2:09:19and then we do an update that's gradient
  3692. 2:09:21descent
  3693. 2:09:26and then we should be able to iterate
  3694. 2:09:27this and we should be able to print the
  3695. 2:09:29current step
  3696. 2:09:30the current loss um let's just print the
  3697. 2:09:33sort of
  3698. 2:09:34number of the loss
  3699. 2:09:36and
  3700. 2:09:38that should be it
  3701. 2:09:40and then the learning rate 0.01 is a
  3702. 2:09:42little too small 0.1 we saw is like a
  3703. 2:09:44little bit dangerously too high let's go
  3704. 2:09:46somewhere in between
  3705. 2:09:47and we'll optimize this for
  3706. 2:09:50not 10 steps but let's go for say 20
  3707. 2:09:52steps
  3708. 2:09:54let me erase all of this junk
  3709. 2:09:59and uh let's run the optimization
  3710. 2:10:03and you see how we've actually converged
  3711. 2:10:05slower in a more controlled manner and
  3712. 2:10:08got to a loss that is very low
  3713. 2:10:11so
  3714. 2:10:12i expect white bread to be quite good
  3715. 2:10:15there we go
  3716. 2:10:19um
  3717. 2:10:22and
  3718. 2:10:23that's it
  3719. 2:10:24okay so this is kind of embarrassing but
  3720. 2:10:25we actually have a really terrible bug
  3721. 2:10:28in here and it's a subtle bug and it's a
  3722. 2:10:31very common bug and i can't believe i've
  3723. 2:10:33done it for the 20th time in my life
  3724. 2:10:36especially on camera and i could have
  3725. 2:10:38reshot the whole thing but i think it's
  3726. 2:10:39pretty funny and you know you get to
  3727. 2:10:41appreciate a bit what um working with
  3728. 2:10:44neural nets maybe
  3729. 2:10:45is like sometimes
  3730. 2:10:47we are guilty of
  3731. 2:10:50come bug i've actually tweeted
  3732. 2:10:52the most common neural net mistakes a
  3733. 2:10:54long time ago now
  3734. 2:10:56uh and
  3735. 2:10:57i'm not really
  3736. 2:10:59gonna explain any of these except for we
  3737. 2:11:01are guilty of number three you forgot to
  3738. 2:11:03zero grad
  3739. 2:11:04before that backward what is that
  3740. 2:11:09basically what's happening and it's a
  3741. 2:11:10subtle bug and i'm not sure if you saw
  3742. 2:11:12it
  3743. 2:11:12is that
  3744. 2:11:14all of these
  3745. 2:11:15weights here have a dot data and a dot
  3746. 2:11:17grad
  3747. 2:11:19and that grad starts at zero
  3748. 2:11:22and then we do backward and we fill in
  3749. 2:11:24the gradients
  3750. 2:11:25and then we do an update on the data but
  3751. 2:11:27we don't flush the grad
  3752. 2:11:29it stays there
  3753. 2:11:31so when we do the second
  3754. 2:11:33forward pass and we do backward again
  3755. 2:11:35remember that all the backward
  3756. 2:11:36operations do a plus equals on the grad
  3757. 2:11:39and so these gradients just
  3758. 2:11:41add up and they never get reset to zero
  3759. 2:11:44so basically we didn't zero grad so
  3760. 2:11:47here's how we zero grad before
  3761. 2:11:50backward
  3762. 2:11:51we need to iterate over all the
  3763. 2:11:52parameters
  3764. 2:11:54and we need to make sure that p dot grad
  3765. 2:11:56is set to zero
  3766. 2:11:58we need to reset it to zero just like it
  3767. 2:12:00is in the constructor
  3768. 2:12:02so remember all the way here for all
  3769. 2:12:04these value nodes grad is reset to zero
  3770. 2:12:07and then all these backward passes do a
  3771. 2:12:09plus equals from that grad
  3772. 2:12:11but we need to make sure that
  3773. 2:12:13we reset these graphs to zero so that
  3774. 2:12:15when we do backward
  3775. 2:12:17all of them start at zero and the actual
  3776. 2:12:18backward pass accumulates um
  3777. 2:12:21the loss derivatives into the grads
  3778. 2:12:25so this is zero grad in pytorch
  3779. 2:12:28and uh
  3780. 2:12:30we will slightly get we'll get a
  3781. 2:12:31slightly different optimization let's
  3782. 2:12:33reset the neural net
  3783. 2:12:34the data is the same this is now i think
  3784. 2:12:37correct
  3785. 2:12:38and we get a much more
  3786. 2:12:40you know we get a much more
  3787. 2:12:42slower descent
  3788. 2:12:44we still end up with pretty good results
  3789. 2:12:46and we can continue this a bit more
  3790. 2:12:48to get down lower
  3791. 2:12:50and lower
  3792. 2:12:51and lower
  3793. 2:12:54yeah
  3794. 2:12:56so the only reason that the previous
  3795. 2:12:57thing worked it's extremely buggy um the
  3796. 2:12:59only reason that worked is that
  3797. 2:13:03this is a very very simple problem
  3798. 2:13:05and it's very easy for this neural net
  3799. 2:13:07to fit this data
  3800. 2:13:09and so the grads ended up accumulating
  3801. 2:13:12and it effectively gave us a massive
  3802. 2:13:13step size and it made us converge
  3803. 2:13:16extremely fast
  3804. 2:13:19but basically now we have to do more
  3805. 2:13:20steps to get to very low values of loss
  3806. 2:13:24and get wipe red to be really good we
  3807. 2:13:26can try to
  3808. 2:13:27step a bit greater
  3809. 2:13:34yeah we're gonna get closer and closer
  3810. 2:13:36to one minus one and one
  3811. 2:13:38so
  3812. 2:13:39working with neural nets is sometimes
  3813. 2:13:41tricky because
  3814. 2:13:43uh
  3815. 2:13:44you may have lots of bugs in the code
  3816. 2:13:47and uh your network might actually work
  3817. 2:13:49just like ours worked
  3818. 2:13:51but chances are is that if we had a more
  3819. 2:13:53complex problem then actually this bug
  3820. 2:13:55would have made us not optimize the loss
  3821. 2:13:57very well and we were only able to get
  3822. 2:13:59away with it because
  3823. 2:14:01the problem is very simple
  3824. 2:14:03so let's now bring everything together
  3825. 2:14:04and summarize what we learned
  3826. 2:14:06what are neural nets neural nets are
  3827. 2:14:09these mathematical expressions
  3828. 2:14:11fairly simple mathematical expressions
  3829. 2:14:13in the case of multi-layer perceptron
  3830. 2:14:15that take
  3831. 2:14:16input as the data and they take input
  3832. 2:14:19the weights and the parameters of the
  3833. 2:14:20neural net mathematical expression for
  3834. 2:14:22the forward pass followed by a loss
  3835. 2:14:24function and the loss function tries to
  3836. 2:14:26measure the accuracy of the predictions
  3837. 2:14:29and usually the loss will be low when
  3838. 2:14:31your predictions are matching your
  3839. 2:14:32targets or where the network is
  3840. 2:14:34basically behaving well so we we
  3841. 2:14:37manipulate the loss function so that
  3842. 2:14:38when the loss is low the network is
  3843. 2:14:40doing what you want it to do on your
  3844. 2:14:42problem
  3845. 2:14:44and then we backward the loss
  3846. 2:14:46use backpropagation to get the gradient
  3847. 2:14:48and then we know how to tune all the
  3848. 2:14:50parameters to decrease the loss locally
  3849. 2:14:52but then we have to iterate that process
  3850. 2:14:54many times in what's called the gradient
  3851. 2:14:55descent
  3852. 2:14:56so we simply follow the gradient
  3853. 2:14:58information and that minimizes the loss
  3854. 2:15:01and the loss is arranged so that when
  3855. 2:15:02the loss is minimized the network is
  3856. 2:15:04doing what you want it to do
  3857. 2:15:06and yeah so we just have a blob of
  3858. 2:15:09neural stuff and we can make it do
  3859. 2:15:11arbitrary things and that's what gives
  3860. 2:15:13neural nets their power um
  3861. 2:15:15it's you know this is a very tiny
  3862. 2:15:16network with 41 parameters
  3863. 2:15:19but you can build significantly more
  3864. 2:15:20complicated neural nets with billions
  3865. 2:15:24at this point almost trillions of
  3866. 2:15:25parameters and it's a massive blob of
  3867. 2:15:28neural tissue simulated neural tissue
  3868. 2:15:31roughly speaking
  3869. 2:15:32and you can make it do extremely complex
  3870. 2:15:34problems and these neurons then have all
  3871. 2:15:37kinds of very fascinating emergent
  3872. 2:15:39properties
  3873. 2:15:40in
  3874. 2:15:41when you try to make them do
  3875. 2:15:43significantly hard problems as in the
  3876. 2:15:45case of gpt for example
  3877. 2:15:47we have massive amounts of text from the
  3878. 2:15:49internet and we're trying to get a
  3879. 2:15:51neural net to predict to take like a few
  3880. 2:15:53words and try to predict the next word
  3881. 2:15:55in a sequence that's the learning
  3882. 2:15:56problem
  3883. 2:15:57and it turns out that when you train
  3884. 2:15:58this on all of internet the neural net
  3885. 2:16:00actually has like really remarkable
  3886. 2:16:02emergent properties but that neural net
  3887. 2:16:04would have hundreds of billions of
  3888. 2:16:05parameters
  3889. 2:16:07but it works on fundamentally the exact
  3890. 2:16:09same principles
  3891. 2:16:10the neural net of course will be a bit
  3892. 2:16:12more complex but otherwise the
  3893. 2:16:15value in the gradient is there
  3894. 2:16:17and would be identical and the gradient
  3895. 2:16:19descent would be there and would be
  3896. 2:16:21basically identical but people usually
  3897. 2:16:23use slightly different updates this is a
  3898. 2:16:25very simple stochastic gradient descent
  3899. 2:16:27update
  3900. 2:16:28um
  3901. 2:16:29and the loss function would not be mean
  3902. 2:16:30squared error they would be using
  3903. 2:16:32something called the cross-entropy loss
  3904. 2:16:34for predicting the next token so there's
  3905. 2:16:36a few more details but fundamentally the
  3906. 2:16:37neural network setup and neural network
  3907. 2:16:39training is identical and pervasive and
  3908. 2:16:42now you understand intuitively
  3909. 2:16:44how that works under the hood in the
  3910. 2:16:46beginning of this video i told you that
  3911. 2:16:47by the end of it you would understand
  3912. 2:16:48everything in micrograd and then we'd
  3913. 2:16:50slowly build it up let me briefly prove
  3914. 2:16:52that to you
  3915. 2:16:54so i'm going to step through all the
  3916. 2:16:55code that is in micrograd as of today
  3917. 2:16:57actually potentially some of the code
  3918. 2:16:59will change by the time you watch this
  3919. 2:17:00video because i intend to continue
  3920. 2:17:01developing micrograd
  3921. 2:17:03but let's look at what we have so far at
  3922. 2:17:05least init.pi is empty when you go to
  3923. 2:17:07engine.pi that has the value
  3924. 2:17:10everything here you should mostly
  3925. 2:17:11recognize so we have the data.grad
  3926. 2:17:13attributes we have the backward function
  3927. 2:17:15uh we have the previous set of children
  3928. 2:17:17and the operation that produced this
  3929. 2:17:19value
  3930. 2:17:20we have addition multiplication and
  3931. 2:17:22raising to a scalar power
  3932. 2:17:25we have the relu non-linearity which is
  3933. 2:17:27slightly different type of nonlinearity
  3934. 2:17:28than 10h that we used in this video
  3935. 2:17:30both of them are non-linearities and
  3936. 2:17:32notably 10h is not actually present in
  3937. 2:17:34micrograd as of right now but i intend
  3938. 2:17:37to add it later
  3939. 2:17:38with the backward which is identical and
  3940. 2:17:40then all of these other operations which
  3941. 2:17:42are built up on top of operations here
  3942. 2:17:45so values should be very recognizable
  3943. 2:17:47except for the non-linearity used in
  3944. 2:17:48this video
  3945. 2:17:50um there's no massive difference between
  3946. 2:17:52relu and 10h and sigmoid and these other
  3947. 2:17:54non-linearities they're all roughly
  3948. 2:17:55equivalent and can be used in mlps so i
  3949. 2:17:58use 10h because it's a bit smoother and
  3950. 2:18:00because it's a little bit more
  3951. 2:18:01complicated than relu and therefore it's
  3952. 2:18:03stressed a little bit more the
  3953. 2:18:05local gradients and working with those
  3954. 2:18:07derivatives which i thought would be
  3955. 2:18:09useful
  3956. 2:18:10and then that pi is the neural networks
  3957. 2:18:12library as i mentioned so you should
  3958. 2:18:14recognize identical implementation of
  3959. 2:18:16neuron layer and mlp
  3960. 2:18:18notably or not so much
  3961. 2:18:20we have a class module here there is a
  3962. 2:18:22parent class of all these modules i did
  3963. 2:18:24that because there's an nn.module class
  3964. 2:18:27in pytorch and so this exactly matches
  3965. 2:18:29that api and end.module and pytorch has
  3966. 2:18:31also a zero grad which i've refactored
  3967. 2:18:33out here
  3968. 2:18:36so that's the end of micrograd really
  3969. 2:18:38then there's a test
  3970. 2:18:40which you'll see
  3971. 2:18:41basically creates
  3972. 2:18:42two chunks of code one in micrograd and
  3973. 2:18:45one in pi torch and we'll make sure that
  3974. 2:18:47the forward and the backward pass agree
  3975. 2:18:49identically
  3976. 2:18:50for a slightly less complicated
  3977. 2:18:51expression a slightly more complicated
  3978. 2:18:53expression everything
  3979. 2:18:55agrees so we agree with pytorch on all
  3980. 2:18:57of these operations
  3981. 2:18:58and finally there's a demo.ipymb here
  3982. 2:19:01and it's a bit more complicated binary
  3983. 2:19:03classification demo than the one i
  3984. 2:19:04covered in this lecture so we only had a
  3985. 2:19:07tiny data set of four examples um here
  3986. 2:19:09we have a bit more complicated example
  3987. 2:19:11with lots of blue points and lots of red
  3988. 2:19:13points and we're trying to again build a
  3989. 2:19:15binary classifier to distinguish uh two
  3990. 2:19:18dimensional points as red or blue
  3991. 2:19:20it's a bit more complicated mlp here
  3992. 2:19:22with it's a bigger mlp
  3993. 2:19:24the loss is a bit more complicated
  3994. 2:19:26because
  3995. 2:19:27it supports batches
  3996. 2:19:29so because our dataset was so tiny we
  3997. 2:19:31always did a forward pass on the entire
  3998. 2:19:32data set of four examples but when your
  3999. 2:19:35data set is like a million examples what
  4000. 2:19:37we usually do in practice is we chair we
  4001. 2:19:39basically pick out some random subset we
  4002. 2:19:41call that a batch and then we only
  4003. 2:19:43process the batch forward backward and
  4004. 2:19:45update so we don't have to forward the
  4005. 2:19:47entire training set
  4006. 2:19:49so this supports batching because
  4007. 2:19:51there's a lot more examples here
  4008. 2:19:53we do a forward pass the loss is
  4009. 2:19:55slightly more different this is a max
  4010. 2:19:57margin loss that i implement here
  4011. 2:20:00the one that we used was the mean
  4012. 2:20:01squared error loss because it's the
  4013. 2:20:03simplest one
  4014. 2:20:04there's also the binary cross entropy
  4015. 2:20:06loss all of them can be used for binary
  4016. 2:20:08classification and don't make too much
  4017. 2:20:10of a difference in the simple examples
  4018. 2:20:11that we looked at so far
  4019. 2:20:13there's something called l2
  4020. 2:20:14regularization used here this has to do
  4021. 2:20:17with generalization of the neural net
  4022. 2:20:19and controls the overfitting in machine
  4023. 2:20:21learning setting but i did not cover
  4024. 2:20:23these concepts and concepts in this
  4025. 2:20:24video potentially later
  4026. 2:20:26and the training loop you should
  4027. 2:20:27recognize so forward backward with zero
  4028. 2:20:31grad
  4029. 2:20:32and update and so on you'll notice that
  4030. 2:20:35in the update here the learning rate is
  4031. 2:20:36scaled as a function of number of
  4032. 2:20:38iterations and it
  4033. 2:20:40shrinks
  4034. 2:20:41and this is something called learning
  4035. 2:20:43rate decay so in the beginning you have
  4036. 2:20:44a high learning rate and as the network
  4037. 2:20:47sort of stabilizes near the end you
  4038. 2:20:49bring down the learning rate to get some
  4039. 2:20:50of the fine details in the end
  4040. 2:20:53and in the end we see the decision
  4041. 2:20:54surface of the neural net and we see
  4042. 2:20:56that it learns to separate out the red
  4043. 2:20:58and the blue area based on the data
  4044. 2:21:00points
  4045. 2:21:01so that's the slightly more complicated
  4046. 2:21:03example and then we'll demo that hyper
  4047. 2:21:05ymb that you're free to go over
  4048. 2:21:07but yeah as of today that is micrograd i
  4049. 2:21:10also wanted to show you a little bit of
  4050. 2:21:11real stuff so that you get to see how
  4051. 2:21:13this is actually implemented in
  4052. 2:21:14production grade library like by torch
  4053. 2:21:16uh so in particular i wanted to show i
  4054. 2:21:18wanted to find and show you the backward
  4055. 2:21:20pass for 10h in pytorch so here in
  4056. 2:21:23micrograd we see that the backward
  4057. 2:21:25password 10h is one minus t square
  4058. 2:21:28where t is the output of the tanh of x
  4059. 2:21:33times of that grad which is the chain
  4060. 2:21:34rule so we're looking for something that
  4061. 2:21:36looks like this
  4062. 2:21:38now
  4063. 2:21:39i went to pytorch um which has an open
  4064. 2:21:42source github codebase and uh i looked
  4065. 2:21:45through a lot of its code
  4066. 2:21:47and honestly i i i spent about 15
  4067. 2:21:49minutes and i couldn't find 10h
  4068. 2:21:51and that's because these libraries
  4069. 2:21:53unfortunately they grow in size and
  4070. 2:21:55entropy and if you just search for 10h
  4071. 2:21:57you get apparently 2 800 results and 400
  4072. 2:22:01and 406 files so i don't know what these
  4073. 2:22:04files are doing honestly
  4074. 2:22:07and why there are so many mentions of
  4075. 2:22:0910h but unfortunately these libraries
  4076. 2:22:11are quite complex they're meant to be
  4077. 2:22:12used not really inspected um
  4078. 2:22:15eventually i did stumble on someone
  4079. 2:22:18who tries to change the 10 h backward
  4080. 2:22:21code for some reason
  4081. 2:22:22and someone here pointed to the cpu
  4082. 2:22:24kernel and the kuda kernel for 10 inch
  4083. 2:22:26backward
  4084. 2:22:27so this so basically depends on if
  4085. 2:22:29you're using pi torch on a cpu device or
  4086. 2:22:31on a gpu which these are different
  4087. 2:22:33devices and i haven't covered this but
  4088. 2:22:35this is the 10 h backwards kernel
  4089. 2:22:37for uh cpu
  4090. 2:22:40and the reason it's so large is that
  4091. 2:22:43number one this is like if you're using
  4092. 2:22:45a complex type which we haven't even
  4093. 2:22:46talked about if you're using a specific
  4094. 2:22:48data type of b-float 16 which we haven't
  4095. 2:22:50talked about
  4096. 2:22:52and then if you're not then this is the
  4097. 2:22:54kernel and deep here we see something
  4098. 2:22:57that resembles our backward pass so they
  4099. 2:23:00have a times one minus
  4100. 2:23:02b square uh so this b
  4101. 2:23:05b here must be the output of the 10h and
  4102. 2:23:07this is the health.grad so here we found
  4103. 2:23:10it
  4104. 2:23:11uh deep inside
  4105. 2:23:14pi torch from this location for some
  4106. 2:23:15reason inside binaryops kernel when 10h
  4107. 2:23:18is not actually a binary op
  4108. 2:23:21and then this is the gpu kernel
  4109. 2:23:25we're not complex
  4110. 2:23:26we're
  4111. 2:23:27here and here we go with one line of
  4112. 2:23:29code
  4113. 2:23:30so we did find it but basically
  4114. 2:23:33unfortunately these codepieces are very
  4115. 2:23:34large and
  4116. 2:23:36micrograd is very very simple but if you
  4117. 2:23:38actually want to use real stuff uh
  4118. 2:23:40finding the code for it you'll actually
  4119. 2:23:41find that difficult
  4120. 2:23:43i also wanted to show you a little
  4121. 2:23:45example here where pytorch is showing
  4122. 2:23:47you how can you can register a new type
  4123. 2:23:49of function that you want to add to
  4124. 2:23:51pytorch as a lego building block
  4125. 2:23:53so here if you want to for example add a
  4126. 2:23:55gender polynomial 3
  4127. 2:23:59here's how you could do it you will
  4128. 2:24:00register it as a class that
  4129. 2:24:03subclasses storage.org that function
  4130. 2:24:06and then you have to tell pytorch how to
  4131. 2:24:07forward your new function
  4132. 2:24:10and how to backward through it
  4133. 2:24:12so as long as you can do the forward
  4134. 2:24:14pass of this little function piece that
  4135. 2:24:15you want to add and as long as you know
  4136. 2:24:17the the local derivative the local
  4137. 2:24:19gradients which are implemented in the
  4138. 2:24:20backward pi torch will be able to back
  4139. 2:24:22propagate through your function and then
  4140. 2:24:24you can use this as a lego block in a
  4141. 2:24:26larger lego castle of all the different
  4142. 2:24:28lego blocks that pytorch already has
  4143. 2:24:31and so that's the only thing you have to
  4144. 2:24:32tell pytorch and everything would just
  4145. 2:24:33work and you can register new types of
  4146. 2:24:35functions
  4147. 2:24:36in this way following this example
  4148. 2:24:38and that is everything that i wanted to
  4149. 2:24:40cover in this lecture
  4150. 2:24:41so i hope you enjoyed building out
  4151. 2:24:42micrograd with me i hope you find it
  4152. 2:24:44interesting insightful
  4153. 2:24:46and
  4154. 2:24:47yeah i will post a lot of the links
  4155. 2:24:50that are related to this video in the
  4156. 2:24:51video description below i will also
  4157. 2:24:53probably post a link to a discussion
  4158. 2:24:55forum
  4159. 2:24:56or discussion group where you can ask
  4160. 2:24:58questions related to this video and then
  4161. 2:25:00i can answer or someone else can answer
  4162. 2:25:02your questions and i may also do a
  4163. 2:25:04follow-up video that answers some of the
  4164. 2:25:06most common questions
  4165. 2:25:08but for now that's it i hope you enjoyed
  4166. 2:25:10it if you did then please like and
  4167. 2:25:11subscribe so that youtube knows to
  4168. 2:25:13feature this video to more people
  4169. 2:25:15and that's it for now i'll see you later
  4170. 2:25:22now here's the problem
  4171. 2:25:24we know
  4172. 2:25:25dl by
  4173. 2:25:28wait what is the problem
  4174. 2:25:31and that's everything i wanted to cover
  4175. 2:25:33in this lecture
  4176. 2:25:34so i hope
  4177. 2:25:35you enjoyed us building up microcraft
  4178. 2:25:38micro crab
  4179. 2:25:42okay now let's do the exact same thing
  4180. 2:25:43for multiply because we can't do
  4181. 2:25:44something like a times two
  4182. 2:25:47oops
  4183. 2:25:50i know what happened there

About this transcript

This page contains the full transcript of The spelled-out intro to neural networks and backpropagation: building micrograd by Andrej Karpathy, generated from the public captions YouTube serves with the video. The transcript has 24,352 words across 4,183 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.