YouTube2Text

YouTube transcript (DATnpGoGhM8) — Transcript

7,407 words · 1,015 segments · language en · Watch on YouTube

Full transcript

  1. 0:05This is a very chill um low-key first
  2. 0:08lecture. You know, as usual, like I'm
  3. 0:09not going to talk about anything deep.
  4. 0:10It's mostly just a introduction about,
  5. 0:13you know, this course and some of the
  6. 0:14materials um that we're going to cover
  7. 0:17some of the topics. We're going to cover
  8. 0:19um and a very very high level
  9. 0:20introduction of machine learning. Uh um
  10. 0:24um still from a somewhat traditional
  11. 0:26viewpoint. So, the teaching staff, this
  12. 0:28is me and Chris. I guess Chris is not
  13. 0:30here today but he will start to teach
  14. 0:32from the next lecture. So basically
  15. 0:34either me or Chris will come to the
  16. 0:36lecture to give the lecture. Um and um
  17. 0:41um you know I guess [laughter]
  18. 0:43you can look look him up know I don't
  19. 0:44necessarily have to introduce him that
  20. 0:47much you know u um um he probably have
  21. 0:49introduction for himself but he works on
  22. 0:51a lot on like all kind of like AI large
  23. 0:53language models you know efficiency
  24. 0:55database you know everything. So um um
  25. 0:58and I um uh in my PhD I worked a lot on
  26. 1:02u theory for deep learning u and over
  27. 1:04the years I start to work on more
  28. 1:05empirical stuff um so um these days I'm
  29. 1:08working on AI for sciences AI for math
  30. 1:10you know like self-improving algorithms
  31. 1:13you know um um so forth basically like
  32. 1:16training large language models so
  33. 1:18prerequisites right so I think um as
  34. 1:20usual we require um some background in
  35. 1:23probability and linear algebra uh and
  36. 1:25these are some of the um CS 109 or CS
  37. 1:291116 um um this list of like a courses
  38. 1:33might be a little out of date you know
  39. 1:34there are maybe equivalent courses you
  40. 1:36know on on campus which I'm not aware of
  41. 1:38but generally speaking you need to
  42. 1:40understand some of these buzzwords a
  43. 1:41little bit you know if you don't know
  44. 1:43them you can go to the um uh there are
  45. 1:45some kind of like there's a something
  46. 1:47called section there's something called
  47. 1:49um Friday uh TA lectures so those are
  48. 1:53for some of the backgrounds that you
  49. 1:56uh uh uh that you can learn uh there. So
  50. 1:58um but generally speaking, if you don't
  51. 2:00know anything about this, you know, you
  52. 2:02probably should take some of this take
  53. 2:04this course later. Uh but if you are
  54. 2:06just missing some aspects, then you can
  55. 2:08come to some of these Friday lectures uh
  56. 2:10to catch up with some of the uh the
  57. 2:12basics. Yeah. And and I think one
  58. 2:15highlight I would say is that it's still
  59. 2:17a pretty mathematical intense course. uh
  60. 2:20uh we are trying to teach the basics you
  61. 2:22know in some sense the foundation of
  62. 2:23machine learning not necessarily proving
  63. 2:25theorems you know um but just there are
  64. 2:27derivations you know mathematical kind
  65. 2:28of like derivations you know
  66. 2:30mathematical modeling formulations you
  67. 2:32know we focus more on those as opposed
  68. 2:34to programming so uh so so we try to
  69. 2:38understand the basics behind the
  70. 2:40algorithms and code so um um so
  71. 2:44probability and linear algebra are
  72. 2:45pretty important um for understanding
  73. 2:48some of Yes. Yeah. So, uh AI tools
  74. 2:51policy which is uh [laughter]
  75. 2:53increasingly important these days. You
  76. 2:55know, I use AI tools every day. So, so
  77. 2:58we do allow you to use tools but there
  78. 3:00are some restrictions. You know, for
  79. 3:01example, you cannot copy the results
  80. 3:02from the tools to um to your uh
  81. 3:06submission, right? So um in some sense
  82. 3:08you can just only use them as a human
  83. 3:10collaborator as if they are human
  84. 3:12collaborator or maybe superhuman
  85. 3:13collaborator um um but not like
  86. 3:16[clears throat] let them generate codes
  87. 3:17that directly uh go into your u
  88. 3:20questions. So um they may there are a
  89. 3:22little more details in on the website
  90. 3:24and and we're happy to clarify if there
  91. 3:26are any questions. Um so so you know um
  92. 3:30these rules probably sounds a little bit
  93. 3:31restrictive you know uh uh to some
  94. 3:33degree but I think this is to in some
  95. 3:36sense um what's the right word like to
  96. 3:39encourage you or or to enforce the
  97. 3:42learning like the updates of your your
  98. 3:44brain if you're just training like if
  99. 3:46you're building an agent that the agents
  100. 3:48understands the the materials then uh
  101. 3:50you probably don't have to take this
  102. 3:52course right so this is more about not
  103. 3:54about training your own agent it's about
  104. 3:55training yourself like our synapsis. So
  105. 3:58[laughter]
  106. 3:58um um at least that's my interpretation
  107. 4:01so far. Um but we are still taking a
  108. 4:03somewhat kind of like uh um relatively
  109. 4:06traditional approach to think about
  110. 4:07machine learning which I think still
  111. 4:08applies because you know I think in
  112. 4:10terms of in my probably biased view. I
  113. 4:14think the uh the the methodology of
  114. 4:17training the models are still very
  115. 4:20similar to the old days. Uh but the way
  116. 4:23we are using the models are very
  117. 4:24different.
  118. 4:25>> [snorts]
  119. 4:25>> So uh so so for example in the old days
  120. 4:28I think uh if you use the machine
  121. 4:29learning in an company you have to get
  122. 4:32complicated kind of like pipelines get
  123. 4:34the data tune the model so and so forth
  124. 4:36but now you just prompt you just ask
  125. 4:38questions directly you build some kind
  126. 4:39of like skills in cloud code so and so
  127. 4:42forth right so um I think we are
  128. 4:44probably not going to focus too much on
  129. 4:45using these models you know u um you
  130. 4:48know we are not going to teach that much
  131. 4:49about how to use cloud code work or
  132. 4:52cloud code to um to write code um but we
  133. 4:55are more focusing on how to tune uh
  134. 4:57what's the fundamental technology to
  135. 4:59tune these models uh so that's why I
  136. 5:01think most of the existing context still
  137. 5:03apply we did drop some of the uh uh the
  138. 5:06outdated stuff so and uh when I was
  139. 5:09reviewing this uh uh some of these
  140. 5:11lectures you know uh just uh last week
  141. 5:14uh um I I found surprisingly some of
  142. 5:16these definitions still kind of host you
  143. 5:18know this is 1959 it's like a like a
  144. 5:21basically like a [laughter] ages ago for
  145. 5:24machine learning. U but somehow the
  146. 5:26study the the definition still seems to
  147. 5:28be pretty much the right thing. You know
  148. 5:30machine learning is a field of study
  149. 5:31that gives computer the ability to learn
  150. 5:34without being explicitly programmed. Um
  151. 5:37I think at that time I think explicit
  152. 5:39programmed means that you program some
  153. 5:41kind of like code to play chess and
  154. 5:44that's that's what explicit program mean
  155. 5:46and unarning means like you somehow use
  156. 5:49some data and you train some model to um
  157. 5:51to to predict or to make the decisions
  158. 5:54on what moves you should make. Um and
  159. 5:57then in 1998
  160. 5:59um that's like 28 years ago. So u Tom
  161. 6:02Mitchell uh he's he's a professor at CMU
  162. 6:05uh uh he said that you know computer
  163. 6:07program is set to learn from experiences
  164. 6:09E with respect to some class of tasks T
  165. 6:13and performance measure P if its
  166. 6:15performance at tasks T uh at tasks T as
  167. 6:20measured by P improves with the
  168. 6:22experience E. I think there's a um
  169. 6:24probably a small typo there. Um anyway
  170. 6:26so um I think this definition is very
  171. 6:29interesting because it mentions a few
  172. 6:31important component here um or concept
  173. 6:34here right one is experiences this is
  174. 6:35kind of like a different way to say data
  175. 6:37right uh and these days experiences is
  176. 6:39even more broad right because there are
  177. 6:41synthetic data you generate right there
  178. 6:42the the in the in the thinking tokens
  179. 6:44you know I'm not sure whether you're all
  180. 6:46familiar with that but like like the
  181. 6:47thinking tokens generate by the models
  182. 6:49you know um there are data from the web
  183. 6:52there are data from the real world there
  184. 6:53are human label data all of these are
  185. 6:55kind of experiences um and tasks
  186. 6:58[snorts] um um in the old days the tasks
  187. 7:00are very specific. you have like image
  188. 7:01classification, you know, you have like
  189. 7:03a u um you know some kind of like price
  190. 7:06prediction problem so and so forth. Each
  191. 7:07of them is a task. These days you still
  192. 7:09have tasks, the concept of tasks, but uh
  193. 7:12but it's more general purpose. You can
  194. 7:13solve like you have one model that can
  195. 7:15solve a million tasks. Um and there's
  196. 7:17performance measure P uh which is also
  197. 7:19very important because this is the
  198. 7:20accuracy you know so and so forth
  199. 7:22because you need do need a goal in some
  200. 7:24sense uh uh to drive the learning of
  201. 7:26these models and I think also um
  202. 7:31[snorts] what after in the if clause I
  203. 7:33think it's trying to emphasize that um
  204. 7:35this model seems to be better and better
  205. 7:37if you train for longer longer with more
  206. 7:39and more data right and that seems to be
  207. 7:41uh that's a very important thing for for
  208. 7:43for learning right so if you have more
  209. 7:44data it's not improving is not learning,
  210. 7:47right? So, um um so I guess for games,
  211. 7:50you know, experience is data. You know,
  212. 7:52the um yeah, the performance measures
  213. 7:54the winning rate um and and the task is
  214. 7:57just a game. So, um and so these are the
  215. 8:01broad definition. Let's try to kind of
  216. 8:02like break down a little bit, right? So,
  217. 8:05if you have a taxonomy, uh I don't think
  218. 8:07there's a consensus on, you know, what's
  219. 8:08the right taxonomy. I think the field is
  220. 8:10moving so fast so that there's no really
  221. 8:12consensus on anything because there's
  222. 8:14there's no even time to reach a
  223. 8:16consensus in some sense, right? Because
  224. 8:18sometimes you reach at the moment you
  225. 8:20reach a consensus the thing you are
  226. 8:22trying to make have a consensus on no
  227. 8:24longer important, right? Like like so uh
  228. 8:27so because people have moved to
  229. 8:28different techniques [laughter] um but
  230. 8:30but these kind of things you know at
  231. 8:31least you know if if I if you you coach
  232. 8:34me on how how do I classify machine
  233. 8:36learning techniques you know uh nobody
  234. 8:38will find crazy. So I'm not saying this
  235. 8:40is a universal ground truth but this is
  236. 8:42a reasonable view. So I think you have
  237. 8:44supervised learning, enterprise learning
  238. 8:45and reinforcement learning. Um and uh of
  239. 8:49course they're intersecting there are
  240. 8:50very a lot of intersections and even
  241. 8:51more intersections these days. Um
  242. 8:53[clears throat]
  243. 8:54and also you know you can think of this
  244. 8:56as tasks. You can also think of this as
  245. 8:58kind of like tools or methods, right?
  246. 9:00So, especially these days, I think the
  247. 9:04tasks, you know, like I think people
  248. 9:06mostly think of this as probably methods
  249. 9:08or or paradigms uh uh as opposed to the
  250. 9:11the end goal, right? Because in some
  251. 9:13sense, nobody has the end goal of just
  252. 9:15doing super learning for one task,
  253. 9:17right? So, like the end goal is always
  254. 9:18having general purpose, you know,
  255. 9:19automatic automatic learning kind of
  256. 9:21agent so forth, right? So but
  257. 9:23supervising is still used uh uh in
  258. 9:25building these agents right so uh you
  259. 9:27probably heard of the word SFT that's
  260. 9:29called supervised fine-tuning which is
  261. 9:31one form of supervised learning so but
  262. 9:34they are more like tools but not end
  263. 9:35goes uh anyway but uh so um I'm going to
  264. 9:39introduce all of this a little bit on
  265. 9:41the high level so um so I guess you know
  266. 9:44supering maybe let's think about this
  267. 9:46you know uh you know 15 years ago if you
  268. 9:48study machine learning this is probably
  269. 9:49one of the interesting application house
  270. 9:51price prediction So g data set contains
  271. 9:53some examples. Um and you want to uh
  272. 9:57predict what's the price for the uh for
  273. 9:59the house. So basically it's kind of
  274. 10:01like each example is one uh property and
  275. 10:04you have like xaxis is the square feet
  276. 10:06and y axis is the price. So given the
  277. 10:07square feet you want to predict the
  278. 10:09price and you know if this is the only
  279. 10:11thing you are given then um okay I guess
  280. 10:13the I'm going to introduce some
  281. 10:14mathematical notation which we're going
  282. 10:16to use uh throughout the uh the lecture.
  283. 10:19You know of course we're going to remind
  284. 10:20you about the notation. So often x is
  285. 10:22used to denote inputs. Y is used to
  286. 10:24denote output. So here square feet is
  287. 10:27the input you know and price is the
  288. 10:29output and you have many um pairs of
  289. 10:31this because you have many properties
  290. 10:34and each property is a is a point um and
  291. 10:37the question is just that you know if it
  292. 10:39has x square feet what's the price um
  293. 10:41and and when when you when you test it
  294. 10:43you just say I have a x and I want to
  295. 10:45know what is the corresponding y and um
  296. 10:48um as you can imagine you know the
  297. 10:49simplest way is that you just face a
  298. 10:51line uh and then when you predict you
  299. 10:53just say you read off the the value of
  300. 10:55the one at a position of x is equal to
  301. 10:58800 u and you in this case probably
  302. 11:01fitting a square quadratic function is a
  303. 11:04better then you get a different
  304. 11:05prediction so [gasps]
  305. 11:07um and in lecture two and three I think
  306. 11:08we're going to cover some of this you
  307. 11:09know how to fit a linear quadratic
  308. 11:11functions to a data set and what are the
  309. 11:13training algorithms you know what's the
  310. 11:14optimizers you know what's the loss
  311. 11:15functions you know um and and what's the
  312. 11:20the algorithms so forth so um of course
  313. 11:23you know this is very simplistic big
  314. 11:25because you know it's not only about the
  315. 11:26square feet you know you probably also
  316. 11:28need to know know the loss size then you
  317. 11:30have a two dimensional problem now you
  318. 11:31have x1 x2 as inputs and you have y and
  319. 11:34now if you visualize you you're in 3D
  320. 11:36settings um and just to um clarify some
  321. 11:40of the the the the
  322. 11:44notations or the kind of like the sorry
  323. 11:46the the the wording so so in machine
  324. 11:49learning every every concept has
  325. 11:52probably one more than one [laughter]
  326. 11:54name like a you know like a like there's
  327. 11:56so much confusions in the in the in the
  328. 11:57past now actually there are fewer
  329. 11:59confusions because some of these words
  330. 12:01are no longer used in the old days I
  331. 12:03think uh features and inputs are both
  332. 12:05referring to uh the values you are given
  333. 12:08uh and and and the price is the the
  334. 12:11output you want to predict and also time
  335. 12:12sometimes people call it labels I think
  336. 12:14I just tend to use it as use input and
  337. 12:16output these days u to avoid kind of
  338. 12:19like the uh the confusion but sometimes
  339. 12:21people still use the word labeling which
  340. 12:23means that you give a label or give a
  341. 12:26designed output you know uh uh to um you
  342. 12:30you you you provide output for certain
  343. 12:32input right you can have human labels
  344. 12:34you can have machine label and so on and
  345. 12:35so forth okay so and uh and there's a
  346. 12:39concept between uh uh about the the
  347. 12:42output type so uh and it's called
  348. 12:45regression versus classification so if
  349. 12:47the output is a continuous variable as a
  350. 12:49pri for example price then it's called
  351. 12:51regression so if the output is a
  352. 12:53discrete variable. for example a label
  353. 12:56something like you know flamingo like
  354. 12:58what for example you can label the image
  355. 13:00with some kind of like tags right so if
  356. 13:02the output is a tag then it's a discrete
  357. 13:04concept and then uh it's called labels
  358. 13:07and the problem to assign labels to the
  359. 13:10input uh is called classification uh so
  360. 13:13for example if the if you are trying to
  361. 13:15know whether it's a house or townhouse
  362. 13:17then that's two discrete variables then
  363. 13:20uh is one discrete variable with two
  364. 13:22options then that's called
  365. 13:23classification
  366. 13:25And if you have a classification problem
  367. 13:26then you cannot draw the you cannot
  368. 13:28visualize in the same way. So you you I
  369. 13:30think this is one of the ways to
  370. 13:31visualize. You have two dimensional
  371. 13:33inputs right? So lo size and um and
  372. 13:36square feet and then uh you can label um
  373. 13:40uh with like a triangle or circle and
  374. 13:43and you can see like the way to classify
  375. 13:45whether this is a house or townhouse is
  376. 13:47probably by building some kind of
  377. 13:48classifier that separates the plane into
  378. 13:50two parts and you say everything above
  379. 13:52the line is a house and everything below
  380. 13:54the line is a townhouse. Uh something
  381. 13:56like that. And in lecture three to five,
  382. 13:58we are talking about classification. And
  383. 13:59this is actually very important uh even
  384. 14:02for these days because all of these
  385. 14:03large language models fundamentally is a
  386. 14:05classification problem. When you're
  387. 14:06predicting the next tax, next token or
  388. 14:08the next word, the word is a discrete
  389. 14:11label. So it's just you have a lot of
  390. 14:12labels. You have like 50k these days
  391. 14:14probably 250k different uh words or
  392. 14:18tokens. Um and and you you're
  393. 14:20classifying into all of those uh u
  394. 14:23discrete choices, right? like you have
  395. 14:25to choose one next word among 50k
  396. 14:29possible options. Um so um and um and
  397. 14:34then you can talk about you know even
  398. 14:35high dimensional inputs right so like
  399. 14:37the inputs you know eventually you want
  400. 14:39to know everything about this property
  401. 14:41to be able to predict the price. So you
  402. 14:44can have living size loss size floors
  403. 14:46conditions zip code so on and so forth.
  404. 14:48Um so so so eventually everything is a
  405. 14:49high dimensional problem. Um and also
  406. 14:51the outputs can be also high dimensional
  407. 14:53as well. you know here I'm only having
  408. 14:54the price but you can have even much
  409. 14:56more complicated uh outputs so um and uh
  410. 15:02and some just you know you can as you
  411. 15:04can imagine there are some kind of
  412. 15:05applications you know image
  413. 15:06classification so this is a
  414. 15:07classification problem because each
  415. 15:09image is has a has a label like you know
  416. 15:13uh which is the main object in the uh in
  417. 15:16the image flamingo quark you know u
  418. 15:19quail you know so and so forth Egyptian
  419. 15:21cat um and uh this This is actually at
  420. 15:24the beginning of the of this AI uh era,
  421. 15:26right? So like basically um I think Fay
  422. 15:30uh she's a professor at Stanford. So uh
  423. 15:32she uh and her team collected this 1
  424. 15:35million uh this data set called image
  425. 15:37night which has 1 million image and
  426. 15:38label pairs. Uh and 1 million pairs at
  427. 15:41that time is really big. The second
  428. 15:43largest data set was like 50k. It's like
  429. 15:4520 lakhs bigger than the second biggest
  430. 15:47one. Uh and uh um and and then um uh uh
  431. 15:52Jeff Hinton and and uh Alex Krosvski and
  432. 15:56and um and Leah Sask they they they
  433. 15:58published the famous Alex Net paper
  434. 16:00which was trained on this data set um
  435. 16:03and with new works and gets fantastic
  436. 16:05performance and that's the pretty much
  437. 16:07the beginning of this era of AI. Of
  438. 16:09course there are several waves in the
  439. 16:10last 10 years right there was like the
  440. 16:12vision uh uh innovation and then people
  441. 16:14moved to language and then uh at the
  442. 16:16beginning the language is models is of
  443. 16:18one form and then later on we have the
  444. 16:20large language models um and that goes
  445. 16:22to uh so basically that's how this last
  446. 16:25kind of 10 years have uh has evolved but
  447. 16:28everything started with this this is
  448. 16:30very very big data set uh so that you
  449. 16:32can do this very simple classification
  450. 16:33problem at this point right like I in
  451. 16:362026 this is a problem that's supposed
  452. 16:37to
  453. 16:38It's one task first of all, right? So
  454. 16:40like you can only do that the only thing
  455. 16:42you can do with this is like you can
  456. 16:43label it nothing else. You cannot ask
  457. 16:45any questions about these images, right?
  458. 16:46So it's only one million pairs, right?
  459. 16:48Uh and the accuracy is only like at that
  460. 16:50point the first paper I think I remember
  461. 16:52the accuracy was on 50% and now it's
  462. 16:53like 95% maybe 97%. Um um so um anyway
  463. 16:58so the X is the raw pixels of the image
  464. 17:00and Y is the main object which is a
  465. 17:02discrete object. So and and then in
  466. 17:05vision there was a a few years people
  467. 17:07study you know people actually still use
  468. 17:09these kind of tools to uh do bounding
  469. 17:11box localization you can because each
  470. 17:14image is not only one object you can uh
  471. 17:16look um locate every objects and have
  472. 17:19bounding box for them and in this case
  473. 17:21the the y is the bounding boxes the the
  474. 17:24output is the bounding boxes and how how
  475. 17:25do you describe the bounding boxes you
  476. 17:27describe it by uh uh uh four numbers one
  477. 17:30number for the left bottom corner and
  478. 17:32two sorry two numbers for the left
  479. 17:33bottom corner that's like X or Y
  480. 17:35coordinates for the bottom left corner
  481. 17:37and the two numbers for the top right
  482. 17:39corner and with these two corners you
  483. 17:40you you uniquely identify the bounding
  484. 17:42box [snorts] right so so basically
  485. 17:44outputs are four numerical numbers and
  486. 17:46the input is still the raw pixels and uh
  487. 17:49and machine translation is one of the
  488. 17:51super problems so basically you just
  489. 17:52have the input which is you know English
  490. 17:54the output it could be Chinese or any
  491. 17:56other languages or vice versa and u um
  492. 18:00so this is x to y um Um I think in and
  493. 18:03by the way in this course we we're not
  494. 18:05going to cover too much about
  495. 18:06applications. We are mostly trying to
  496. 18:08teach a little bit about the fundamental
  497. 18:10techniques u behind these applications
  498. 18:13um um and and use more toy examples to
  499. 18:15demonstrate the techniques. U if you
  500. 18:17really cover the applications I think
  501. 18:19there are other lectures other courses
  502. 18:21on campus that are more useful. So and
  503. 18:24uh um and deep learning you know as I
  504. 18:26said you know it's is one of the is is
  505. 18:28the technique that enables all of this.
  506. 18:30So basically it's a new network you know
  507. 18:32which is very complicated you know I'm
  508. 18:34showing a small examples but actually
  509. 18:36there are like millions of neurons or
  510. 18:38billions of neurons in it and uh okay
  511. 18:41maybe not billions of neurons but like
  512. 18:42there are there are there are billions
  513. 18:44of parameters each parameter is actually
  514. 18:46like some kind of like edge um you know
  515. 18:48we're not going to details today but um
  516. 18:51um so in some sense this is kind of like
  517. 18:53simulating to some degree simulating the
  518. 18:55human brains right so the neurons are
  519. 18:57the nose and the and and the edges are
  520. 18:59the synapses
  521. 19:00Um and and you give this image to it and
  522. 19:03then it outputs uh the labels and we're
  523. 19:06going to talk about the basics about
  524. 19:07this like what's the how how do you
  525. 19:09define your artworks what are some of
  526. 19:11the uh um uh the basics on how how you
  527. 19:14define and what are the important things
  528. 19:16to pay attention to and then most
  529. 19:17importantly how do you train uh the
  530. 19:19models by computing the gradient because
  531. 19:21you're going to define loss functions
  532. 19:22and then you do the so-called back
  533. 19:24propagation and and this back
  534. 19:25propagation is how you automatically
  535. 19:28compute the gradient of such a complex
  536. 19:30uh model or or function um um and and
  537. 19:33that's lecture seven and eight. Um so
  538. 19:37okay so that's super learning u let me
  539. 19:39see how much time I spending okay um any
  540. 19:44questions no feel free to interrupt me
  541. 19:45at any point you know this is a very
  542. 19:47low-key lecture I think I've only spent
  543. 19:49probably 15 minutes if there's no
  544. 19:51question um um yeah okay anyway so let's
  545. 19:56move on to an express learning um so
  546. 19:58what does that mean so so so it means
  547. 20:01that you only have uh a data set with
  548. 20:04only inputs But there's no outputs. So
  549. 20:06there's no clear tasks, right? Because
  550. 20:07you don't know what they are trying to
  551. 20:09predict. You just only have inputs.
  552. 20:11So for example, you um um for super case
  553. 20:15you have the label which is like house
  554. 20:17versus townhouse, right? So you have
  555. 20:18this triangle versus circle and in the
  556. 20:20plot. But if you are unsurprised, then
  557. 20:23you don't have that information. You
  558. 20:24just have a bunch of like a um data,
  559. 20:28right? And you don't know what the type
  560. 20:29of the [clears throat] house uh it is.
  561. 20:32So um so just a b of sketch uh points on
  562. 20:34the on the plot uh and you still want to
  563. 20:38understand something about it right so
  564. 20:39there's still some patterns on the right
  565. 20:40hand side figure um you can clearly see
  566. 20:43you know if you're human right um um so
  567. 20:46so in some sense you know we are trying
  568. 20:48to get as much as possible from the data
  569. 20:50without the labels that's pretty much
  570. 20:52what learning is doing it's it's a very
  571. 20:54vaguely defined um goal because what
  572. 20:56does interesting structure means in data
  573. 20:59right what does that mean uh In some
  574. 21:01sense actually depends on your
  575. 21:02methodology right like different methods
  576. 21:05will give you different structures right
  577. 21:07in some sense um um um so for example
  578. 21:11you know for this data point if you
  579. 21:12there's a technique called clustering
  580. 21:14which just means that you look you find
  581. 21:16out you know data points that are
  582. 21:18somewhat similar right uh and in 2D is
  583. 21:21actually pretty obvious you just I think
  584. 21:23if you ask humans to cluster you
  585. 21:25probably get similar uh results or maybe
  586. 21:27you can cluster in this way as well um I
  587. 21:30guess this is probably value clustering
  588. 21:32because there's a clearer pattern. Um
  589. 21:34and uh uh in lecture 9 10 we're going to
  590. 21:37talk about some some of these clustering
  591. 21:38algorithms right like what's the uh the
  592. 21:41way to design these algorithms what's
  593. 21:42the the the loss function or what's the
  594. 21:44objectives here so and so forth so and
  595. 21:48[laughter] in terms of applications I
  596. 21:50guess these are some of the uh the
  597. 21:52classical applications right so for
  598. 21:54example uh this is about classroom genes
  599. 21:56right so every column is the gene
  600. 21:59expression uh the genes for uh one
  601. 22:02individual and you have so individuals
  602. 22:04they have different genes and uh and you
  603. 22:07you put them into a table and you can
  604. 22:09see if you sort them in the right way
  605. 22:10this is actually somewhat sorted. So um
  606. 22:13you can see clearly patterns right so
  607. 22:14there are a bunch of individuals which
  608. 22:16have those kind of genes and there's a
  609. 22:17bunch of individuals which has these
  610. 22:19kind of genes u and and then you can
  611. 22:21classify them into different uh types
  612. 22:24and understand you know what genes can
  613. 22:25cause what kind of like disease or
  614. 22:27symptoms so and so forth. So um and uh
  615. 22:30and these are the clusters probably and
  616. 22:33you can also cluster uh uh words or
  617. 22:36documents. So um uh so I think this
  618. 22:39example is like you have um on the
  619. 22:44horizontal dimension is documents. So
  620. 22:47each document contains a number of words
  621. 22:49and you count how many words. uh for
  622. 22:52example this means that pollution show
  623. 22:54up in document D0 for a lot of times and
  624. 22:57that's why it's darker and uh and you
  625. 22:59see this so basically you count how many
  626. 23:01times each word shows up in the document
  627. 23:04that's it that's how you build this
  628. 23:06table and you don't see much pattern
  629. 23:07here right but if you um
  630. 23:11regroup reorder them and uh let me see
  631. 23:14whether you can see oh sorry let's wait
  632. 23:18for the animation okay yeah so so after
  633. 23:20If you sort in the right way, you can
  634. 23:22find groups of like documents which have
  635. 23:24similar patterns, right? For example,
  636. 23:26these three documents all have a lot of
  637. 23:29mentions of shuttle, space launch, and
  638. 23:31booster that seems to be about, you
  639. 23:32know, um maybe SpaceX or something like
  640. 23:35that. And uh and there's the top left
  641. 23:37kind of like corner is about air
  642. 23:39pollution, environmental, it's about
  643. 23:41environments. So, you can figure out
  644. 23:43similar words and also similar
  645. 23:44documents, right? Because these
  646. 23:46documents are similar and these
  647. 23:47documents are similar so and so forth.
  648. 23:50So um and this proper techniques called
  649. 23:53LSA um so and we're going to talk about
  650. 23:56how to uh do some of this in lecture
  651. 23:59nine. And then this is like a post deep
  652. 24:02learning uh your techniques you know you
  653. 24:04can also represent uh words you know by
  654. 24:08vectors. So um so basically every word
  655. 24:11you can encode it into a numerical
  656. 24:13vector which is like sometimes thousand
  657. 24:15dimensional a thousand dimensional
  658. 24:16vectors or sometimes even longer. Um and
  659. 24:19uh and it turns out that actually these
  660. 24:21vectors not only have um have a lot of
  661. 24:24semantical meanings. First of all like
  662. 24:25similar words can have similar vectors
  663. 24:28right? So Rome, Paris and Berlin they
  664. 24:30have like um vectors in high space that
  665. 24:32are very close and also sometimes the
  666. 24:34the directions can corresponds to
  667. 24:36relationship and for example this
  668. 24:38direction from left to right um
  669. 24:41corresponds to relationship between the
  670. 24:42capital and the country. So basically if
  671. 24:45you uh find Italy and you don't know
  672. 24:47what's the uh the capital of it, you
  673. 24:49just go in this direction and find
  674. 24:51what's which point you head uh and that
  675. 24:54that word you had uh uh will likely be
  676. 24:57the capital of Italy and the same thing
  677. 25:00for France and Germany. So you have one
  678. 25:01direction where you can use to search
  679. 25:02for the capital of the the country. Um
  680. 25:06and there are many relationships like
  681. 25:07this people have found. Um and we're
  682. 25:10going to talk about how to kind of like
  683. 25:11learn this kind of embeddings or
  684. 25:13encodings of the the natural language.
  685. 25:15Um and all of these are learned um by
  686. 25:18training you know in those days they're
  687. 25:202014 2015 right so they're training on
  688. 25:23uh um uh unlabelled uh data sets you
  689. 25:26know basically Wikipedia uh I think
  690. 25:28actually I got this image from the web
  691. 25:30so so the free in uh I think this
  692. 25:33becomes the free coers right because
  693. 25:35people are using this uh to train models
  694. 25:38of course these days wiki is a very
  695. 25:39small corpus it's like probably like 1%
  696. 25:42or like.1% of the your the the the total
  697. 25:45corpus um um and these days. So um and
  698. 25:50and then let's talk about large we'll
  699. 25:52talk about large language models. I
  700. 25:53guess um uh uh this is u um I think
  701. 25:58large language models also have like
  702. 25:59different kind of like nuances you know
  703. 26:02um uh we're going to talk about the so
  704. 26:04so the basically the general idea is
  705. 26:06that it's a general purpose model learn
  706. 26:08unlabelled massive data sets and it's
  707. 26:10general purpose because it's not only
  708. 26:12for for example before here you know
  709. 26:15what's the purpose here you are trying
  710. 26:16to understand which words are similar to
  711. 26:18the other words you know you understand
  712. 26:19the relationship right so here you are
  713. 26:21trying to understand the the
  714. 26:22relationship ship between documents, the
  715. 26:24relationship between words, so on and so
  716. 26:25forth, right? So that's one task in some
  717. 26:27sense or maybe two tasks. Um but but
  718. 26:29here when you have large language
  719. 26:31models, we have chat GBT, it says what's
  720. 26:34the agenda? What's on agenda today? You
  721. 26:36can you can ask anything. I guess that's
  722. 26:37the right. That's the [laughter]
  723. 26:39what they have here, right? Ask
  724. 26:40anything, right? You can ask anything
  725. 26:42about um anything, right? So like it's
  726. 26:45not only one task or two task, it's
  727. 26:47about um general purpose models. Um and
  728. 26:51uh we're going to cover something about
  729. 26:52the architecture like how what new
  730. 26:54artworks uh you're going to use for this
  731. 26:56kind of like uh pretrain large language
  732. 26:58models and what loss functions um people
  733. 27:01use to train these models. uh and then
  734. 27:03we're going to talk about some of the
  735. 27:04little more advanced things like how do
  736. 27:06you prompt these models or how do you
  737. 27:07train the models to be more responsive
  738. 27:10to prompting right so I guess when you
  739. 27:11use arish models you probably have seen
  740. 27:13that it does depend on how you ask the
  741. 27:15question right so if you ask the wrong
  742. 27:17question then it doesn't respond in the
  743. 27:19right way and and different models these
  744. 27:21days probably all the models are pretty
  745. 27:22responsive to instructions but if you um
  746. 27:25two years ago you know not every model
  747. 27:27actually responds to your instructions
  748. 27:28you know and also depending on how you
  749. 27:30instruct them so and the reason of The
  750. 27:32change is just that um this frontier
  751. 27:35labs open anthropic they train the
  752. 27:37models to be more responsive to your
  753. 27:38instructions so they follow your uh your
  754. 27:40order. Um um and we're going to cover a
  755. 27:44little bit about image generation. So
  756. 27:46this is about you know how to generate
  757. 27:48realistic image you know from noise or
  758. 27:50from instructions. Um so this is the
  759. 27:52diffusion model. We're going to have
  760. 27:53like one lecture on this. Um so this is
  761. 27:55this is also unsurprised because you
  762. 27:57know you don't have labels for these
  763. 27:58images. you just have a lot of realistic
  764. 28:00image and then you can come up with a
  765. 28:02model such that you can generate more
  766. 28:04realistic images um um so I think one of
  767. 28:06the some of these data sets actually
  768. 28:08celebrity data sets I don't know why
  769. 28:09they use celebrity but apparently you
  770. 28:12can get a lot of celebrity data sets and
  771. 28:14and you can generate more images which
  772. 28:15looks like celebrities to some degree
  773. 28:17right with the right makeup you know the
  774. 28:19right style so and so forth um very
  775. 28:21realistic [snorts] um anyway so we have
  776. 28:24one lecture on this um so and then we're
  777. 28:28going to move on to uh reinforcement
  778. 28:30learning. So uh this bullet is what we I
  779. 28:33had uh uh three years ago and I I just
  780. 28:36add another bullet. Uh actually I didn't
  781. 28:39make the the form consistent. I probably
  782. 28:41should use training. Um [gasps] so uh
  783. 28:43anyway so I guess uh I think the typical
  784. 28:45way to in think about reinforcement
  785. 28:47learning is that you are trying to uh
  786. 28:49use this is a task right the task is
  787. 28:51that you want to make sequential
  788. 28:52decisions right you want to make
  789. 28:54decisions and sequential decisions and
  790. 28:55the typical applications are like alpha
  791. 28:57go you know like learning robots you
  792. 29:00know maybe actually it's kind of
  793. 29:01interesting to uh play this right so
  794. 29:03this is at iteration 10 the robots
  795. 29:06cannot really walk that much and then
  796. 29:08eventually uh you can Um, oh, it I think
  797. 29:13iteration 20 it doesn't do anything yet.
  798. 29:16And then at 80, I think it's working a
  799. 29:19little bit.
  800. 29:23[clears throat]
  801. 29:24And uh, let me see. And I I plan I
  802. 29:28remember for this example now it's it's
  803. 29:31doing pretty well. And okay, I know the
  804. 29:32point here is that this is the
  805. 29:34sequential decision task because at any
  806. 29:36moment you have to decide what actions
  807. 29:38you take, right? So how do you move your
  808. 29:40arm? how do you move uh the legs so and
  809. 29:42so forth. So every time you have to take
  810. 29:44a action and it's sequential because you
  811. 29:46know what you what actions you take
  812. 29:48right now does affect your future right
  813. 29:50if you take the wrong move then maybe
  814. 29:51you cannot recover you're going to fall
  815. 29:53down and then you you can never move
  816. 29:54again. So, so, so there's a sequential
  817. 29:57dependency and there is action aspect.
  818. 30:00So, so that's the uh the relatively kind
  819. 30:02of like traditional way to think about
  820. 30:03reinforce learning but um but I think
  821. 30:05these days people also use this as a
  822. 30:07tool to trim models especially when the
  823. 30:10model have stocastic uh decisions or
  824. 30:12stoastic kind of outputs in the middle.
  825. 30:15So um probably you know a little bit you
  826. 30:17have probably played with CHBT right or
  827. 30:20or cloud. So uh if you ask a question
  828. 30:22right it will generate a lot of like
  829. 30:23tokens right a lot of tax right and some
  830. 30:26of the tax actually not revealed to you
  831. 30:27because they're saying I'm thinking
  832. 30:29right but actually the during the
  833. 30:30thinking process is kind of like
  834. 30:32generating a lot of like um tokens and
  835. 30:34all of these generation are stoastic
  836. 30:37sampling right so each time you generate
  837. 30:39a probability distribution and you
  838. 30:41sample from the probability distribution
  839. 30:42one of the words right and then you
  840. 30:44condition that word or you you give the
  841. 30:46word back to the model and then you
  842. 30:47generate next word and you generate next
  843. 30:49word so and so forth and each time is
  844. 30:50stocastic sampling process. It's not a
  845. 30:52differentiable operation. So, so um so
  846. 30:56that's why to deal with this kind of
  847. 30:57like um generation you cannot really
  848. 31:00back through it easily. In some cases
  849. 31:03you can but in some other cases you
  850. 31:04cannot. So uh and you have to use
  851. 31:07reinforce learning which is a way is a
  852. 31:09technique to deal with this kind of
  853. 31:11stoastic sampling steps in the models.
  854. 31:13Um and uh I think that's a uh and you
  855. 31:16you can interpret each of the generation
  856. 31:18step as a decision. So and uh uh and
  857. 31:20it's a sequential decision because what
  858. 31:22you generate in the the first word
  859. 31:24depend affects what you generate in the
  860. 31:25second word so on and so forth.
  861. 31:26>> So would you say that the first bullet
  862. 31:28point is how um original RL was and like
  863. 31:31the second is how we're using oil for
  864. 31:33LC.
  865. 31:35>> Yeah. More or less. Yes. Yeah. Yeah.
  866. 31:37Yeah. Exactly. Exactly. Um you know they
  867. 31:39are not fundamentally different. It's
  868. 31:41not like a [laughter] they are
  869. 31:41contradicted to each other. It's more
  870. 31:43like a slightly different viewpoint in
  871. 31:44some sense. Yeah. Um so um
  872. 31:49yeah okay um I I think another aspect of
  873. 31:53reinforcement learning is I guess my
  874. 31:55automation is broken. So um another
  875. 31:58aspect of the reinforce learning is that
  876. 32:00it also allows you to collect data
  877. 32:02interactively. This is also very
  878. 32:04different from super learning and
  879. 32:05unpress learning. I'm not sure whether
  880. 32:06you focused on you know I didn't really
  881. 32:08emphasize those right in all of the
  882. 32:09super and super sections at least
  883. 32:12traditionally you're giving a large data
  884. 32:14set um and and then you this is a you
  885. 32:16work with uh and if you're giving a
  886. 32:19bigger one you're going to work with the
  887. 32:20bigger one but but but you but you don't
  888. 32:22but your model don't generate more data
  889. 32:25so uh the data set is given right so you
  890. 32:27can you can probably try to collect more
  891. 32:28and more so and so forth right that's on
  892. 32:30a different layer but once you have the
  893. 32:32data your model only sees those data
  894. 32:34right so um but For reinforce learning,
  895. 32:37uh it's kind of like a trial and error
  896. 32:38kind of like a type of algorithm where
  897. 32:40you iterate between data collection and
  898. 32:42training, right? So you try some
  899. 32:43strategy and collect feedbacks. Uh for
  900. 32:45example, you can try how to walk and and
  901. 32:47you fall down and even it falls down,
  902. 32:49right? So like uh there's still some
  903. 32:51data you collected, right? You collect
  904. 32:53some trajectory of data, a sequence of
  905. 32:55actions and observations uh for the
  906. 32:57trials and those are new data for your
  907. 32:59training for the next round. So, so you
  908. 33:01can improve your uh uh performance based
  909. 33:04on the data you collected yourself uh as
  910. 33:06opposed to just a given data. Uh for
  911. 33:09another example is that you know um so
  912. 33:12you can say automatic theorem proving
  913. 33:14this is something that I kind of like
  914. 33:15work on uh quite a bit these days. So,
  915. 33:18so you can you given some statements,
  916. 33:20you don't have the proofs, right? You
  917. 33:22don't have the proofs, but you gener
  918. 33:23first you generate some proofs yourself
  919. 33:25and you try to verify and you select the
  920. 33:27correct ones and those are the new
  921. 33:28training data and now you have like some
  922. 33:30more training data because you have you
  923. 33:31generate some of them and then you can
  924. 33:33tune or reinforce the large language
  925. 33:35model on the correct proofs with the
  926. 33:37hope that after reinforce then you can
  927. 33:39generate more correct proofs in the next
  928. 33:40round and you get more data and the next
  929. 33:42time you're going to tune on more data
  930. 33:43and then you keep going and you have a
  931. 33:45bootstrapping kind of effect.
  932. 33:47>> [snorts]
  933. 33:48>> The same thing happens for robotics as
  934. 33:49well where you can collect more kind of
  935. 33:51trajectories and then you train on the a
  936. 33:53good trajectories. So um uh and you can
  937. 33:57also use this for training large
  938. 33:58language models. Um and uh um and the
  939. 34:01kind of the the rough idea is that you
  940. 34:03have some inputs which is for example
  941. 34:05user prompt and the logical machine
  942. 34:07model will generate a lot of tax and
  943. 34:08then you can have another u module
  944. 34:12either it's human reward or or human
  945. 34:14reward model or like kind of like
  946. 34:17another large language model that's
  947. 34:19separate discussion but there's another
  948. 34:20module which given a tax you generate so
  949. 34:24something called a reward. So basically
  950. 34:25it's indicating whether this is a good
  951. 34:27tax or not. Um and and you can use this
  952. 34:29reward as a guidance for updating the
  953. 34:31large and push model to generate better
  954. 34:32text in next round. Uh and and uh so so
  955. 34:36so and the reason why we use
  956. 34:38reinforcement here is that um uh it's
  957. 34:41needed because this whole process is not
  958. 34:44antiffiable
  959. 34:46in the most traditional sense because
  960. 34:48the generation requires stoastic
  961. 34:50sampling. Um and that's why we're going
  962. 34:52to talk about how to deal with this
  963. 34:54case. I think the technique is called
  964. 34:55policy gradient and you can use this and
  965. 34:57extend it uh and some of the buzzwords
  966. 34:59is like rovr
  967. 35:01with verify reward with human uh in the
  968. 35:04loop ro human human wait human feedback
  969. 35:08yes u and you can use this to train the
  970. 35:10longchain thought you know these are
  971. 35:11just passwords you don't have to
  972. 35:12understand um anyway so that's probably
  973. 35:15uh uh most of the lectures uh we're
  974. 35:18going to uh cover uh and I think on the
  975. 35:20website we have a link to uh the
  976. 35:22syllabus uh Uh this is my internal
  977. 35:25version. I just take a screenshot. I'm
  978. 35:27hoping that cloud code can make it more
  979. 35:28pretty, but it didn't really do the
  980. 35:30work. So I just have the the ugly
  981. 35:33version. Uh uh so um yeah. Anyway, so
  982. 35:37but this is just a summarizing uh what I
  983. 35:39have talked about. Um yeah, I think
  984. 35:41we're going to have probably what what
  985. 35:43what I didn't do is that there's another
  986. 35:45lecture we planned for ML system. So
  987. 35:47this is about how do you u u make the
  988. 35:50software and hardware um more
  989. 35:52compatible. You know actually this is
  990. 35:54very important because you know if you
  991. 35:55can make your algorithm run 2x faster
  992. 35:58you're going to save like billions of
  993. 35:59dollars for open eye right you don't
  994. 36:01even have to do it 2x faster you can
  995. 36:02just only do 20x 20% faster it's going
  996. 36:05to be because everything is so expensive
  997. 36:06right so and [snorts] and even it's not
  998. 36:08like a saving millions of dollars it's
  999. 36:11actually think about if I have algorithm
  1000. 36:13is 2x faster than you then it's one year
  1001. 36:14for me and two years for you and one
  1002. 36:16year versus two year that's just day and
  1003. 36:17night like for for these days right like
  1004. 36:20uh um so Um so that's why it's very
  1005. 36:23important and uh we'll talk a little bit
  1006. 36:26about we're trying to find a guest
  1007. 36:27lecturer who can lecturer who can uh uh
  1008. 36:30cover fairness and uh uh the the the
  1009. 36:34social impact aspect of AI which I think
  1010. 36:36will be very important as well u just
  1011. 36:38because AI is replacing so much jobs uh
  1012. 36:40and we have to think about the uh the
  1013. 36:43social impact of them. Um um yeah, I
  1014. 36:47guess that's the uh that's the overview
  1015. 36:49for the for the course.

About this transcript

This page contains the full transcript of YouTube transcript (DATnpGoGhM8) , generated from the public captions YouTube serves with the video. The transcript has 7,407 words across 1,015 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.