YouTube2Text

Why predict-then-optimize and end-to-end learning won't fix your optimization under uncertainty. — Transcript

by InsideOpt Tutorials · 7,184 words · 1,183 segments · language en · Watch on YouTube

Full transcript

  1. 0:01Today we want to talk about two
  2. 0:04approaches that people typically take
  3. 0:07when dealing with uncertain forecasts in
  4. 0:09operational research.
  5. 0:11The first one is predict and optimize
  6. 0:13and the other one is end-to-end
  7. 0:15learning, sometimes also known as
  8. 0:18predict and optimize.
  9. 0:21And today we're going to look at a
  10. 0:22couple of examples, very simple
  11. 0:24examples, that will show
  12. 0:27that both of them
  13. 0:28are not good approaches and that in fact
  14. 0:31the whole idea of end-to-end learning is
  15. 0:33a misguided research agenda.
  16. 0:37We need to talk. Welcome to Inside Opt.
  17. 0:50Let's start by looking at the typical
  18. 0:52setup as we find it in many many
  19. 0:55businesses that use optimization.
  20. 0:58The data that goes into our optimization
  21. 1:01approach, which we
  22. 1:03illustrate here by the beautiful rocket,
  23. 1:07needs to come from somewhere.
  24. 1:09And it usually comes from estimates.
  25. 1:12Think about prices, think about demand,
  26. 1:16think about traffic times, think about
  27. 1:18lead times when you order something,
  28. 1:20when it will arrive.
  29. 1:21All of these are typically estimates.
  30. 1:25And in more sophisticated departments,
  31. 1:28you will find that those estimates are
  32. 1:30not taken from thin air or simple
  33. 1:32statistics over historical data, but
  34. 1:34they actually come from a machine
  35. 1:36learning model. So, you have the
  36. 1:38historical data,
  37. 1:40then you build a machine learning model
  38. 1:41that that comes up with
  39. 1:44a
  40. 1:45scenario, a prediction of what might
  41. 1:48happen,
  42. 1:49and then you take this predicted data
  43. 1:51and hand it over to the OR department
  44. 1:53and then you get your plan.
  45. 1:55Very very common machine learning to
  46. 1:58optimization flow as you will find it in
  47. 2:01many many organizations.
  48. 2:06Now, there is of course that disconnect
  49. 2:09that you have here. On the left-hand
  50. 2:12side, you have a machine learning
  51. 2:14department. On the right-hand side, you
  52. 2:15have the OR department. And it's not
  53. 2:18just that we're handing over the
  54. 2:20predicted data that marks the
  55. 2:22delineation between the two departments.
  56. 2:25There's also something profoundly
  57. 2:26different going on on both sides of this
  58. 2:28wall.
  59. 2:30The machine learners,
  60. 2:31their task is to handle outliers, to
  61. 2:34handle errors, to handle bias, missing
  62. 2:37values. In one word,
  63. 2:39most of what they're concerned with is
  64. 2:41uncertainty.
  65. 2:43Then they make a prediction and somehow
  66. 2:46on the other side of the wall,
  67. 2:48everything is beautiful. We have binary
  68. 2:51constraints, meaning on the OR side we
  69. 2:53can say exactly whether a constraint is
  70. 2:56violated or whether it is fine.
  71. 2:59We have measurable objectives. We can
  72. 3:01look at two different plans and say
  73. 3:03exactly and minutely whether one is
  74. 3:06better than the other by evaluating the
  75. 3:08objective function.
  76. 3:10We have fixed coefficients
  77. 3:12in the matrix on the right-hand side for
  78. 3:14the objective function and in the end
  79. 3:17maybe even provable
  80. 3:20optimality.
  81. 3:22If you look at this picture, you get a
  82. 3:24feeling that there must be something
  83. 3:26magical happening
  84. 3:28when you hand over that data via that
  85. 3:31wall cuz somehow all the uncertainty has
  86. 3:34vanished.
  87. 3:37And that is exactly what those two
  88. 3:39approaches, predict and optimize and
  89. 3:41predict and optimize, are trying to
  90. 3:43reconcile.
  91. 3:45So, what's what's the issue? If we look
  92. 3:47at our optimization approach here, the
  93. 3:50legacy solvers, they want one scenario,
  94. 3:53they want one input
  95. 3:56of data
  96. 3:57and then they're going to produce a plan
  97. 4:00and they're going to tell you what the
  98. 4:01objective function value for that plan
  99. 4:03is. This case here, 1150.098.
  100. 4:07Beautiful.
  101. 4:09Cuz if you ask the machine learners on
  102. 4:13how reasonable it is to do that, they
  103. 4:15will say, "Well,
  104. 4:17we gave you one example. We gave you the
  105. 4:20scenario that you wanted,
  106. 4:23right? Like an expected case, for
  107. 4:25example. But we know that there are many
  108. 4:28different potential futures that could
  109. 4:31happen.
  110. 4:32So, what we would actually prefer to do
  111. 4:35is to give you a distribution of
  112. 4:37scenarios. So, each you know, number of
  113. 4:40scenarios and each one associated with a
  114. 4:42certain probability so that you, when
  115. 4:44you make the decision, can come up with
  116. 4:46something that's reasonable
  117. 4:48for that entire cloud of potential
  118. 4:51futures rather than just one of them.
  119. 4:54Now,
  120. 4:55no matter what you do in order to handle
  121. 4:57this because legacy solvers can't just
  122. 4:59take in, you know, 10,000 scenarios.
  123. 5:02It's not something you can do. You can't
  124. 5:03give them a posterior distribution or
  125. 5:05samples thereof and then say, "Please
  126. 5:07optimize for this."
  127. 5:09But no matter what you do in order to
  128. 5:12arrive there, and we will discuss
  129. 5:13certain ways on how you could handle
  130. 5:15this,
  131. 5:16or how you might try to handle this, I
  132. 5:19should say.
  133. 5:20Um no matter what you do, the solution
  134. 5:22you're going to get,
  135. 5:24yes, the optimizer is going to say,
  136. 5:27"Well, this is um for the expected
  137. 5:29scenario or whatever we used as input
  138. 5:31for our optimizer now is 1150.098,
  139. 5:35but in reality, even for that one plan
  140. 5:38that you're producing,
  141. 5:40the outcomes will vary.
  142. 5:43Because
  143. 5:44the future that's going to hit you later
  144. 5:46is going to vary. And that means you're
  145. 5:49going to get with a certain probability,
  146. 5:51you're going to get something much less
  147. 5:53than 1150 and then with some
  148. 5:55probability, you're going to get
  149. 5:56something that's maybe closer to 3,000.
  150. 5:59And there is a certain
  151. 6:01likelihood associated with these
  152. 6:03particular outcomes.
  153. 6:05We summarize them all by saying, "Well,
  154. 6:08we're expecting this to be 1150, but in
  155. 6:11reality,
  156. 6:12the results for that one plan that you
  157. 6:15generated are going to vary."
  158. 6:21So, what can we do?
  159. 6:24Given the fact that the machine learners
  160. 6:26have a whole cloud of potential
  161. 6:29scenarios, but our legacy optimizer only
  162. 6:32wants one input.
  163. 6:34Well, first idea is, how about we take
  164. 6:38some scenarios and for each one of those
  165. 6:40scenarios, we use our optimizer to
  166. 6:43create an optimal plan for that
  167. 6:45scenario.
  168. 6:47And then we're going to worry about what
  169. 6:48plan we're going to to use later.
  170. 6:53So, this is what it would look like,
  171. 6:55right? So, for every scenario that you
  172. 6:57can imagine,
  173. 6:58um you know, typically these are macro
  174. 7:01scenarios, so take anywhere between 5
  175. 7:03and 20. You use your optimizer, you
  176. 7:05generate a plan for that respective
  177. 7:07scenario, and now you have the problem
  178. 7:10that you somehow need to go into this
  179. 7:12set of optimal solutions that you found,
  180. 7:16optimal for the respective scenario, of
  181. 7:18course,
  182. 7:19and
  183. 7:21generate one that will work
  184. 7:23overall.
  185. 7:26So, this can be a daunting task. And if
  186. 7:28you've ever said in executive meetings
  187. 7:31where scenario planning happened, so you
  188. 7:34have optimal plans for each one of them,
  189. 7:36you know what political discussions can
  190. 7:39can unfold um where people start
  191. 7:43reasoning about the likelihood of
  192. 7:44individual scenarios happening based on
  193. 7:46the features that they like or dislike
  194. 7:49about the plans that would be optimal
  195. 7:50for that respective scenario.
  196. 7:53So, it's like an after-the-fact
  197. 7:56discussion of, "Well, I like this
  198. 7:58solution. Now I'm going to argue why
  199. 8:00it's the right one." Um that can be very
  200. 8:03very frustrating um if you're sitting in
  201. 8:06in one of those political meetings.
  202. 8:10So, let's look at a particular example
  203. 8:12on why the whole idea is rather strange.
  204. 8:16Um this is of course a contrived
  205. 8:18example, but it perfectly illustrates
  206. 8:20the point.
  207. 8:21So, let's say here,
  208. 8:23um we have a stochastic linear program.
  209. 8:26And it's a very simple linear program
  210. 8:27with just two variables. We have two
  211. 8:30constraints, 10x + y is lower equal to
  212. 8:321,000, x + 1,000y is lower equal to
  213. 8:351,000, x and y are greater or equal to
  214. 8:380. The feasible region here is basically
  215. 8:40everything between 0 and 1, except that
  216. 8:43this point here isn't 1 1, but 0.999 and
  217. 8:470.999 and then some.
  218. 8:50Right?
  219. 8:52Now, what makes this LP stochastic?
  220. 8:55Well, we don't know
  221. 8:57what we are supposed to optimize for.
  222. 9:00Um with probability 1/2, we might have
  223. 9:03to optimize for y and with probability
  224. 9:051/2, we might have to optimize for x.
  225. 9:09So, we have two scenarios, not more, two
  226. 9:11scenarios, two variables, very simple,
  227. 9:14right?
  228. 9:15Uh everything is symmetric uh in in this
  229. 9:18example.
  230. 9:19It's really not very complicated.
  231. 9:22But we don't know whether we're going to
  232. 9:23optimize for x or for y. Now, imagine
  233. 9:25that you had created optimal solutions
  234. 9:28for the two scenarios that might hit
  235. 9:30you.
  236. 9:31Well, if you're optimizing for x, well,
  237. 9:33a good idea is to set x to 1, right?
  238. 9:36Uh and y to 0. And that's doable. You
  239. 9:39can do that.
  240. 9:41And similarly, if you're supposed to
  241. 9:42maximize for y, you're going to set x to
  242. 9:450 and y to 1.
  243. 9:47Beautiful.
  244. 9:49But take any one of those two solutions,
  245. 9:52which were optimal for the two scenarios
  246. 9:54that can actually happen, any one of
  247. 9:56them. And suddenly the expected case
  248. 10:00is or the expected result is .5.
  249. 10:03And the distribution of outcomes looks
  250. 10:05very daunting cuz with 50% you're going
  251. 10:08to get one and with 50% you're going to
  252. 10:09get nada. Zits, zilch, nothing.
  253. 10:13Okay? So, take either one of those two
  254. 10:16solutions
  255. 10:18and you're going to end up with a very
  256. 10:20poor expected value
  257. 10:22and at the same time also um a a very
  258. 10:26brittle result, right? So, you might be
  259. 10:28lucky and you were in the right
  260. 10:29scenario, there's 50% chance of that,
  261. 10:32but with 50% chance you you're suddenly
  262. 10:34ending up with with nothing, right?
  263. 10:37Where you were expecting one.
  264. 10:40And that even though a compromise
  265. 10:42candidate is perfectly available to you.
  266. 10:46You could have taken .999 and .999 as a
  267. 10:49solution
  268. 10:51and it's feasible with respect to both
  269. 10:53constraints.
  270. 10:54And no matter which scenario later
  271. 10:56happens, you're going to get .999.
  272. 11:00The only reason why you're not seeing
  273. 11:01that is because it's not optimal for
  274. 11:04either of the two scenarios. You could
  275. 11:07have gone to one instead of .999. And
  276. 11:11since you were so greedy
  277. 11:13and created the provably optimal
  278. 11:15solution for each scenario,
  279. 11:17you lack any form of robustness with
  280. 11:21respect to the other scenarios that are
  281. 11:23there.
  282. 11:24Right? So, the the very reason that you
  283. 11:26wouldn't forego that 1/1000 that is in
  284. 11:29there
  285. 11:31made it made it so that you didn't see
  286. 11:34that compromise candidate. The optimal
  287. 11:36solutions that you're seeing here for
  288. 11:38each scenario do not contain that
  289. 11:40solution and no matter which one of
  290. 11:42those two you're choosing,
  291. 11:43you're not really doing well.
  292. 11:45Right? This solution here is much much
  293. 11:48better, especially if you're in a
  294. 11:49business environment.
  295. 11:52Now, this is of course a contrived
  296. 11:54mathematical example to illustrate the
  297. 11:56point that the very fact that you're
  298. 11:59optimizing, maybe even to the to the
  299. 12:02last bit provable optimality for each
  300. 12:04scenario, makes the set of solutions
  301. 12:07that you're going to get as a whole very
  302. 12:09very brittle.
  303. 12:11So, let's have a look at a real example.
  304. 12:14So, this this is a network design
  305. 12:17problem. Our client here has four
  306. 12:19distribution centers already in New
  307. 12:21Jersey, in Ohio, in Georgia, and in
  308. 12:24California and is thinking of opening
  309. 12:27two new ones.
  310. 12:29Now, we're running a thousand scenarios.
  311. 12:31For the first 500 scenarios, it's a good
  312. 12:33idea to open um the two new facilities,
  313. 12:37one in Oregon, the other one in
  314. 12:38Pennsylvania.
  315. 12:41Then we go to the next 500 scenarios and
  316. 12:44we find, ah, it would be a better idea
  317. 12:46for you went to Pennsylvania and Texas.
  318. 12:49So, what we have is
  319. 12:51Pennsylvania and Texas or Pennsylvania
  320. 12:53and Oregon.
  321. 12:55And now, you know, if we if we take
  322. 12:57those two solutions, I mean, there were
  323. 12:59a thousand scenarios, but we only get
  324. 13:00two solutions,
  325. 13:02we somehow have to find, well, which one
  326. 13:04of them is better.
  327. 13:06Well, as it turns out for this data that
  328. 13:10we ran here, the optimal solution would
  329. 13:12have been to open in Oregon and in
  330. 13:16Texas.
  331. 13:18So, even though you found that both
  332. 13:20optimal solutions wanted to open a
  333. 13:22distribution center in Pennsylvania,
  334. 13:25for the for the totality of the thousand
  335. 13:28scenarios, it's actually a good idea not
  336. 13:29to do that at all
  337. 13:31and instead to go to the to the other
  338. 13:34two options that were on the table,
  339. 13:36which were Oregon and Texas.
  340. 13:39Okay? So, you get 500 times this
  341. 13:42solution A, 500 times solution B, and
  342. 13:46somehow the one thing that seemed to be
  343. 13:48clear that you're going to open in
  344. 13:50Pennsylvania is actually not a good
  345. 13:53thing to do. You should just open here
  346. 13:55or here. Now, the kicker is, well, if
  347. 13:57the data had been slightly different,
  348. 14:00you would have still gotten this
  349. 14:02solution for the first 500
  350. 14:04for the first 500 scenarios and this
  351. 14:06solution for the next 500 scenarios, but
  352. 14:09now the optimal solution would have been
  353. 14:11to open in Pennsylvania and in Colorado.
  354. 14:16So,
  355. 14:19even though you find common structure
  356. 14:21among all feasible solutions, all all
  357. 14:23optimal solutions that you have run,
  358. 14:26doesn't mean that you should do that,
  359. 14:27right? So, as the Pennsylvania example
  360. 14:29show you, nor does it mean that
  361. 14:32something that was never opened before
  362. 14:35wouldn't be a good idea to open
  363. 14:38in the compromise as you're building a
  364. 14:40compromise scenario for all of the
  365. 14:42thousand scenarios.
  366. 14:44And that's what we're talking about,
  367. 14:45right? You need a compromise candidate
  368. 14:48that works against the whole set of
  369. 14:50scenarios, not something that was
  370. 14:52optimized for one of them.
  371. 14:55So, this idea of aggregating multiple
  372. 14:57solutions for multiple scenarios is
  373. 15:00bonkers.
  374. 15:01You should not do that.
  375. 15:04Okay. Now, that leaves us with the
  376. 15:07obvious other choice, which is, well, I
  377. 15:09have all of my different scenarios, but
  378. 15:12my solver only wants one.
  379. 15:15Fine. I'll somehow aggregate the input
  380. 15:18data.
  381. 15:19And I can do that, of course, by asking
  382. 15:22my machine learning department to please
  383. 15:24give me the expected values
  384. 15:27for each of the data points that I'm
  385. 15:29going to need.
  386. 15:30Okay? So,
  387. 15:32that is called predict then optimize and
  388. 15:36it is fundamentally flawed.
  389. 15:39Um so, I mean, what what you're doing
  390. 15:41here is this, right? So, you have this
  391. 15:43plethora of scenarios, you aggregate
  392. 15:45them to one, hand it to the solver, get
  393. 15:47one solution, and you're a happy
  394. 15:49trooper.
  395. 15:50Um this is called predict then optimize,
  396. 15:53right? As I was mentioning.
  397. 15:55Um and but, you know, it it's it's going
  398. 15:57to work and it's in fact the thing that
  399. 15:59most companies do.
  400. 16:01It's a fundamentally flawed approach.
  401. 16:04Why?
  402. 16:05Well, before I'm going to run through
  403. 16:07another actual example with you
  404. 16:09with numbers and everything,
  405. 16:11um here's a mental picture that I want
  406. 16:13to to open. Think penalty shots in
  407. 16:16soccer. If you've never done this, it
  408. 16:18means that, you know, somebody gets to
  409. 16:20kick the ball against this very big goal
  410. 16:23and you put a goalie in there and if
  411. 16:25it's in, it's a goal and if if not, then
  412. 16:27it's not a goal, right? So, that's a
  413. 16:29penalty shot. You're alone against the
  414. 16:32goalie and you have a very very good
  415. 16:33chance of actually scoring a goal, which
  416. 16:36is a big deal in soccer.
  417. 16:38Now, we have data. The Economist
  418. 16:41actually analyzed 434 individual
  419. 16:44penalties
  420. 16:45that were shot in 44 World Cups and
  421. 16:47European Championships games. Okay? And
  422. 16:51you see here that some of them were
  423. 16:52misses, right? Where it didn't even hit
  424. 16:54the goal. Uh some of them were saves,
  425. 16:57right? Where the goalie got them and all
  426. 16:59of the other ones are actually in.
  427. 17:01Okay? So, now from there, of course, we
  428. 17:04can compute
  429. 17:06where the average penalty shot will be
  430. 17:08sent. And that's what you're asking the
  431. 17:11machine learning department to do. They
  432. 17:13see this wealth of different outcomes,
  433. 17:16but you asked them, would you please
  434. 17:19compress this to one scenario?
  435. 17:23Well, they have little other choice than
  436. 17:25to take the average. So, lo and behold,
  437. 17:29here is the average
  438. 17:31penalty shot
  439. 17:33uh in one of the 44 World Cups and
  440. 17:36European Championships um between '76
  441. 17:40and 2016. Okay? This is what it looks
  442. 17:43like.
  443. 17:48Very very boring shot.
  444. 17:50Slightly up and into the middle.
  445. 17:54Now, if you base your optimal strategy
  446. 17:58on that prediction, on that one scenario
  447. 18:02that you forced your machine learners to
  448. 18:03give you, well, the optimal strategy for
  449. 18:06your goalie would be just stay put.
  450. 18:08The ball will come straight to you. All
  451. 18:10you have to do is take it and you're
  452. 18:12going to do this.
  453. 18:14We know what the reality of that
  454. 18:16strategy looks like. It looks like this.
  455. 18:20You're not going to catch a single ball.
  456. 18:24Now, you
  457. 18:25you know, again, this is in to create a
  458. 18:27mental picture in in your head, but this
  459. 18:30happens in businesses all the time and
  460. 18:33it costs businesses millions of dollars.
  461. 18:36So, here's an example. We do inventory
  462. 18:38relocation. We have a big warehouse here
  463. 18:41somewhere in the port where most of our
  464. 18:43inventory is and we can send a truck to,
  465. 18:47you know, a warehouse somewhere
  466. 18:50somewhere um in the middle of the
  467. 18:52country,
  468. 18:53which is smaller and we you know,
  469. 18:55there's only 110 units. I think these
  470. 18:57were air conditionings that we can put
  471. 18:59onto the truck.
  472. 19:01Costs us $350 to run the truck, but, you
  473. 19:04know,
  474. 19:08and then it's much much cheaper to
  475. 19:11service clients that are in the vicinity
  476. 19:13of that smaller warehouse.
  477. 19:16Okay? So, if I do this, if I actually
  478. 19:19send a unit of product one to this
  479. 19:23remote warehouse here
  480. 19:24in red, then I'm saving on shipping
  481. 19:27costs, I'm saving $10
  482. 19:29for each of those products. And for
  483. 19:31product two, it's still $5. On the other
  484. 19:33hand, if I send it and then it doesn't
  485. 19:36get sold within a certain period of
  486. 19:37time, then I have to pay overstock
  487. 19:40costs, right? So, it costs me money to
  488. 19:42actually um put the inventory here into
  489. 19:45this remote location. It costs me three
  490. 19:47bucks more than it would have cost me at
  491. 19:49the at the port to have it there.
  492. 19:52And similarly for product two, um I pay
  493. 19:55a $2 penalty. Running the trucks cost
  494. 19:57350, by the way.
  495. 19:59So, now the question is
  496. 20:02do we send a truck? And if we do send a
  497. 20:04truck, how many of each product are we
  498. 20:06going to put in there?
  499. 20:08Now, this is obviously a question that
  500. 20:09you cannot answer unless you estimate
  501. 20:12how many of each product are going to be
  502. 20:14sold.
  503. 20:16So, let's go to the machine learning
  504. 20:18department. And because we don't want
  505. 20:21multiple futures, we're going to do
  506. 20:23predict and optimize. We're going to ask
  507. 20:25them to give us the expected number, the
  508. 20:27expected demand for each product
  509. 20:31at this remote location.
  510. 20:35Now, these guys are awesome. They give
  511. 20:37you numbers that are 100%
  512. 20:40accurate.
  513. 20:41This is correct. For product one, our
  514. 20:44expected demand is 19 units. And for
  515. 20:47product two, the expected demand is 91
  516. 20:51units.
  517. 20:53100% correct.
  518. 20:55Telling you this right now. There's no
  519. 20:57There's no ambiguity here. The expected
  520. 21:00demand for product one is 19. The
  521. 21:01expected demand for product two is 91.
  522. 21:05Beautiful. Let's hand this over to the
  523. 21:07OR department.
  524. 21:09So, what the OR department is going to
  525. 21:11do is going to set up a simple MIP.
  526. 21:13Right? So, you have one zero one
  527. 21:15variable T, which tells us whether we're
  528. 21:17going to run the truck. If we do that,
  529. 21:18costs us $350.
  530. 21:20And then we have variables P1 and P2,
  531. 21:23which tells us well, how much of product
  532. 21:24one and how much of product two are we
  533. 21:26going to relocate.
  534. 21:27Right? If we relocate it and it gets
  535. 21:30sold, then we save 10 bucks for product
  536. 21:32one, five bucks for product two.
  537. 21:33Beautiful. On the other hand, if we have
  538. 21:35overage, so we're going to compute
  539. 21:38whether something is going to be stuck
  540. 21:40in that warehouse, then we have to
  541. 21:42subtract the 10 bucks that I just gave
  542. 21:44you optimistically. Um
  543. 21:47and then pay the three bucks extra for
  544. 21:51for the additional storage cost in that
  545. 21:53remote warehouse. And similarly for OR
  546. 21:55two, we have to subtract the five again
  547. 21:57and then the two for the for the
  548. 22:00storing.
  549. 22:01>> [snorts]
  550. 22:01>> And here
  551. 22:02you And And of course, you know, the the
  552. 22:04total number of units that we can
  553. 22:06relocate is bounded by 110.
  554. 22:09And we do have our demand estimate here.
  555. 22:13Right? So, we're going to get overage if
  556. 22:15we're going to send more than 19 units
  557. 22:18for product one. And for product two, if
  558. 22:21we're sending more than 91 units.
  559. 22:24Unsurprisingly, the optimal solution
  560. 22:27here is, well, send 19 units of product
  561. 22:29one and 91 units of product two. You
  562. 22:32expect absolutely no overage, right?
  563. 22:35Because this is exactly the exact the
  564. 22:37the
  565. 22:38expected demand. And you're going to
  566. 22:41send the truck. And overall, you're
  567. 22:42going to save $295.
  568. 22:47Beautiful. Now, we do that.
  569. 22:50And
  570. 22:51we observe over time how much money
  571. 22:53we're actually saving. And it turns out,
  572. 22:56well, actually, we're not saving $295
  573. 22:59over just serving everything from the
  574. 23:01port.
  575. 23:03In reality, we're only saving $131.
  576. 23:07And to add insult to injury, the local
  577. 23:10planner the the manual planner that we
  578. 23:11used to have used to send 10 units of
  579. 23:14product one and 100 units of product
  580. 23:16two.
  581. 23:18And this person would get $187 in
  582. 23:22savings.
  583. 23:24So, what's going on? We had a perfect
  584. 23:26forecast. This is the expected number of
  585. 23:29units that will be sold as 19 and 91.
  586. 23:33And then we we we used those exact
  587. 23:36numbers, those correct numbers, and gave
  588. 23:38them to the OR department. And they came
  589. 23:41back with a provably optimal solution.
  590. 23:44There's no discussion about it. This is
  591. 23:46the correct solution for that MIP.
  592. 23:50And nevertheless, you're getting
  593. 23:53you know, the the the the hand planner
  594. 23:55makes 40% savings more
  595. 23:58than the MIP.
  596. 23:59What What's going on here? Well, what's
  597. 24:02going on is that nobody bothered to look
  598. 24:05at the demand distribution.
  599. 24:07Right? You compressed this distribution
  600. 24:11and made it one number.
  601. 24:1419 units of product one.
  602. 24:1791 units of product two.
  603. 24:20This is the same thing as taking that
  604. 24:21economist data on the penalty shots and
  605. 24:24saying, well, the expected penalty shot
  606. 24:26will go directly in the middle. It's the
  607. 24:27exact same thing.
  608. 24:29So, what you see here is that, well,
  609. 24:31actually, you know, there is a 90%
  610. 24:33chance that you're around 10 and 100
  611. 24:37for the demand, right? Product one, 10
  612. 24:39units. Product two, 100 units. And but
  613. 24:42there's a 10% chance that it flips
  614. 24:43because some influencer put out a video
  615. 24:46and says, well, I really like product
  616. 24:47one, and then suddenly it flips. Right?
  617. 24:50So, here we go.
  618. 24:51This is the individual
  619. 24:53probability density for the two
  620. 24:55products. It makes more sense to look at
  621. 24:58them
  622. 24:59in the joint forecast. So, you can see
  623. 25:01here that you're somewhere here around
  624. 25:03the diagonal, which is always 110 units
  625. 25:06that that you're going to to to sell. Um
  626. 25:10but you see that this 10% outlier over
  627. 25:12here skews the whole average.
  628. 25:16And I told you that those are the
  629. 25:17correct averages. It's 19 and 91. But it
  630. 25:20skews it
  631. 25:22towards this point over here, away from
  632. 25:24where 90% of the cases are actually
  633. 25:26happening, which are always 10 and 100.
  634. 25:30In fact, 19 and 91 never happens.
  635. 25:34Right? This is This is not something
  636. 25:35that has actually in the cards. This has
  637. 25:37never happened before. It's just that
  638. 25:39you you kind of skewed this because
  639. 25:41sometimes it could be instead of 10 100,
  640. 25:43it could be 110.
  641. 25:45Okay?
  642. 25:46So, if you look at that distribution,
  643. 25:49you understand why the discrepancy is.
  644. 25:52So,
  645. 25:54the curious thing is that the folks who
  646. 25:56have these
  647. 25:58legacy solvers who say, well, give me
  648. 26:00one input and I give you one provably
  649. 26:02optimal output, are going to sneer at
  650. 26:04you if you're telling them, well, you're
  651. 26:06going to use something that doesn't come
  652. 26:07with a proof of optimality cuz they will
  653. 26:09say, "Ha, what if you're 3% suboptimal?
  654. 26:12What if you're 5% suboptimal? You're
  655. 26:14losing so much money." Well, turns out
  656. 26:17that they just lost you 30%.
  657. 26:21Not because the optimization was wrong,
  658. 26:24but because it forced that you're
  659. 26:27compressing
  660. 26:29the the the wealth of futures that could
  661. 26:31fit you
  662. 26:33into that one scenario forecast.
  663. 26:38So, you should definitely
  664. 26:40prefer a heuristic solution such as, you
  665. 26:43know, for example, 5% suboptimal here
  666. 26:45for the actual real model,
  667. 26:48you'd still make $177 instead of 131.
  668. 26:52You should always prefer heuristic
  669. 26:54solution to a realistic model,
  670. 26:56particularly a model that can handle the
  671. 26:58stochasticity of your problem, over an
  672. 27:00exact solution to the approximated
  673. 27:03model, which forces you to put one
  674. 27:05compressed scenario inside.
  675. 27:07Very important to remember that.
  676. 27:11So, forget predict and optimize. It is a
  677. 27:13destroyer of businesses. It costs you
  678. 27:16millions to follow this approach.
  679. 27:19Now, people know this. People in OR have
  680. 27:21known this for a long time that this is
  681. 27:23a terrible idea, even though legacy
  682. 27:26optimizer representatives will still
  683. 27:28tell you to do exactly that.
  684. 27:30I can guess why because they can only
  685. 27:33handle one scenario inputs.
  686. 27:36So, the idea came up,
  687. 27:39what if we kind of skew the input?
  688. 27:41So, maybe we can aggregate the whole
  689. 27:43thing,
  690. 27:44but in such a way that the solver the
  691. 27:47the the optimizer is somehow nudged in
  692. 27:50the right direction to give you a
  693. 27:51solution which would actually do well
  694. 27:55against the whole cloud of potential
  695. 27:57futures. That is the idea
  696. 28:02of end-to-end learning or predict and
  697. 28:04optimize. Why is it called end-to-end
  698. 28:06learning? Well, you see that before
  699. 28:07you're actually doing the real thing.
  700. 28:09But it basically means that you know, as
  701. 28:12you're looking at the at the history,
  702. 28:14right? So, if if you're looking at the
  703. 28:15data
  704. 28:16and you build a model that makes the
  705. 28:18forecast, which you're going to shove
  706. 28:19into the optimization model.
  707. 28:22As you're doing this, you're going to
  708. 28:23modify this model
  709. 28:25um here um
  710. 28:27taking into account what the regret is
  711. 28:31with respect to some of these scenarios
  712. 28:33that could have also happened um
  713. 28:36when you look at the plan that comes out
  714. 28:38if you gave it a certain input.
  715. 28:42Right? So, so you do this offline
  716. 28:44because I'm making this online makes no
  717. 28:46sense at all because it means that you
  718. 28:48would actually solve the whole thing.
  719. 28:49But you do this offline to somehow bias
  720. 28:52the forecasting model in such a way that
  721. 28:54it's going to nudge the solver
  722. 28:57to give you something that's actually
  723. 28:59way more robust than you would get
  724. 29:01otherwise. And hopefully get better
  725. 29:03performance that way. It's called
  726. 29:04predict and optimize. And this process
  727. 29:07here of modifying the forecasting method
  728. 29:10is called end-to-end learning.
  729. 29:14So, we can use the exact same
  730. 29:17same example that I just gave you
  731. 29:21to show you that this is bonkers. The
  732. 29:24whole idea is bonkers to do it this way.
  733. 29:26Why?
  734. 29:27Look at this example.
  735. 29:29Well, naturally, as long as your
  736. 29:32forecast is a total demand of 110 units,
  737. 29:35which is, you know, historically, that's
  738. 29:38always what happened. It's always around
  739. 29:39110 units.
  740. 29:41No matter what you put there as your
  741. 29:43forecasted demand for product one and
  742. 29:46product two,
  743. 29:47you're going to get as an optimum
  744. 29:51that same number.
  745. 29:54Let me say that again. Yes, you can
  746. 29:56nudge the solver to give you the correct
  747. 29:58answer, which is 10 and 100,
  748. 30:01by essentially telling the solver, "Hey,
  749. 30:04the correct answer would be 10 and 100."
  750. 30:08So, what's the point of optimizing if
  751. 30:10the machine learning itself has to know
  752. 30:13the optimal solution to nudge the solver
  753. 30:16properly to give you the right answer?
  754. 30:20Forecast would need to predict the
  755. 30:22optimal solution if this end-to-end
  756. 30:24learning was supposed to be working.
  757. 30:27Nuts. It's It's just nuts. I I can't say
  758. 30:30it any differently. Now, there's there's
  759. 30:33one
  760. 30:34case
  761. 30:35where you might think, "Hey, you know,
  762. 30:37aggregating is actually not a bad idea."
  763. 30:39And that is if the uncertainty in your
  764. 30:43optimization problem lies purely in the
  765. 30:46objective function. Right? So, in this
  766. 30:48case here, I use capital C and D in
  767. 30:50order to
  768. 30:52make clear that these are random
  769. 30:54variables, so we don't know exactly
  770. 30:57what the profit coefficients for X and
  771. 31:00for Y are, and there's no uncertainty in
  772. 31:03the constraint structure at all.
  773. 31:05Right? Note that before we had
  774. 31:06uncertainty in the constraint structure,
  775. 31:08right? This the example that we had.
  776. 31:10Imagine that that wasn't the case. The
  777. 31:11uncertainty lies purely in there. Well,
  778. 31:13then mathematically, you can show very
  779. 31:15easily that you can move that
  780. 31:16expectation directly into the
  781. 31:18coefficients of each variable.
  782. 31:20Right? So, everything is linear over
  783. 31:22here. So, you can put the expectation of
  784. 31:25C and the expectation of D in there, and
  785. 31:27now if you solve this problem,
  786. 31:30by the way, no end-to-end learning
  787. 31:31required, right? You You just take the
  788. 31:33expectation, you're done with it.
  789. 31:35Um
  790. 31:36if you if you do that, you're going to
  791. 31:38get you're going to get a solution that
  792. 31:40maximizes the expected profit. Right?
  793. 31:44So, this is for maximizing profit.
  794. 31:48But there's one caveat.
  795. 31:50And that is
  796. 31:52that in business, you
  797. 31:53rarely care about just expectations. You
  798. 31:57also care about the variability of your
  799. 32:00solutions. Remember when I said, well,
  800. 32:02whatever solution you're going to get,
  801. 32:03you're going to get a distribution of
  802. 32:05outcomes, matters a great deal for
  803. 32:08businesses. So, in this perfect world,
  804. 32:12where everything is linear and our our
  805. 32:16uncertainty lies purely in the objective
  806. 32:18function, let's look at this example
  807. 32:20over here.
  808. 32:22So, what what are we supposed to do
  809. 32:23here? Essentially, we're supposed to
  810. 32:24maximize all the Z eyes. And the Z eyes
  811. 32:28basically have to be lower equal one,
  812. 32:29but we can make some additional room
  813. 32:32here. So, for example, if we set the Y
  814. 32:35eyes to minus one, so the Y eyes can
  815. 32:38live anywhere between minus one and one,
  816. 32:40then Z I could go uh for every I in I1,
  817. 32:44Z I could go to two. And similarly, if
  818. 32:46we're setting a Y I to one, plus one,
  819. 32:51for every I in the other set, um I2, uh
  820. 32:55then again, we can set the corresponding
  821. 32:57Z eyes to two.
  822. 32:59So, we can make additional room by
  823. 33:02setting the Y eyes accordingly. Okay? Um
  824. 33:06thing is,
  825. 33:08we don't know how much it costs us. This
  826. 33:10could be a good thing
  827. 33:12um if if we're actually um
  828. 33:15setting Y I to one or minus one, but we
  829. 33:17don't know beforehand because this part
  830. 33:20of the objective function gets
  831. 33:21multiplied with a
  832. 33:23with random coefficients that are drawn
  833. 33:27uh from Gaussian distributions with
  834. 33:29expected value zero. So, on
  835. 33:31expectations, we're expecting the Y eyes
  836. 33:33to cost nothing,
  837. 33:34but they do have a standard deviation of
  838. 33:36one.
  839. 33:37Okay?
  840. 33:38So, this is the problem. Now, I just
  841. 33:41showed it to you mathematically, it is
  842. 33:43correct to just take the expected value
  843. 33:45for these X's, so just make it zero.
  844. 33:48Now, you don't even have any uncertainty
  845. 33:50here anymore.
  846. 33:52It means you can just maximize the Z
  847. 33:54eyes. So, what you're going to do is set
  848. 33:56the Y eyes to the corresponding values
  849. 33:58all the Z eyes are being set to, and now
  850. 34:01you're going to get n twos divided by n,
  851. 34:03that gives you two. Right? 2n divided by
  852. 34:06n gives two.
  853. 34:07And you get an expected value of two.
  854. 34:11Okay? That's the optimum for the
  855. 34:13expectation that you're going to get.
  856. 34:16But look at the variance of your
  857. 34:18solution. The variance of this is going
  858. 34:20to be n.
  859. 34:24So, if you have n variables over here,
  860. 34:28you're going to get a a a massive spread
  861. 34:32of outcomes
  862. 34:34of for your
  863. 34:37for your final result. And if this is
  864. 34:39somehow profit, right? So, let's say
  865. 34:41this is $2 million,
  866. 34:43you don't suddenly want to end up with
  867. 34:45minus $18 million.
  868. 34:48Right? For a business, this is this is
  869. 34:50breaking your neck. You can't have that.
  870. 34:53So, what you would actually want to do,
  871. 34:56depending on how sensitive to risk you
  872. 34:58are over here, um what you would
  873. 35:00actually like to do is to set the whole
  874. 35:03set every Z I to one and set all the Y's
  875. 35:07to zero in order to take the risk out of
  876. 35:10the equation.
  877. 35:12Now, you see this this image that I give
  878. 35:15you over here. What is that? Well, we
  879. 35:16have n Gaussian variables. And And I
  880. 35:19want to challenge your intuition a
  881. 35:21little bit.
  882. 35:23Yes, these n Gaussian distributed
  883. 35:26variables all have an expected value of
  884. 35:28zero.
  885. 35:29So, now you might think, well, most of
  886. 35:31the time I'm going to get something that
  887. 35:33is kind of like really close to zero,
  888. 35:35and then as I move out here um further
  889. 35:39away from the expected case, by the way,
  890. 35:42this is really bad terminology, as
  891. 35:44you'll see in a moment, um the more
  892. 35:46unlikely it gets that I'm going to see
  893. 35:48that. The reality is completely
  894. 35:51different. What you're going to have is
  895. 35:53that you will you will have a an X
  896. 35:56vector that has length square root of n.
  897. 35:59So, if the if there are 100 Z eyes and
  898. 36:01100 Y eyes, you're going to get a length
  899. 36:0410 vector
  900. 36:06for the X's that will be there, right?
  901. 36:08So, square root of 100 is 10.
  902. 36:10Um and
  903. 36:12most of the probability density is going
  904. 36:14to lie on this very, very small shell
  905. 36:19out here. Right? So, you kind of you
  906. 36:21know what the radius is going to be with
  907. 36:23very high probability, and you have
  908. 36:25basically no chance, not in your
  909. 36:28lifetime, that you're going to ever see
  910. 36:31the expected case, quote unquote, where
  911. 36:34all the X's would be zero.
  912. 36:37Just never happens.
  913. 36:39And but But this is the the case that
  914. 36:42you were preparing for.
  915. 36:45And the shell here,
  916. 36:47where those where those vectors actually
  917. 36:49are lying, right? So, this is like an
  918. 36:51n-dimensional sphere, of course,
  919. 36:53that shell is what you should have
  920. 36:55prepared for.
  921. 36:57Now, think back to the goalie, right?
  922. 37:00You have the chance that the ball will
  923. 37:01go left, that the ball will go right.
  924. 37:03You stayed right in the middle. And this
  925. 37:04is exactly what you keep doing when you
  926. 37:07use predict and optimize,
  927. 37:10um or predict then optimize, because you
  928. 37:13you kind of you're nudging the whole
  929. 37:15thing in the middle. There's There's
  930. 37:17really no other way to nudge this
  931. 37:19because you you just don't know where
  932. 37:21the X's are going to fall, right?
  933. 37:23So,
  934. 37:24there you go. You You just cannot
  935. 37:26control
  936. 37:28for losses
  937. 37:30if you whenever you compress the input
  938. 37:32data to one scenario, you have no
  939. 37:36ability to control the outcome, to
  940. 37:38control the risk that is associated with
  941. 37:41it.
  942. 37:42So, here's another example that
  943. 37:44illustrates this very simple two
  944. 37:46variables. Essentially, we're supposed
  945. 37:48to maximize Y2. So, Y2 is is vertical
  946. 37:51and Y1 is horizontal. Um and the
  947. 37:54feasible thing here is is a very thin
  948. 37:56strip here that goes from minus 60 to
  949. 37:58250,
  950. 38:00um and then that, you know,
  951. 38:02small little tilt here that's taken
  952. 38:04away, and the maximum you can do for Y2
  953. 38:07is five. Right? So, it's a very thin
  954. 38:09strip.
  955. 38:10No matter where you're going to go
  956. 38:12for any feasible Y1,
  957. 38:15you're always going to get an expected
  958. 38:17value of five
  959. 38:20for this optimization problem.
  960. 38:22Right? This
  961. 38:23You can't do better than five for this
  962. 38:25one here.
  963. 38:26You could hope that the Y1 contributes
  964. 38:29something, but again, its coefficient is
  965. 38:32drawn sample drawn randomly from a
  966. 38:35Gaussian with expected value zero. So,
  967. 38:38on expectation, you're not going to get
  968. 38:39anything else for Y1.
  969. 38:43So, where would you going to go?
  970. 38:45Where would you go? I mean, if you take
  971. 38:46a simplex algorithm to solve this thing,
  972. 38:49no matter how you nudge it, you can set
  973. 38:51this to zero, you can set this to minus
  974. 38:52one, you can set this to plus one. It
  975. 38:54doesn't matter where you nudge it,
  976. 38:55you're going to end up with one of those
  977. 38:57two solutions. Either you're going to be
  978. 39:00plus five for Y2 and then whatever that
  979. 39:03is here, minus 47 or whatever. Um I'm
  980. 39:07sorry, my minus minus 57 or something
  981. 39:10like this. Um or
  982. 39:12plus
  983. 39:14241 or something like this. Um
  984. 39:20you're going to get one of those two
  985. 39:21corners.
  986. 39:23And the right thing to do here, the
  987. 39:25right way to control the risk that is
  988. 39:27associated with the whole thing, is of
  989. 39:30course to put it right here where Y1 is
  990. 39:34zero.
  991. 39:35Just take out the risk all together. It
  992. 39:37doesn't hurt you one bit to do so, and
  993. 39:41predict and optimize, even in this super
  994. 39:43simplified case where all the
  995. 39:45uncertainties in the objective function
  996. 39:47and the whole problem is linear
  997. 39:49cannot give that solution to you.
  998. 39:52Which is why you should forget about the
  999. 39:54whole idea of predict and optimize right
  1000. 39:56now.
  1001. 39:59So, I mentioned before that in order to
  1002. 40:02make predict and optimize work in
  1003. 40:04general, you would have to be able to
  1004. 40:06predict the solution directly.
  1005. 40:10Now, in most cases this is not doable,
  1006. 40:12of course, because you have usually
  1007. 40:15complex constraint structures and
  1008. 40:17whatever the machine learning is going
  1009. 40:18to suggest will likely be infeasible.
  1010. 40:22So, this is not something that works.
  1011. 40:23But, the question is, well, what if you
  1012. 40:25have a problem where the constraint
  1013. 40:27structure isn't totally crazy?
  1014. 40:29Couldn't I directly forecast what I
  1015. 40:33ought to be doing?
  1016. 40:34Well, let's look at an example of this.
  1017. 40:37In fact, the example that spawned the
  1018. 40:39idea of
  1019. 40:40of starting InsideOut in the first
  1020. 40:42place. We were participating in a
  1021. 40:45competition, the H Guy 21 competition
  1022. 40:48on a price collection TSP.
  1023. 40:51Um, so, we are supposed to collect the
  1024. 40:53rewards. We can say which clients we're
  1025. 40:55going to visit. We have to do that in a
  1026. 40:57certain time window.
  1027. 40:59Um, and um, then we have to meet a
  1028. 41:02cutoff time at the end of the day when
  1029. 41:04we have to be back. Problem is, we don't
  1030. 41:06actually know how long it's going to
  1031. 41:08take us to go from A to B. We were given
  1032. 41:11a distribution that would tell us how
  1033. 41:13long that would take.
  1034. 41:15And then there were two tracks. In track
  1035. 41:17one, we were supposed to find one tour
  1036. 41:20that you had to stick to.
  1037. 41:23You couldn't modify it once you started
  1038. 41:25it, no matter how the travel times were
  1039. 41:27evolving during the day. You had to
  1040. 41:29stick to it. And you were supposed to
  1041. 41:31come up with a tour that has great
  1042. 41:33expected value. And then in track two,
  1043. 41:36we were supposed to learn a policy. We
  1044. 41:39were supposed to use reinforcement
  1045. 41:41learning in order to say, "Hey, if
  1046. 41:43you're in this state, maybe this is the
  1047. 41:45right client to go to next in order to
  1048. 41:47pick up a reward." And then, of course,
  1049. 41:50make sure that you will be back at the
  1050. 41:51depot at the end of the
  1051. 41:54of the day.
  1052. 41:56So, um,
  1053. 41:57if you use the track one solution where
  1054. 42:00you're stuck
  1055. 42:01using the same solution every single
  1056. 42:03time, no matter how the day unfolds, um,
  1057. 42:07we could get a solution that is like
  1058. 42:1010.81, right? So, this is price
  1059. 42:11collection. So, you wanted to maximize
  1060. 42:13this whole thing. Um, and this was the
  1061. 42:16value that you would have. Now, if you
  1062. 42:18had had perfect clairvoyance for each
  1063. 42:20instance,
  1064. 42:21um, where you know beforehand these will
  1065. 42:24be the travel times that are going to
  1066. 42:25hit you, well, then you could have
  1067. 42:28computed a the best tour for every
  1068. 42:31single day for every scenario that was
  1069. 42:33there, right? So, there's a theoretical
  1070. 42:35maximum what you can achieve.
  1071. 42:37Now, you would expect that the track two
  1072. 42:39solution, reinforcement learning, which
  1073. 42:41can adapt its solution with more
  1074. 42:44information how long it did actually
  1075. 42:46take to get so far, right? That that
  1076. 42:49solution would be able to, you know, lie
  1077. 42:52somewhere here between the Uncle Otto
  1078. 42:55solution, which stubbornly always takes
  1079. 42:57the same tour no matter how late it
  1080. 42:59gets,
  1081. 43:00um, and the perfect clairvoyant
  1082. 43:03solution.
  1083. 43:04Right? Somewhere in between here you
  1084. 43:05would expect the reinforcement learning
  1085. 43:07to lie.
  1086. 43:08And we submitted the the solution to
  1087. 43:10this, um,
  1088. 43:13uh, which won the competition. So, we
  1089. 43:15did something that was very close to the
  1090. 43:16state of the art. A deep learning base,
  1091. 43:18we used an instance graph encoding. Um,
  1092. 43:21we did active learning. We wouldn't
  1093. 43:22just, you know, just take the whole
  1094. 43:24thing. We did rollout, so there was
  1095. 43:26dynamic search involved. And
  1096. 43:28nevertheless,
  1097. 43:29the reinforcement learning that we
  1098. 43:31submitted
  1099. 43:32would only get us to 10.7734.
  1100. 43:36Now, we knew that before we submitted
  1101. 43:38it, which is why we asked the, uh,
  1102. 43:41competition organizers whether we could
  1103. 43:43also just, you know, do a very simple
  1104. 43:46policy, which is like, well, stick to
  1105. 43:47Uncle Otto.
  1106. 43:49Um, they I wasn't allowed. So, they
  1107. 43:51said, "No, it has to be reinforcement
  1108. 43:53learning. You have to submit the code
  1109. 43:55and everything." So, we had to do that,
  1110. 43:57even though we knew that just sticking
  1111. 43:59to a stupid default was better than
  1112. 44:02reinforcement learning. So, now the
  1113. 44:03question is, why is that? I mean, this
  1114. 44:05is the perfect application scenario for
  1115. 44:07reinforcement learning. There are no
  1116. 44:09crazy side constraints that would
  1117. 44:11suddenly say, "Oh, your solution isn't
  1118. 44:12isn't isn't good." Nothing.
  1119. 44:15Well, it is because the reinforcement
  1120. 44:17learning is sometimes amazingly good. It
  1121. 44:20it it just, you know, for many of those
  1122. 44:22instances, it just found a perfect tour
  1123. 44:25for them.
  1124. 44:27But, every now and then it screwed up so
  1125. 44:29tremendously
  1126. 44:31that it was just too brittle to work
  1127. 44:33robustly. And that is what ruined the
  1128. 44:36the average performance in the
  1129. 44:38reinforcement learning.
  1130. 44:39Right? So, for the time being at least,
  1131. 44:43even if you're in a in a scenario where
  1132. 44:45you don't have complex side constraints
  1133. 44:48to deal with and things that just
  1134. 44:49require you to do to use optimization,
  1135. 44:52trying to forecast the optimal action
  1136. 44:54directly with machine learning
  1137. 44:57is a very bad idea
  1138. 44:59due to the brittleness of, uh,
  1139. 45:01reinforcement learning.
  1140. 45:03So, forget about predicting the solution
  1141. 45:06directly. Either you need optimization,
  1142. 45:09you need to use search in order to find
  1143. 45:11good plans.
  1144. 45:13So, what's left?
  1145. 45:14What what can we possibly do? Well, we
  1146. 45:17have to break out of the confines of
  1147. 45:19legacy solvers, which would only take
  1148. 45:21one input.
  1149. 45:23What you actually want to do is this.
  1150. 45:25You want to maximize in the feasible
  1151. 45:28space some non-linear function and then
  1152. 45:31some aggregation thereof, right? So,
  1153. 45:33remember that this function here under
  1154. 45:35the,
  1155. 45:36uh, under the different scenarios that
  1156. 45:38can hit you will have a variability to
  1157. 45:41it. No matter which plan you're going to
  1158. 45:43choose,
  1159. 45:45even if you use a legacy optimizer, even
  1160. 45:47if you do predict and optimize, you are
  1161. 45:49going to be hit with a variable outcome,
  1162. 45:52right? So, okay, now you get to decide
  1163. 45:56which outcome you want to optimize for.
  1164. 45:58Is it the expected value?
  1165. 46:01Is it the expected plus one standard
  1166. 46:03deviation? Is it the CVaR 5%? Right? It
  1167. 46:07is up to you to do that, which is why
  1168. 46:08sometimes we actually write it like
  1169. 46:10this, right? So, use some aggregator
  1170. 46:12over your objective function, which can
  1171. 46:14be non-linear and whatever you want.
  1172. 46:17Right? This is what you want to be
  1173. 46:18doing. You want your solver to be able
  1174. 46:22to look at all of these scenarios as
  1175. 46:24input and evaluate each course of action
  1176. 46:28each course of action internally
  1177. 46:31when weighing the different options that
  1178. 46:33are available to then spit out one good
  1179. 46:37compromise candidate that is going to
  1180. 46:39work well against this cloud of
  1181. 46:40solutions.
  1182. 46:42And that is what InsideOut Seeker is
  1183. 46:45going to give you.

About this transcript

This page contains the full transcript of Why predict-then-optimize and end-to-end learning won't fix your optimization under uncertainty. by InsideOpt Tutorials, generated from the public captions YouTube serves with the video. The transcript has 7,184 words across 1,183 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.