YouTube2Text

Deep Hedging for Quantitative Finance — Transcript

by Roman Paolucci · 5,691 words · 845 segments · language en · Watch on YouTube

Full transcript

  1. 0:00Roman, if we have all of this data, why
  2. 0:02aren't we just training artificial
  3. 0:04intelligence, computational agents to go
  4. 0:07about making trading decisions for us?
  5. 0:10We can give it a state vector, a vector
  6. 0:12of information containing news, spot
  7. 0:15prices, compressions of information that
  8. 0:17have already happened, path signatures,
  9. 0:20volatility, the list goes on. And that
  10. 0:23computational agent is going to make way
  11. 0:26better decisions than us. Why aren't we
  12. 0:30why aren't we doing that? Well, we're
  13. 0:33working on it. It is the frontier of the
  14. 0:35academic literature and that's exactly
  15. 0:38what we're going to talk about today.
  16. 0:40We're going to talk about this idea of
  17. 0:42deep hedging for quantitative finance
  18. 0:44pioneered of course by Hans Buer who I
  19. 0:48saw at the Bachier conference giving a
  20. 0:50presentation when I was a young quant
  21. 0:52and I didn't know my ash for my elbow
  22. 0:54and I was trying to figure out what the
  23. 0:56hell he was talking about from a
  24. 0:58marketmaking perspective and I had never
  25. 1:00seen the reinforcement learning
  26. 1:01literature at the time and I was trying
  27. 1:03to figure out what the hell he was
  28. 1:04talking about there as well. But
  29. 1:06fortunately since then after a whole
  30. 1:09bunch of whiteboarding with my
  31. 1:11colleagues at the time and a whole bunch
  32. 1:13of reading and of course implementation
  33. 1:16I have been able to understand the
  34. 1:20landscape and construct this lesson for
  35. 1:23you to save you all of the headaches
  36. 1:26that I had trying to understand this
  37. 1:29over the course of several years. Now,
  38. 1:31of course, the problem that we face is
  39. 1:34very intricate because reinforcement
  40. 1:36learning requires a substantial amount
  41. 1:38of data and iteration. It is a very
  42. 1:43there's really no other way to put it
  43. 1:45finicky problem. It's very difficult to
  44. 1:49approximate these policy functions as
  45. 1:52we're going to discuss. So even if you
  46. 1:55have a a state with a whole bunch of
  47. 1:57information, it's discretized. We exist
  48. 2:01in the continuous sense, but we can only
  49. 2:04feed a computational agent so much
  50. 2:07information over so many different time
  51. 2:09intervals. And very quickly, that's
  52. 2:13going to explode in terms of a problem,
  53. 2:16but we'll get there. Let's start on the
  54. 2:19trading desk. We're going to talk about
  55. 2:21this notion of riskneutral pricing. Why
  56. 2:24should we care at all about riskneutral
  57. 2:27pricing? Well, you want to make money,
  58. 2:29don't you? If you're a market maker,
  59. 2:31your objective is to provide liquidity
  60. 2:33to the market. We can debate about
  61. 2:36whether the role of a market maker is
  62. 2:38productive or not. I really don't care
  63. 2:40about that right now. I'm just talking
  64. 2:41about the broader business function. If
  65. 2:44that market maker is going to quote both
  66. 2:47sides of a market, right, a two-way,
  67. 2:50then what should the mid price be? What
  68. 2:54is the fair value of a derivative
  69. 2:56contract right now anyway? Well, we
  70. 3:00don't have a crystal ball. So, the best
  71. 3:02we can do is what? An expectation.
  72. 3:05Shocker. That is quite literally
  73. 3:08mathematically the best guess in the
  74. 3:10mean squared error sense. But where does
  75. 3:13that expectation fit into the broader
  76. 3:15notion of riskneutral pricing? Well,
  77. 3:18Black Scholes gives us this delta
  78. 3:20hedging argument in their framework.
  79. 3:23Assuming the underlying follows a
  80. 3:25geometric brownian motion, the market is
  81. 3:27complete. And if we can hedge away that
  82. 3:30randomness via a delta hedge, then we
  83. 3:33are going to be able to construct what
  84. 3:35is known as the value of a replicating
  85. 3:38portfolio. That is the fair value of the
  86. 3:41option contract today. There's no
  87. 3:43speculation here. All right? We're not
  88. 3:46talking about, oh, you know, the market
  89. 3:48risk premium is 7 and a half% on average
  90. 3:51risk adjusted or inflation adjusted
  91. 3:53year-over-year, so that's what we're
  92. 3:55going to use in the P measure sense to
  93. 3:57price our derivatives. No, no, no, no.
  94. 4:00This is riskneutral pricing. We are
  95. 4:03hedging that risk in the complete market
  96. 4:06sense. And on average, if we have no
  97. 4:09risk, we must earn the risk-free rate.
  98. 4:13So that's exactly what you see here as
  99. 4:15the fair price, the mid price of our
  100. 4:18derivative today. Because if we go out
  101. 4:21and we buy these contracts and we hedge
  102. 4:24all of that risk, then this is what we
  103. 4:27are going to earn on average. It's the
  104. 4:29riskneutral
  105. 4:31discounted expected payoff and that's
  106. 4:34exactly what you see here. Now,
  107. 4:37typically when I introduce this idea to
  108. 4:40students, it is a cluster I'm not
  109. 4:43gonna put it lightly. It's a cluster
  110. 4:45because they're like, "So, you
  111. 4:48always earn the risk-free rate?" It's
  112. 4:50like, it's like, "No, you earn it in
  113. 4:52expectation on average." So, where do I
  114. 4:55like to start? I always like to start in
  115. 4:57the stochcastic world in in the fineman
  116. 4:59CAC sense and then take that
  117. 5:03visualization and move it over to the
  118. 5:06delta hedging argument to show exactly
  119. 5:09what's going on. And that's exactly what
  120. 5:11I have for you here today in terms of an
  121. 5:13animation. So on the right you're going
  122. 5:15to observe the option price. Okay. Now,
  123. 5:18what we're going to do to try to figure
  124. 5:20out the value of the option contract is
  125. 5:22we're going to simulate the underlying
  126. 5:24asset and we're going to compute the
  127. 5:27payoff of that option for this one
  128. 5:29sample path. Then what we're going to do
  129. 5:32is effectively add it to a list. And
  130. 5:34then we're going to do it again and
  131. 5:36again and again and just take the
  132. 5:37average of that list. This is Monte
  133. 5:39Carlos simulating the options price
  134. 5:42which we know is going to converge to
  135. 5:43the black trolls price which you see
  136. 5:45here as a red dash line. That's fineman
  137. 5:48CAC. So if I play this animation for you
  138. 5:50here, it's not going to be a shocker. As
  139. 5:52we realize more and more and more and
  140. 5:54more price paths, you're going to see
  141. 5:56the option price converge on the right
  142. 5:58to the black Scholes price. Again, not a
  143. 6:01shocker. That's fine CAC. So what you're
  144. 6:04observing here is essentially an
  145. 6:07approximation for the fair value of the
  146. 6:10option price today. But this is not the
  147. 6:14black schles delta hedging argument.
  148. 6:16Okay, I need to be very clear here. I
  149. 6:19introduce Monte Carlo simulating the
  150. 6:22option price first because it's very
  151. 6:23easy to understand and mathematically
  152. 6:26again thanks to Fman and CAC that piece
  153. 6:29of the puzzle is automatically
  154. 6:31equivalent to the result from Black
  155. 6:34Scholes. They have to be equivalent.
  156. 6:37That partial differential equation
  157. 6:38representation is equivalent to the
  158. 6:41conditional expectation in the
  159. 6:43stochastic world. it has to be and I
  160. 6:46actually have a dedicated video deriving
  161. 6:49that result if you would like to check
  162. 6:52it out. So understanding that this is
  163. 6:55the approximation for this riskneutral
  164. 6:57conditional expectation
  165. 6:59discounted of course for the option
  166. 7:01price now we can go about validating it
  167. 7:04in the black scholes framework. All
  168. 7:08right so what are we going to do in a
  169. 7:10black trolls framework? Well, what we're
  170. 7:12going to do is we are going to go out
  171. 7:14and we are going to buy or sell this
  172. 7:18contract. All right? And the contract
  173. 7:21that we are going to be holding then is
  174. 7:24going to be delta hedged. So on the left
  175. 7:27here, what I want to show you is I want
  176. 7:29to show you the underlying asset. So
  177. 7:31this is the same chart that you saw
  178. 7:33above. It's the underlying asset
  179. 7:35simulation. But on the right, I want to
  180. 7:38show you our equity curve. All right, we
  181. 7:40go out and we're going to go long or
  182. 7:43short the option at this theoretical
  183. 7:46fair price. And then throughout the life
  184. 7:48of the option, we are going to be
  185. 7:51continuously delta hedging. What does
  186. 7:55this actually look like? If I play this
  187. 7:57animation for you here, I want you to
  188. 7:59observe on the left, right, we have all
  189. 8:02of these market paths and you can see
  190. 8:04that the drift of the market paths on
  191. 8:07average is going to be roughly what?
  192. 8:10It's going to be roughly the risk-free
  193. 8:11rate. All right, that is the dynamics
  194. 8:14that we're using for the simulation here
  195. 8:15of the geometric brownie emotion. On the
  196. 8:18right, you can see as we experience more
  197. 8:21and more and more and more delta hedged
  198. 8:24paths, what happens to the average P&L
  199. 8:27for each trade? It becomes zero. That is
  200. 8:32exactly what this replicating portfolio
  201. 8:34argument is suggesting. Functionally and
  202. 8:38visually, this is exactly what the
  203. 8:40implication is. Okay. So if we go out
  204. 8:43and we transact at the fair value of
  205. 8:46this financial derivative, not fair
  206. 8:48value out of thin air, fair value based
  207. 8:51on this framework in the stochcastic
  208. 8:54sense or this framework in the black
  209. 8:56shaw sense, then we're going to if we
  210. 8:59delta hedge throughout the life of that
  211. 9:02trade realize net zero P&L on average.
  212. 9:08All right. On average,
  213. 9:11this makes it very easy to see where
  214. 9:14we're going and the sense of market
  215. 9:16making. What are you going to be a
  216. 9:17liquidity provider if on average this
  217. 9:20yellow line your empirical expected P&L
  218. 9:24is going to be zero? No, you're not
  219. 9:28going to transact at mid. Everyone pulls
  220. 9:29up Yahoo Finance or Bloomberg
  221. 9:31Terminal Interactive Brokers and we go,
  222. 9:34"Oh, you know, Nvidia is $220."
  223. 9:37um oh you know the spy ETF it's like
  224. 9:39it's like no that's that's the mid price
  225. 9:42right that's the mid price last week I
  226. 9:44did a video on making a market on a onem
  227. 9:47run to show you exactly the implications
  228. 9:49of this from a market making perspective
  229. 9:52and what we're seeing here is if desk
  230. 9:54P&L is flat on average in expectation
  231. 9:59nobody's going to provide liquidity so
  232. 10:01what are you going to do you're going to
  233. 10:02quote two way you're going to be willing
  234. 10:04to what buy at the bid
  235. 10:07sell at the ask. And what that's going
  236. 10:09to do is it's going to artificially
  237. 10:11induce
  238. 10:13an edge. And that edge means that this
  239. 10:17desk P&L is going to have a positive
  240. 10:20trajectory. All right.
  241. 10:23But what's the catch? The catch is this
  242. 10:27only holds with continuous delta
  243. 10:29hedging. In practice, you can't delta
  244. 10:32hedge continuously.
  245. 10:34You have to hedge what? Indiscreet time.
  246. 10:38Otherwise, you're going to accumulate an
  247. 10:39infinite number of transaction costs,
  248. 10:40right? And that's going to eat away at
  249. 10:43your profits. Okay. Moreover, you have
  250. 10:46model risk, inventory risk, blah blah
  251. 10:48blah, all the other stuff, too. But
  252. 10:49we'll talk about that down the line. All
  253. 10:52right. Let's talk about market and the
  254. 10:54hedging problem. If we understand the
  255. 10:56idea of risk neutral pricing and why we
  256. 10:57care about it from a desk perspective,
  257. 10:59well, if we know what the fair price is,
  258. 11:01we know we can get to net zero P&L. And
  259. 11:03then if we have a spread on top of that,
  260. 11:05we're going to accumulate an edge. But
  261. 11:06what is an edge anyway, let's quickly go
  262. 11:09to the casino. All right, this is an
  263. 11:12example of a fixed edge. All right, in a
  264. 11:17casino, if we look at something like
  265. 11:18roulette, the players will always have a
  266. 11:22negative edge. The casino will always
  267. 11:25have a positive edge. All right? As long
  268. 11:29as the casino doesn't reach an absorbing
  269. 11:32state of zero wealth where they're in
  270. 11:34ruin and they can't come back from it,
  271. 11:36they are guaranteed to accumulate
  272. 11:38infinite wealth. Man, that's the
  273. 11:42government's wet dream in terms of tax
  274. 11:44dollars. I'll tell you that much. And
  275. 11:46you know, we see everything with sports
  276. 11:48books and casinos online nowadays. It's
  277. 11:51vile, but it's just the case.
  278. 11:53All right, let's take a look at this
  279. 11:54simulation here. On the left, you can
  280. 11:58see player wealth paths, negative edge,
  281. 12:01over enough plays, every single player,
  282. 12:05even with variance, right, is going to
  283. 12:08eventually lose all their money. That's
  284. 12:11statistically certain. There's no
  285. 12:13uncertainty there. And on the right,
  286. 12:15damn, it certainly pays to be the
  287. 12:17casino. You can see here that the casino
  288. 12:19literally takes all of the expected
  289. 12:22value, literally takes all of the
  290. 12:24wealth. It's a zero- sum game. If you
  291. 12:26sum up all of the wealth gained and
  292. 12:28lost, right, it would add to zero. But
  293. 12:30wow, take a look at that. The casino
  294. 12:32just takes every single dollar from all
  295. 12:34of these players. I don't want to talk
  296. 12:36about arrogodicity. Not going to talk
  297. 12:37about the markets in the context of
  298. 12:39arrogodicity. The Kelly criterion. This
  299. 12:41is arithmetic. Sure, geometric is
  300. 12:44exactly how we would compound wealth in
  301. 12:47practice. I don't want to hear
  302. 12:48it right now. We're just going to go
  303. 12:50with this simulation to illustrate the
  304. 12:52notion of an edge. All right.
  305. 12:55In practice, market-making desks are
  306. 12:58doing the exact same thing. Okay, their
  307. 13:01edge comes from then the bid ask spread.
  308. 13:05So, what do we do? Well, we're going to
  309. 13:08quote the fair price of that [snorts]
  310. 13:10particular product, security, whatever.
  311. 13:14What is that fair price? Oh, yeah. It's
  312. 13:16based on that hedge portfolio argument.
  313. 13:20Okay. If we understand that hedge
  314. 13:21portfolio argument and how to get to net
  315. 13:24zero expected P&L, then how do we get to
  316. 13:28positive P&L? Well, if we know what the
  317. 13:30fair price is, doesn't it make it easy
  318. 13:33to buy low and sell high all the time?
  319. 13:35Buy at the bid, sell at the ask, that's
  320. 13:38exactly what we're doing. And then all
  321. 13:40we have to do is hedge and try to
  322. 13:43collect the spread. So what you're going
  323. 13:44to see here is of course in in reality,
  324. 13:46you carry a lot of risks. You have model
  325. 13:48risk, sequence of returns risk,
  326. 13:50inventory risk, whatever. There are all
  327. 13:51of these risks that you're going to face
  328. 13:53in the real world, but we're not going
  329. 13:54to talk about that. We're just going to
  330. 13:56talk about the accumulated desk edge.
  331. 13:58So, take a look. This is the exact same
  332. 13:59animation that you saw above, but now
  333. 14:01we're transacting at the bid and ask.
  334. 14:03So, every single time we realize the
  335. 14:06outcome of a trade, check it out. Our
  336. 14:08P&L in expectation is not net zero
  337. 14:11anymore. There is an edge.
  338. 14:15Well, That's pretty good. This is
  339. 14:18the entire framework for making a
  340. 14:20market. Literally, this is everything.
  341. 14:22You have the riskneutral pricing
  342. 14:24argument that constructs the theoretical
  343. 14:26price, the mid price. And then you have
  344. 14:28this broader notion of a statistical
  345. 14:30edge where if theoretically you're
  346. 14:31always going to be buying low or selling
  347. 14:34high, then so long as you hedge
  348. 14:36effectively, you'll be able to what?
  349. 14:39Accumulate that edge over time. Just
  350. 14:42like a casino steals wealth from all of
  351. 14:44the players, you're going to be able to
  352. 14:47accumulate that wealth by providing
  353. 14:48liquidity. But there's a problem.
  354. 14:52There's a there's a lot of problems
  355. 14:53actually. And the probably biggest two
  356. 14:58or at least the the biggest two
  357. 15:00elephants in the room are this idea of
  358. 15:04discretization and model risk. Okay.
  359. 15:08Number one, what if our model doesn't
  360. 15:10capture forward dynamics? Well, I hear
  361. 15:12people moan, and cry all the
  362. 15:14time. Oh, the real world doesn't
  363. 15:17follow a geometric brownie emotion. Oh,
  364. 15:20H model doesn't actually capture the
  365. 15:23dynamics of volatility, which are
  366. 15:25rougher than blah blah blah blah blah
  367. 15:26blah blah. It's like, yeah, join the
  368. 15:29party, dude. Join the party. All
  369. 15:32right. If we had true dynamics, we
  370. 15:35wouldn't be building models in the first
  371. 15:37place. So if our model doesn't capture
  372. 15:40forward dynamics well, then we need a
  373. 15:43better model. All right? And that's what
  374. 15:45we're always trying to do. And then
  375. 15:47there's this battle between efficiency
  376. 15:50and of course the effectiveness. And
  377. 15:52then it turns into this problem of
  378. 15:54computational tractability and and the
  379. 15:56list goes on and on and on. But you
  380. 15:58know, I talk about this all the time in
  381. 16:00the context of of path signatures, in
  382. 16:02the context of of rough volatility
  383. 16:04non-marovian models. I have videos on
  384. 16:06those if you want to check those out.
  385. 16:08All right, so that's probably the
  386. 16:10biggest elephant in the room. And that
  387. 16:12really goes for anything, not just
  388. 16:13sellside trading, but buyside trading,
  389. 16:15too. It's like everyone is just subject
  390. 16:17to model risk. It's like what risks are
  391. 16:20you willing to assume to try to generate
  392. 16:23a return? And you know, the the retailer
  393. 16:26is going to focus on overfitting, but in
  394. 16:28reality, everyone's just stepping up to
  395. 16:30the plate trying to assume model risk to
  396. 16:32the best of their ability. It's really
  397. 16:34not as sexy as everyone makes it out to
  398. 16:37be. Um, okay. So, that's the biggest
  399. 16:39elephant in the room. Number two, we
  400. 16:41can't continuously hedge or we'd
  401. 16:43accumulate infinite frictions
  402. 16:45transaction costs. right now because we
  403. 16:49can't continuously hedge it makes
  404. 16:51realizing the actual P&L in terms of
  405. 16:56the delta hedge close to zero a bit more
  406. 17:01difficult because we are going to what
  407. 17:03accumulate transaction costs and
  408. 17:05increase variance with deterministic or
  409. 17:09not deterministic but discrete hedging.
  410. 17:12Okay, I'm going to play this animation
  411. 17:14for you. Here on the left is the market
  412. 17:17making desk like we saw before with the
  413. 17:19idea of continuous hedging to some
  414. 17:21capacity. And on the right, we have
  415. 17:23rehedging every 02 across this one-year
  416. 17:27simulation. And take a look, our
  417. 17:31expectation
  418. 17:32is less and we have a greater variance
  419. 17:35because of the error in our hedging.
  420. 17:40Okay,
  421. 17:42this chart on the right is much more
  422. 17:44indicative of what we're going to have
  423. 17:45to face in reality because we can't
  424. 17:46continuously hedge. We're always going
  425. 17:48to have some error and we're always
  426. 17:50going to have to accumulate frictions
  427. 17:51like transaction costs. Okay, when I was
  428. 17:55a student a very long time ago, I asked
  429. 17:58my professor,
  430. 18:00I understand
  431. 18:02the argument, the riskneutral pricing
  432. 18:05argument from a delta hydrogen
  433. 18:06perspective, but what the the hell are
  434. 18:08we going to do? What are we gonna do?
  435. 18:10When is it best to delta hedge? Not just
  436. 18:14delta hedge. If you're dynamically
  437. 18:16hedging other Greeks, too. That's, you
  438. 18:18know, that's what I'm talking about
  439. 18:20here. It's it's all synonymous with this
  440. 18:22broader idea of hedging. Okay? So, don't
  441. 18:24just say it's about all only about
  442. 18:25Delta. It's like you have other Greeks,
  443. 18:26too. I understand that. Um, and then in
  444. 18:28practice, they're sticky, shadow Greeks,
  445. 18:30whatever. You know, we have this idea of
  446. 18:32um gamma bands, no trade regions,
  447. 18:35whatever. Um but but that's like the
  448. 18:37more practical deskoriented approach you
  449. 18:40know the my my professor when I was
  450. 18:41asking him he's like yeah
  451. 18:45that that was what I got from uh as a
  452. 18:48response to that question like yeah so I
  453. 18:50did a lot of research at the time um and
  454. 18:52I was looking into a lot of generative
  455. 18:54methods um and I I really appreciate
  456. 18:57this method from Hans Buer which we're
  457. 18:59we're building toward now so hopefully
  458. 19:01you see why we need better prescriptions
  459. 19:06for hedging. We can't hedge continuously
  460. 19:09which would be ideal because then it
  461. 19:11would be much easier to collect that
  462. 19:13theoretical spread. And as soon as we
  463. 19:16induce discrete time and the idea of
  464. 19:19discrete hedging, it's like here I'm
  465. 19:21just doing it at fixed time increments.
  466. 19:24But there's no way that's optimal. What
  467. 19:26if you know we have a massive gap down?
  468. 19:29Is it worth rehedging there? All right.
  469. 19:32What if we have a massive gap down and
  470. 19:33then a massive gap up? We didn't need to
  471. 19:35rehedge there, but we didn't know that
  472. 19:37was going to happen. Yeah. Blah blah
  473. 19:39blah. The list goes on. So, a better
  474. 19:41question to ask, and I already alluded
  475. 19:43to a bunch of practical approaches in
  476. 19:44practice about, you know, just guide
  477. 19:46rails for keeping the hedge on track.
  478. 19:49But like what is the best prescription
  479. 19:51for delta hedging, right? When new
  480. 19:53information when new information
  481. 19:55disseminates, do we want to include that
  482. 19:57into our state consideration for
  483. 19:59rehedging? Are we looking at it purely
  484. 20:01from a riskneutral pricing perspective
  485. 20:03and we're just looking at statistical
  486. 20:05dynamics in terms of when to rehedge?
  487. 20:08This is the entire discussion. This is
  488. 20:11the entire discussion. But at the end of
  489. 20:13the day, whatever you choose to do, you
  490. 20:16are going to be assuming model risk. You
  491. 20:18cannot get away from that. We're dealing
  492. 20:20with the real world. Okay? So even if
  493. 20:23you say, "Oh, like you know, I'm going
  494. 20:26to go with an artificial intelligence
  495. 20:28agent to automatically delta hedge for
  496. 20:31me or just automatically hedge for me."
  497. 20:33It's like, that's great. And on average,
  498. 20:36you may outperform,
  499. 20:38but remember, we're stepping up to the
  500. 20:40plate. So it's very possible, purely
  501. 20:44incidentally, in terms of one
  502. 20:45realization that discrete time hedging
  503. 20:49just generates more P&L.
  504. 20:52All right, it's very possible. But
  505. 20:55remember, P&L isn't really everything at
  506. 20:58the end of the day. It's also about your
  507. 21:00tail risk. It's also about, you know,
  508. 21:03how much risk you're assuming in terms
  509. 21:04of inventory, all these other states,
  510. 21:07all these other um state variable
  511. 21:08informations.
  512. 21:10You're not just going to be training an
  513. 21:12artificial intelligence agent in a marov
  514. 21:15decision process context. And we're
  515. 21:16going to talk about reinforcement
  516. 21:17learning in a second to maximize your
  517. 21:19expected P&L. That would be like
  518. 21:22psychotic, you know, because you you
  519. 21:24need to constrain it by the risk that
  520. 21:26your portfolio assumes. Otherwise, we're
  521. 21:30just going to all triple lever up on on
  522. 21:33Dogecoin, on Pepecoin, on on
  523. 21:35memecoins because it has the highest
  524. 21:38expectation. Like, no. Absolutely not.
  525. 21:40Right? So, that that's really what I'm
  526. 21:42trying to articulate to you here. Okay?
  527. 21:44So we need to learn right we need to
  528. 21:47figure out or learn a framework
  529. 21:51for hedging when is it most optimal to
  530. 21:55hedge right we want to accumulate as few
  531. 21:58frictions as possible transaction costs
  532. 22:00so on and so forth what is the best way
  533. 22:03to do that well this is the mathematical
  534. 22:05function of learning effectively and you
  535. 22:09know this isn't really an indepth video
  536. 22:12on marov decision processes
  537. 22:14SARSA Q-learning, Bellman equations and
  538. 22:17deterministic environments or quasi
  539. 22:19deterministic environments with
  540. 22:20expectations. I'm not going to talk
  541. 22:22about game trees game theory um markoff
  542. 22:25chain Monte Carlo Monte Carlo research.
  543. 22:27If there's interested in if there's
  544. 22:29interest in those subjects certainly let
  545. 22:31me know in the comment section below.
  546. 22:32I'm happy to do videos on those topics.
  547. 22:34I love them. I don't know how relevant
  548. 22:36they are to other topics in quantitative
  549. 22:39finance but they are foundational for
  550. 22:42this particular idea. the broader notion
  551. 22:44of reinforcement learning. So, I
  552. 22:46mentioned all of them for you. If you'd
  553. 22:48like to conduct, you know, research on
  554. 22:50your own, I'm happy to do videos on
  555. 22:52them. Let me know. But this is what it
  556. 22:54builds toward. All right. The idea here
  557. 22:58is we have a series of rewards based on
  558. 23:01states and actions. And our objective is
  559. 23:04going to be to what? Maximize that
  560. 23:07reward. And the reward in this context
  561. 23:10is not just P&L but the reward is risk
  562. 23:14adjusted P&L typically some form of
  563. 23:19conditional value at risk even. Okay. So
  564. 23:23every single state is going to generate
  565. 23:25an action that is dictated by a policy
  566. 23:29function and then it's iterative. That
  567. 23:31action is going to generate a new state
  568. 23:34and that state is going to be plugged
  569. 23:35right back into the policy function.
  570. 23:37It's going to generate another action.
  571. 23:39And of course, our objective is going to
  572. 23:41be because we're dealing with a
  573. 23:42stochastic environment to maximize the
  574. 23:45expectation under that policy. And once
  575. 23:48we have that policy function, we're
  576. 23:49going to be able to derive the policy or
  577. 23:51the action to take given a particular
  578. 23:54state. What does this sound like? It
  579. 23:57should sound like how you've learned and
  580. 23:59acted your entire life. You are
  581. 24:02presented a state and then you choose an
  582. 24:05action. And then you're presented with
  583. 24:07another state and you choose another
  584. 24:10action. You are literally a composite
  585. 24:13policy function. But nobody is feeding
  586. 24:17you a discrete state vector. Your
  587. 24:20biology, your your senses, everything
  588. 24:24you're experiencing
  589. 24:26is more or less a continuous state
  590. 24:30vector. So that is starting to introduce
  591. 24:33some frictions between biological agents
  592. 24:37and artificial intelligence agents. This
  593. 24:40is a broader discussion and there's a
  594. 24:42lot of interesting things in physics
  595. 24:44that I'm working on right now in terms
  596. 24:46of constraints and bounds but it's
  597. 24:49really interesting because biological
  598. 24:51agents like you and I don't have this
  599. 24:54computational intractability of oh we
  600. 24:57need to feed it a deterministic series
  601. 25:00of states in a in a discrete capacity.
  602. 25:04um that's that's nice and those agents
  603. 25:07might be able to extract more
  604. 25:08information and generalize faster in
  605. 25:10some contexts, but it also means that
  606. 25:12there are things it can't do because of
  607. 25:15our sensory capacity. So, a very
  608. 25:17interesting segue for you there, but
  609. 25:19what I have for you here is a beautiful
  610. 25:22animation
  611. 25:24of a Markov decision process, this
  612. 25:26broader idea of reinforcement learning.
  613. 25:28Okay, so on the left here, you're going
  614. 25:31to see the environment. I have a star of
  615. 25:34+ 10 and I have this red square of minus
  616. 25:3810. Okay. Now, the agent is trying to
  617. 25:41figure out this policy function as it
  618. 25:43engages with its environment. It's going
  619. 25:45to walk around and it's going to touch
  620. 25:48the star. It's going to touch the skull.
  621. 25:52And over time, it's going to hopefully
  622. 25:54learn the correct series of actions. In
  623. 25:58fact, you can see right here, this is
  624. 26:00the narrow network that is taking the
  625. 26:03state features as an input and it is
  626. 26:06trying to learn the appropriate output.
  627. 26:08And what have you noticed here? The
  628. 26:10agent is learning, oh we do not
  629. 26:13want to touch this red skull. Whenever
  630. 26:15we start to go this way or if we ever
  631. 26:17spawn on this square, we want to go away
  632. 26:20from that red skull and toward the green
  633. 26:23star. This is exactly how the learning
  634. 26:27process occurs in practice. The only
  635. 26:30difference in practice is what we always
  636. 26:34have this problem as biological agents
  637. 26:36of signal and noise. Right? So when you
  638. 26:40make a decision, you're never acting in
  639. 26:43against your own self-interests, right?
  640. 26:46But when you make that decision, you may
  641. 26:48yield a negative outcome. So good
  642. 26:52decision, bad outcome. But was it a good
  643. 26:54decision? That is the hardest part to
  644. 26:58analyze. This gets into poker theory. I
  645. 27:00recently read um thinking in bets. It's
  646. 27:02a great book and it talks about this
  647. 27:04idea of good decisions can yield bad
  648. 27:08outcomes. Bad decisions can yield good
  649. 27:11outcomes in a random environment. So
  650. 27:15that's why learning in some cases in
  651. 27:18reality as biological agents you and I
  652. 27:21is very challenging when you play a hand
  653. 27:24of poker. Did you make a good decision
  654. 27:27or were you lucky or were you unlucky?
  655. 27:30Did you make a bad decision and you were
  656. 27:32lucky? This is exactly what you need to
  657. 27:34discern. And it is not trivial. And if
  658. 27:37anyone says it's trivial, ask them why
  659. 27:40they're not a multi-time World Poker
  660. 27:42Tour champion. Okay? Because they should
  661. 27:45have no trouble distinguishing signal
  662. 27:48and noise. In reality, it's not a
  663. 27:50trivial problem. We all face it every
  664. 27:52single day. But what you're observing
  665. 27:54here is the idea of a Markov decision
  666. 27:56process, reinforcement learning in this
  667. 27:58environment. Look at this point. All
  668. 28:01right, by episode 90 almost 100, the
  669. 28:05agent has no interest in going towards
  670. 28:07this skull. It has no interest anymore.
  671. 28:10It is simply trying to go toward the
  672. 28:13star. All right. It is learning this
  673. 28:16through experience. The weights in its
  674. 28:18neural network are adjusting. It's
  675. 28:20learning this expected reward based on a
  676. 28:22series of its actions and states. And it
  677. 28:26now knows what to do.
  678. 28:29Okay. So that is this broader notion of
  679. 28:32reinforcement learning. After a series
  680. 28:33of iterations, right, it is now able to
  681. 28:36act on its own. It has the policy
  682. 28:38function and that policy function is
  683. 28:41going to dictate optimal actions for
  684. 28:43wherever it exists on this board in this
  685. 28:46environment as a function of the current
  686. 28:49state. Okay,
  687. 28:54understanding that let's go back to this
  688. 28:56idea now of
  689. 29:00deep hedging. We're going to apply this
  690. 29:02framework of reinforcement learning to
  691. 29:05the market making desk. We are going to
  692. 29:09be trying to train an agent now to hedge
  693. 29:14optimally. And this is the idea of deep
  694. 29:16hedging Hans Beller. All right. Instead
  695. 29:19of arbitrary frameworks, no trade
  696. 29:22regions, gamma bands, deterministic
  697. 29:24rehagging, whatever, the first thing
  698. 29:27that we're going to do is observe. Okay,
  699. 29:30there are more optimal ways in terms of
  700. 29:33risk that the desk is assuming in terms
  701. 29:35of optimizing the P&L at the end of the
  702. 29:38day to hedge and adjust our hedges over
  703. 29:43time.
  704. 29:45Number one, we are always going to have
  705. 29:47to assume model risk. It could be a
  706. 29:50classical model like stocastic
  707. 29:51volatility, jump diffusion. It could be
  708. 29:53data driven models like VAEEs, GANs,
  709. 29:55other generative models. I've used
  710. 29:57generative models in that capacity
  711. 29:59before to simulate tail risk events that
  712. 30:02may have not occurred yet empirically
  713. 30:05just as one example for you. But no
  714. 30:08matter what, even in this context, as
  715. 30:10we're training an artificial
  716. 30:12intelligence, a computational agent to
  717. 30:16make these hedging decisions for us, we
  718. 30:18have to assume a model. We have to give
  719. 30:20it information to learn from. And it's
  720. 30:23possible that we haven't experienced
  721. 30:25something that the agent has seen
  722. 30:28before. Just like us in the real world.
  723. 30:30All right, we've never experienced
  724. 30:32computational agents before. We could
  725. 30:34train on all the historic data we would
  726. 30:36like, right? Nobody knew that was going
  727. 30:38to happen. Then LLMs came and changed
  728. 30:42the world. And you can also have left
  729. 30:44tail world changing events that never
  730. 30:46occurred before either. So that's called
  731. 30:49model risk. There's no way around that.
  732. 30:52We have a data scarcity issue. These
  733. 30:54reinforcement learning agents are
  734. 30:55extremely difficult to train. They take
  735. 30:58a lot of data. They are very finicky to
  736. 31:01learn the policy function you saw in
  737. 31:03that little grid world example. It was
  738. 31:05reasonably easy. All right. But it still
  739. 31:08took a hundred episodes. 100. Okay. We
  740. 31:13can see right off the bat what to do.
  741. 31:16But it took a 100 episodes for that
  742. 31:19agent to learn what to do. So I want you
  743. 31:22to understand that in the context of
  744. 31:23learning for deep hedging, it is a
  745. 31:27tremendous tremendous issue from a
  746. 31:30computational tractability perspective.
  747. 31:33All right. So once we assume a model, we
  748. 31:35can select whatever one we would like in
  749. 31:36terms of a market simulator. We train a
  750. 31:39neural network agent. Okay. Does this
  751. 31:43look familiar? is exactly what we just
  752. 31:44talked about before in the grid world
  753. 31:46example. All right, the only difference
  754. 31:48is our state is not just the
  755. 31:49environment, right? The map effectively.
  756. 31:52It's inventory, current positions, the
  757. 31:55underlying assets, the current prices,
  758. 31:57relevant market observations,
  759. 31:59um path dependent features, realized
  760. 32:02volatility, running minax, time elapse,
  761. 32:04the list goes on. Whatever you think
  762. 32:06might be effective in explaining
  763. 32:08variation in the cross-section here. And
  764. 32:11again, this is for multiasset
  765. 32:12portfolios. you'll have implied
  766. 32:13volatility surfaces so on and so forth.
  767. 32:18The neural network is trying to
  768. 32:20approximate this map. Okay,
  769. 32:24neural networks and this is why it's
  770. 32:26called deep hedging are universal
  771. 32:27function approximators. The function
  772. 32:29we're approximating is what it is the
  773. 32:34policy function. It is an action given a
  774. 32:36current state and that state is just a
  775. 32:39vector. So everything is numerical.
  776. 32:41Everything is mathematical. If you have
  777. 32:42a news article that you want to feed to
  778. 32:44the neural network to approximate the
  779. 32:46policy, then you have to tokenize it
  780. 32:48first, interpret it and feed it into the
  781. 32:52state vector in a reasonable way. Okay,
  782. 32:55so that is literally how it functions.
  783. 32:56Then what do we do? Well, we have this
  784. 32:58reward function and then that reward
  785. 33:01function is what just expected P&L? No,
  786. 33:04we have entropic risk measures. We have
  787. 33:06expected shortfall. We have these risk
  788. 33:08measures that we are going to try to
  789. 33:10optimize for in the reinforcement
  790. 33:14learning context. And now here I can
  791. 33:17show you this beautiful dashboard. Yeah,
  792. 33:20I'm definitely not open sourcing this.
  793. 33:22This thing's awesome. Um, this
  794. 33:24is what it looks like to learn to hedge.
  795. 33:28I want you to see here that on the left
  796. 33:30we're observing one stochcastic path
  797. 33:32from the market. This is literally the
  798. 33:34exact same situation as we had before.
  799. 33:37Okay, it's the exact same situation.
  800. 33:39Meaning instead of a grid world, we're
  801. 33:41now observing a market and a market
  802. 33:43simulator. So we have this one
  803. 33:45stochcastic path, the underlying asset,
  804. 33:47we have whatever position we're taking
  805. 33:49in a European option contract. And here
  806. 33:52we have the hedge. I have the deep hedge
  807. 33:55that is the hedge that the neural
  808. 33:56network is learning and the classical
  809. 33:59delta hedge according to a rule set
  810. 34:01prior. So that delta hedge would be what
  811. 34:04you would need to hold to perfectly
  812. 34:06hedge the exposure. We're learning then
  813. 34:10a more optimal way to hedge as we
  814. 34:13continue to simulate more and more and
  815. 34:14more and more of these trades. And what
  816. 34:17you're going to observe is the tail risk
  817. 34:20changing relative to the arbitrary
  818. 34:23static delta hedge. Eg i.e. We are
  819. 34:28learning a more effective way to hedge
  820. 34:31the portfolio. And that's exactly what
  821. 34:34you are seeing here. You are seeing the
  822. 34:38tail risk during the training process
  823. 34:41fundamentally change relative to the
  824. 34:43fixed distribution of static delta
  825. 34:46hedging. This is the idea of deep
  826. 34:50hedging. In other words, it is more
  827. 34:54effective on average in the model sense
  828. 34:57to use an agent like this to go about
  829. 35:02constructing your hedges and rehedging a
  830. 35:05portfolio, a multiasset portfolio, a
  831. 35:07series of different exposures that you
  832. 35:10are going to be running in the market
  833. 35:13making sense. That's going to do it for
  834. 35:15this video on deep hedging for
  835. 35:18quantitative finance. I hope you
  836. 35:19enjoyed. But I hope you learned
  837. 35:21something. This video certainly took a
  838. 35:22tremendous effort to put together. So if
  839. 35:25you liked it and you want to see more
  840. 35:26like it in the future, please like,
  841. 35:28comment, subscribe, share. It helps me
  842. 35:30out tremendously. It is always greatly
  843. 35:32appreciated. Other than that, I'm going
  844. 35:35to thank you so much for watching and I
  845. 35:38will see you in the next

About this transcript

This page contains the full transcript of Deep Hedging for Quantitative Finance by Roman Paolucci, generated from the public captions YouTube serves with the video. The transcript has 5,691 words across 845 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.

What you can do with it

Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.

Free YouTube transcript tool

YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.