Deep Hedging for Quantitative Finance — Transcript
Full transcript
- 0:00Roman, if we have all of this data, why
- 0:02aren't we just training artificial
- 0:04intelligence, computational agents to go
- 0:07about making trading decisions for us?
- 0:10We can give it a state vector, a vector
- 0:12of information containing news, spot
- 0:15prices, compressions of information that
- 0:17have already happened, path signatures,
- 0:20volatility, the list goes on. And that
- 0:23computational agent is going to make way
- 0:26better decisions than us. Why aren't we
- 0:30why aren't we doing that? Well, we're
- 0:33working on it. It is the frontier of the
- 0:35academic literature and that's exactly
- 0:38what we're going to talk about today.
- 0:40We're going to talk about this idea of
- 0:42deep hedging for quantitative finance
- 0:44pioneered of course by Hans Buer who I
- 0:48saw at the Bachier conference giving a
- 0:50presentation when I was a young quant
- 0:52and I didn't know my ash for my elbow
- 0:54and I was trying to figure out what the
- 0:56hell he was talking about from a
- 0:58marketmaking perspective and I had never
- 1:00seen the reinforcement learning
- 1:01literature at the time and I was trying
- 1:03to figure out what the hell he was
- 1:04talking about there as well. But
- 1:06fortunately since then after a whole
- 1:09bunch of whiteboarding with my
- 1:11colleagues at the time and a whole bunch
- 1:13of reading and of course implementation
- 1:16I have been able to understand the
- 1:20landscape and construct this lesson for
- 1:23you to save you all of the headaches
- 1:26that I had trying to understand this
- 1:29over the course of several years. Now,
- 1:31of course, the problem that we face is
- 1:34very intricate because reinforcement
- 1:36learning requires a substantial amount
- 1:38of data and iteration. It is a very
- 1:43there's really no other way to put it
- 1:45finicky problem. It's very difficult to
- 1:49approximate these policy functions as
- 1:52we're going to discuss. So even if you
- 1:55have a a state with a whole bunch of
- 1:57information, it's discretized. We exist
- 2:01in the continuous sense, but we can only
- 2:04feed a computational agent so much
- 2:07information over so many different time
- 2:09intervals. And very quickly, that's
- 2:13going to explode in terms of a problem,
- 2:16but we'll get there. Let's start on the
- 2:19trading desk. We're going to talk about
- 2:21this notion of riskneutral pricing. Why
- 2:24should we care at all about riskneutral
- 2:27pricing? Well, you want to make money,
- 2:29don't you? If you're a market maker,
- 2:31your objective is to provide liquidity
- 2:33to the market. We can debate about
- 2:36whether the role of a market maker is
- 2:38productive or not. I really don't care
- 2:40about that right now. I'm just talking
- 2:41about the broader business function. If
- 2:44that market maker is going to quote both
- 2:47sides of a market, right, a two-way,
- 2:50then what should the mid price be? What
- 2:54is the fair value of a derivative
- 2:56contract right now anyway? Well, we
- 3:00don't have a crystal ball. So, the best
- 3:02we can do is what? An expectation.
- 3:05Shocker. That is quite literally
- 3:08mathematically the best guess in the
- 3:10mean squared error sense. But where does
- 3:13that expectation fit into the broader
- 3:15notion of riskneutral pricing? Well,
- 3:18Black Scholes gives us this delta
- 3:20hedging argument in their framework.
- 3:23Assuming the underlying follows a
- 3:25geometric brownian motion, the market is
- 3:27complete. And if we can hedge away that
- 3:30randomness via a delta hedge, then we
- 3:33are going to be able to construct what
- 3:35is known as the value of a replicating
- 3:38portfolio. That is the fair value of the
- 3:41option contract today. There's no
- 3:43speculation here. All right? We're not
- 3:46talking about, oh, you know, the market
- 3:48risk premium is 7 and a half% on average
- 3:51risk adjusted or inflation adjusted
- 3:53year-over-year, so that's what we're
- 3:55going to use in the P measure sense to
- 3:57price our derivatives. No, no, no, no.
- 4:00This is riskneutral pricing. We are
- 4:03hedging that risk in the complete market
- 4:06sense. And on average, if we have no
- 4:09risk, we must earn the risk-free rate.
- 4:13So that's exactly what you see here as
- 4:15the fair price, the mid price of our
- 4:18derivative today. Because if we go out
- 4:21and we buy these contracts and we hedge
- 4:24all of that risk, then this is what we
- 4:27are going to earn on average. It's the
- 4:29riskneutral
- 4:31discounted expected payoff and that's
- 4:34exactly what you see here. Now,
- 4:37typically when I introduce this idea to
- 4:40students, it is a cluster I'm not
- 4:43gonna put it lightly. It's a cluster
- 4:45because they're like, "So, you
- 4:48always earn the risk-free rate?" It's
- 4:50like, it's like, "No, you earn it in
- 4:52expectation on average." So, where do I
- 4:55like to start? I always like to start in
- 4:57the stochcastic world in in the fineman
- 4:59CAC sense and then take that
- 5:03visualization and move it over to the
- 5:06delta hedging argument to show exactly
- 5:09what's going on. And that's exactly what
- 5:11I have for you here today in terms of an
- 5:13animation. So on the right you're going
- 5:15to observe the option price. Okay. Now,
- 5:18what we're going to do to try to figure
- 5:20out the value of the option contract is
- 5:22we're going to simulate the underlying
- 5:24asset and we're going to compute the
- 5:27payoff of that option for this one
- 5:29sample path. Then what we're going to do
- 5:32is effectively add it to a list. And
- 5:34then we're going to do it again and
- 5:36again and again and just take the
- 5:37average of that list. This is Monte
- 5:39Carlos simulating the options price
- 5:42which we know is going to converge to
- 5:43the black trolls price which you see
- 5:45here as a red dash line. That's fineman
- 5:48CAC. So if I play this animation for you
- 5:50here, it's not going to be a shocker. As
- 5:52we realize more and more and more and
- 5:54more price paths, you're going to see
- 5:56the option price converge on the right
- 5:58to the black Scholes price. Again, not a
- 6:01shocker. That's fine CAC. So what you're
- 6:04observing here is essentially an
- 6:07approximation for the fair value of the
- 6:10option price today. But this is not the
- 6:14black schles delta hedging argument.
- 6:16Okay, I need to be very clear here. I
- 6:19introduce Monte Carlo simulating the
- 6:22option price first because it's very
- 6:23easy to understand and mathematically
- 6:26again thanks to Fman and CAC that piece
- 6:29of the puzzle is automatically
- 6:31equivalent to the result from Black
- 6:34Scholes. They have to be equivalent.
- 6:37That partial differential equation
- 6:38representation is equivalent to the
- 6:41conditional expectation in the
- 6:43stochastic world. it has to be and I
- 6:46actually have a dedicated video deriving
- 6:49that result if you would like to check
- 6:52it out. So understanding that this is
- 6:55the approximation for this riskneutral
- 6:57conditional expectation
- 6:59discounted of course for the option
- 7:01price now we can go about validating it
- 7:04in the black scholes framework. All
- 7:08right so what are we going to do in a
- 7:10black trolls framework? Well, what we're
- 7:12going to do is we are going to go out
- 7:14and we are going to buy or sell this
- 7:18contract. All right? And the contract
- 7:21that we are going to be holding then is
- 7:24going to be delta hedged. So on the left
- 7:27here, what I want to show you is I want
- 7:29to show you the underlying asset. So
- 7:31this is the same chart that you saw
- 7:33above. It's the underlying asset
- 7:35simulation. But on the right, I want to
- 7:38show you our equity curve. All right, we
- 7:40go out and we're going to go long or
- 7:43short the option at this theoretical
- 7:46fair price. And then throughout the life
- 7:48of the option, we are going to be
- 7:51continuously delta hedging. What does
- 7:55this actually look like? If I play this
- 7:57animation for you here, I want you to
- 7:59observe on the left, right, we have all
- 8:02of these market paths and you can see
- 8:04that the drift of the market paths on
- 8:07average is going to be roughly what?
- 8:10It's going to be roughly the risk-free
- 8:11rate. All right, that is the dynamics
- 8:14that we're using for the simulation here
- 8:15of the geometric brownie emotion. On the
- 8:18right, you can see as we experience more
- 8:21and more and more and more delta hedged
- 8:24paths, what happens to the average P&L
- 8:27for each trade? It becomes zero. That is
- 8:32exactly what this replicating portfolio
- 8:34argument is suggesting. Functionally and
- 8:38visually, this is exactly what the
- 8:40implication is. Okay. So if we go out
- 8:43and we transact at the fair value of
- 8:46this financial derivative, not fair
- 8:48value out of thin air, fair value based
- 8:51on this framework in the stochcastic
- 8:54sense or this framework in the black
- 8:56shaw sense, then we're going to if we
- 8:59delta hedge throughout the life of that
- 9:02trade realize net zero P&L on average.
- 9:08All right. On average,
- 9:11this makes it very easy to see where
- 9:14we're going and the sense of market
- 9:16making. What are you going to be a
- 9:17liquidity provider if on average this
- 9:20yellow line your empirical expected P&L
- 9:24is going to be zero? No, you're not
- 9:28going to transact at mid. Everyone pulls
- 9:29up Yahoo Finance or Bloomberg
- 9:31Terminal Interactive Brokers and we go,
- 9:34"Oh, you know, Nvidia is $220."
- 9:37um oh you know the spy ETF it's like
- 9:39it's like no that's that's the mid price
- 9:42right that's the mid price last week I
- 9:44did a video on making a market on a onem
- 9:47run to show you exactly the implications
- 9:49of this from a market making perspective
- 9:52and what we're seeing here is if desk
- 9:54P&L is flat on average in expectation
- 9:59nobody's going to provide liquidity so
- 10:01what are you going to do you're going to
- 10:02quote two way you're going to be willing
- 10:04to what buy at the bid
- 10:07sell at the ask. And what that's going
- 10:09to do is it's going to artificially
- 10:11induce
- 10:13an edge. And that edge means that this
- 10:17desk P&L is going to have a positive
- 10:20trajectory. All right.
- 10:23But what's the catch? The catch is this
- 10:27only holds with continuous delta
- 10:29hedging. In practice, you can't delta
- 10:32hedge continuously.
- 10:34You have to hedge what? Indiscreet time.
- 10:38Otherwise, you're going to accumulate an
- 10:39infinite number of transaction costs,
- 10:40right? And that's going to eat away at
- 10:43your profits. Okay. Moreover, you have
- 10:46model risk, inventory risk, blah blah
- 10:48blah, all the other stuff, too. But
- 10:49we'll talk about that down the line. All
- 10:52right. Let's talk about market and the
- 10:54hedging problem. If we understand the
- 10:56idea of risk neutral pricing and why we
- 10:57care about it from a desk perspective,
- 10:59well, if we know what the fair price is,
- 11:01we know we can get to net zero P&L. And
- 11:03then if we have a spread on top of that,
- 11:05we're going to accumulate an edge. But
- 11:06what is an edge anyway, let's quickly go
- 11:09to the casino. All right, this is an
- 11:12example of a fixed edge. All right, in a
- 11:17casino, if we look at something like
- 11:18roulette, the players will always have a
- 11:22negative edge. The casino will always
- 11:25have a positive edge. All right? As long
- 11:29as the casino doesn't reach an absorbing
- 11:32state of zero wealth where they're in
- 11:34ruin and they can't come back from it,
- 11:36they are guaranteed to accumulate
- 11:38infinite wealth. Man, that's the
- 11:42government's wet dream in terms of tax
- 11:44dollars. I'll tell you that much. And
- 11:46you know, we see everything with sports
- 11:48books and casinos online nowadays. It's
- 11:51vile, but it's just the case.
- 11:53All right, let's take a look at this
- 11:54simulation here. On the left, you can
- 11:58see player wealth paths, negative edge,
- 12:01over enough plays, every single player,
- 12:05even with variance, right, is going to
- 12:08eventually lose all their money. That's
- 12:11statistically certain. There's no
- 12:13uncertainty there. And on the right,
- 12:15damn, it certainly pays to be the
- 12:17casino. You can see here that the casino
- 12:19literally takes all of the expected
- 12:22value, literally takes all of the
- 12:24wealth. It's a zero- sum game. If you
- 12:26sum up all of the wealth gained and
- 12:28lost, right, it would add to zero. But
- 12:30wow, take a look at that. The casino
- 12:32just takes every single dollar from all
- 12:34of these players. I don't want to talk
- 12:36about arrogodicity. Not going to talk
- 12:37about the markets in the context of
- 12:39arrogodicity. The Kelly criterion. This
- 12:41is arithmetic. Sure, geometric is
- 12:44exactly how we would compound wealth in
- 12:47practice. I don't want to hear
- 12:48it right now. We're just going to go
- 12:50with this simulation to illustrate the
- 12:52notion of an edge. All right.
- 12:55In practice, market-making desks are
- 12:58doing the exact same thing. Okay, their
- 13:01edge comes from then the bid ask spread.
- 13:05So, what do we do? Well, we're going to
- 13:08quote the fair price of that [snorts]
- 13:10particular product, security, whatever.
- 13:14What is that fair price? Oh, yeah. It's
- 13:16based on that hedge portfolio argument.
- 13:20Okay. If we understand that hedge
- 13:21portfolio argument and how to get to net
- 13:24zero expected P&L, then how do we get to
- 13:28positive P&L? Well, if we know what the
- 13:30fair price is, doesn't it make it easy
- 13:33to buy low and sell high all the time?
- 13:35Buy at the bid, sell at the ask, that's
- 13:38exactly what we're doing. And then all
- 13:40we have to do is hedge and try to
- 13:43collect the spread. So what you're going
- 13:44to see here is of course in in reality,
- 13:46you carry a lot of risks. You have model
- 13:48risk, sequence of returns risk,
- 13:50inventory risk, whatever. There are all
- 13:51of these risks that you're going to face
- 13:53in the real world, but we're not going
- 13:54to talk about that. We're just going to
- 13:56talk about the accumulated desk edge.
- 13:58So, take a look. This is the exact same
- 13:59animation that you saw above, but now
- 14:01we're transacting at the bid and ask.
- 14:03So, every single time we realize the
- 14:06outcome of a trade, check it out. Our
- 14:08P&L in expectation is not net zero
- 14:11anymore. There is an edge.
- 14:15Well, That's pretty good. This is
- 14:18the entire framework for making a
- 14:20market. Literally, this is everything.
- 14:22You have the riskneutral pricing
- 14:24argument that constructs the theoretical
- 14:26price, the mid price. And then you have
- 14:28this broader notion of a statistical
- 14:30edge where if theoretically you're
- 14:31always going to be buying low or selling
- 14:34high, then so long as you hedge
- 14:36effectively, you'll be able to what?
- 14:39Accumulate that edge over time. Just
- 14:42like a casino steals wealth from all of
- 14:44the players, you're going to be able to
- 14:47accumulate that wealth by providing
- 14:48liquidity. But there's a problem.
- 14:52There's a there's a lot of problems
- 14:53actually. And the probably biggest two
- 14:58or at least the the biggest two
- 15:00elephants in the room are this idea of
- 15:04discretization and model risk. Okay.
- 15:08Number one, what if our model doesn't
- 15:10capture forward dynamics? Well, I hear
- 15:12people moan, and cry all the
- 15:14time. Oh, the real world doesn't
- 15:17follow a geometric brownie emotion. Oh,
- 15:20H model doesn't actually capture the
- 15:23dynamics of volatility, which are
- 15:25rougher than blah blah blah blah blah
- 15:26blah blah. It's like, yeah, join the
- 15:29party, dude. Join the party. All
- 15:32right. If we had true dynamics, we
- 15:35wouldn't be building models in the first
- 15:37place. So if our model doesn't capture
- 15:40forward dynamics well, then we need a
- 15:43better model. All right? And that's what
- 15:45we're always trying to do. And then
- 15:47there's this battle between efficiency
- 15:50and of course the effectiveness. And
- 15:52then it turns into this problem of
- 15:54computational tractability and and the
- 15:56list goes on and on and on. But you
- 15:58know, I talk about this all the time in
- 16:00the context of of path signatures, in
- 16:02the context of of rough volatility
- 16:04non-marovian models. I have videos on
- 16:06those if you want to check those out.
- 16:08All right, so that's probably the
- 16:10biggest elephant in the room. And that
- 16:12really goes for anything, not just
- 16:13sellside trading, but buyside trading,
- 16:15too. It's like everyone is just subject
- 16:17to model risk. It's like what risks are
- 16:20you willing to assume to try to generate
- 16:23a return? And you know, the the retailer
- 16:26is going to focus on overfitting, but in
- 16:28reality, everyone's just stepping up to
- 16:30the plate trying to assume model risk to
- 16:32the best of their ability. It's really
- 16:34not as sexy as everyone makes it out to
- 16:37be. Um, okay. So, that's the biggest
- 16:39elephant in the room. Number two, we
- 16:41can't continuously hedge or we'd
- 16:43accumulate infinite frictions
- 16:45transaction costs. right now because we
- 16:49can't continuously hedge it makes
- 16:51realizing the actual P&L in terms of
- 16:56the delta hedge close to zero a bit more
- 17:01difficult because we are going to what
- 17:03accumulate transaction costs and
- 17:05increase variance with deterministic or
- 17:09not deterministic but discrete hedging.
- 17:12Okay, I'm going to play this animation
- 17:14for you. Here on the left is the market
- 17:17making desk like we saw before with the
- 17:19idea of continuous hedging to some
- 17:21capacity. And on the right, we have
- 17:23rehedging every 02 across this one-year
- 17:27simulation. And take a look, our
- 17:31expectation
- 17:32is less and we have a greater variance
- 17:35because of the error in our hedging.
- 17:40Okay,
- 17:42this chart on the right is much more
- 17:44indicative of what we're going to have
- 17:45to face in reality because we can't
- 17:46continuously hedge. We're always going
- 17:48to have some error and we're always
- 17:50going to have to accumulate frictions
- 17:51like transaction costs. Okay, when I was
- 17:55a student a very long time ago, I asked
- 17:58my professor,
- 18:00I understand
- 18:02the argument, the riskneutral pricing
- 18:05argument from a delta hydrogen
- 18:06perspective, but what the the hell are
- 18:08we going to do? What are we gonna do?
- 18:10When is it best to delta hedge? Not just
- 18:14delta hedge. If you're dynamically
- 18:16hedging other Greeks, too. That's, you
- 18:18know, that's what I'm talking about
- 18:20here. It's it's all synonymous with this
- 18:22broader idea of hedging. Okay? So, don't
- 18:24just say it's about all only about
- 18:25Delta. It's like you have other Greeks,
- 18:26too. I understand that. Um, and then in
- 18:28practice, they're sticky, shadow Greeks,
- 18:30whatever. You know, we have this idea of
- 18:32um gamma bands, no trade regions,
- 18:35whatever. Um but but that's like the
- 18:37more practical deskoriented approach you
- 18:40know the my my professor when I was
- 18:41asking him he's like yeah
- 18:45that that was what I got from uh as a
- 18:48response to that question like yeah so I
- 18:50did a lot of research at the time um and
- 18:52I was looking into a lot of generative
- 18:54methods um and I I really appreciate
- 18:57this method from Hans Buer which we're
- 18:59we're building toward now so hopefully
- 19:01you see why we need better prescriptions
- 19:06for hedging. We can't hedge continuously
- 19:09which would be ideal because then it
- 19:11would be much easier to collect that
- 19:13theoretical spread. And as soon as we
- 19:16induce discrete time and the idea of
- 19:19discrete hedging, it's like here I'm
- 19:21just doing it at fixed time increments.
- 19:24But there's no way that's optimal. What
- 19:26if you know we have a massive gap down?
- 19:29Is it worth rehedging there? All right.
- 19:32What if we have a massive gap down and
- 19:33then a massive gap up? We didn't need to
- 19:35rehedge there, but we didn't know that
- 19:37was going to happen. Yeah. Blah blah
- 19:39blah. The list goes on. So, a better
- 19:41question to ask, and I already alluded
- 19:43to a bunch of practical approaches in
- 19:44practice about, you know, just guide
- 19:46rails for keeping the hedge on track.
- 19:49But like what is the best prescription
- 19:51for delta hedging, right? When new
- 19:53information when new information
- 19:55disseminates, do we want to include that
- 19:57into our state consideration for
- 19:59rehedging? Are we looking at it purely
- 20:01from a riskneutral pricing perspective
- 20:03and we're just looking at statistical
- 20:05dynamics in terms of when to rehedge?
- 20:08This is the entire discussion. This is
- 20:11the entire discussion. But at the end of
- 20:13the day, whatever you choose to do, you
- 20:16are going to be assuming model risk. You
- 20:18cannot get away from that. We're dealing
- 20:20with the real world. Okay? So even if
- 20:23you say, "Oh, like you know, I'm going
- 20:26to go with an artificial intelligence
- 20:28agent to automatically delta hedge for
- 20:31me or just automatically hedge for me."
- 20:33It's like, that's great. And on average,
- 20:36you may outperform,
- 20:38but remember, we're stepping up to the
- 20:40plate. So it's very possible, purely
- 20:44incidentally, in terms of one
- 20:45realization that discrete time hedging
- 20:49just generates more P&L.
- 20:52All right, it's very possible. But
- 20:55remember, P&L isn't really everything at
- 20:58the end of the day. It's also about your
- 21:00tail risk. It's also about, you know,
- 21:03how much risk you're assuming in terms
- 21:04of inventory, all these other states,
- 21:07all these other um state variable
- 21:08informations.
- 21:10You're not just going to be training an
- 21:12artificial intelligence agent in a marov
- 21:15decision process context. And we're
- 21:16going to talk about reinforcement
- 21:17learning in a second to maximize your
- 21:19expected P&L. That would be like
- 21:22psychotic, you know, because you you
- 21:24need to constrain it by the risk that
- 21:26your portfolio assumes. Otherwise, we're
- 21:30just going to all triple lever up on on
- 21:33Dogecoin, on Pepecoin, on on
- 21:35memecoins because it has the highest
- 21:38expectation. Like, no. Absolutely not.
- 21:40Right? So, that that's really what I'm
- 21:42trying to articulate to you here. Okay?
- 21:44So we need to learn right we need to
- 21:47figure out or learn a framework
- 21:51for hedging when is it most optimal to
- 21:55hedge right we want to accumulate as few
- 21:58frictions as possible transaction costs
- 22:00so on and so forth what is the best way
- 22:03to do that well this is the mathematical
- 22:05function of learning effectively and you
- 22:09know this isn't really an indepth video
- 22:12on marov decision processes
- 22:14SARSA Q-learning, Bellman equations and
- 22:17deterministic environments or quasi
- 22:19deterministic environments with
- 22:20expectations. I'm not going to talk
- 22:22about game trees game theory um markoff
- 22:25chain Monte Carlo Monte Carlo research.
- 22:27If there's interested in if there's
- 22:29interest in those subjects certainly let
- 22:31me know in the comment section below.
- 22:32I'm happy to do videos on those topics.
- 22:34I love them. I don't know how relevant
- 22:36they are to other topics in quantitative
- 22:39finance but they are foundational for
- 22:42this particular idea. the broader notion
- 22:44of reinforcement learning. So, I
- 22:46mentioned all of them for you. If you'd
- 22:48like to conduct, you know, research on
- 22:50your own, I'm happy to do videos on
- 22:52them. Let me know. But this is what it
- 22:54builds toward. All right. The idea here
- 22:58is we have a series of rewards based on
- 23:01states and actions. And our objective is
- 23:04going to be to what? Maximize that
- 23:07reward. And the reward in this context
- 23:10is not just P&L but the reward is risk
- 23:14adjusted P&L typically some form of
- 23:19conditional value at risk even. Okay. So
- 23:23every single state is going to generate
- 23:25an action that is dictated by a policy
- 23:29function and then it's iterative. That
- 23:31action is going to generate a new state
- 23:34and that state is going to be plugged
- 23:35right back into the policy function.
- 23:37It's going to generate another action.
- 23:39And of course, our objective is going to
- 23:41be because we're dealing with a
- 23:42stochastic environment to maximize the
- 23:45expectation under that policy. And once
- 23:48we have that policy function, we're
- 23:49going to be able to derive the policy or
- 23:51the action to take given a particular
- 23:54state. What does this sound like? It
- 23:57should sound like how you've learned and
- 23:59acted your entire life. You are
- 24:02presented a state and then you choose an
- 24:05action. And then you're presented with
- 24:07another state and you choose another
- 24:10action. You are literally a composite
- 24:13policy function. But nobody is feeding
- 24:17you a discrete state vector. Your
- 24:20biology, your your senses, everything
- 24:24you're experiencing
- 24:26is more or less a continuous state
- 24:30vector. So that is starting to introduce
- 24:33some frictions between biological agents
- 24:37and artificial intelligence agents. This
- 24:40is a broader discussion and there's a
- 24:42lot of interesting things in physics
- 24:44that I'm working on right now in terms
- 24:46of constraints and bounds but it's
- 24:49really interesting because biological
- 24:51agents like you and I don't have this
- 24:54computational intractability of oh we
- 24:57need to feed it a deterministic series
- 25:00of states in a in a discrete capacity.
- 25:04um that's that's nice and those agents
- 25:07might be able to extract more
- 25:08information and generalize faster in
- 25:10some contexts, but it also means that
- 25:12there are things it can't do because of
- 25:15our sensory capacity. So, a very
- 25:17interesting segue for you there, but
- 25:19what I have for you here is a beautiful
- 25:22animation
- 25:24of a Markov decision process, this
- 25:26broader idea of reinforcement learning.
- 25:28Okay, so on the left here, you're going
- 25:31to see the environment. I have a star of
- 25:34+ 10 and I have this red square of minus
- 25:3810. Okay. Now, the agent is trying to
- 25:41figure out this policy function as it
- 25:43engages with its environment. It's going
- 25:45to walk around and it's going to touch
- 25:48the star. It's going to touch the skull.
- 25:52And over time, it's going to hopefully
- 25:54learn the correct series of actions. In
- 25:58fact, you can see right here, this is
- 26:00the narrow network that is taking the
- 26:03state features as an input and it is
- 26:06trying to learn the appropriate output.
- 26:08And what have you noticed here? The
- 26:10agent is learning, oh we do not
- 26:13want to touch this red skull. Whenever
- 26:15we start to go this way or if we ever
- 26:17spawn on this square, we want to go away
- 26:20from that red skull and toward the green
- 26:23star. This is exactly how the learning
- 26:27process occurs in practice. The only
- 26:30difference in practice is what we always
- 26:34have this problem as biological agents
- 26:36of signal and noise. Right? So when you
- 26:40make a decision, you're never acting in
- 26:43against your own self-interests, right?
- 26:46But when you make that decision, you may
- 26:48yield a negative outcome. So good
- 26:52decision, bad outcome. But was it a good
- 26:54decision? That is the hardest part to
- 26:58analyze. This gets into poker theory. I
- 27:00recently read um thinking in bets. It's
- 27:02a great book and it talks about this
- 27:04idea of good decisions can yield bad
- 27:08outcomes. Bad decisions can yield good
- 27:11outcomes in a random environment. So
- 27:15that's why learning in some cases in
- 27:18reality as biological agents you and I
- 27:21is very challenging when you play a hand
- 27:24of poker. Did you make a good decision
- 27:27or were you lucky or were you unlucky?
- 27:30Did you make a bad decision and you were
- 27:32lucky? This is exactly what you need to
- 27:34discern. And it is not trivial. And if
- 27:37anyone says it's trivial, ask them why
- 27:40they're not a multi-time World Poker
- 27:42Tour champion. Okay? Because they should
- 27:45have no trouble distinguishing signal
- 27:48and noise. In reality, it's not a
- 27:50trivial problem. We all face it every
- 27:52single day. But what you're observing
- 27:54here is the idea of a Markov decision
- 27:56process, reinforcement learning in this
- 27:58environment. Look at this point. All
- 28:01right, by episode 90 almost 100, the
- 28:05agent has no interest in going towards
- 28:07this skull. It has no interest anymore.
- 28:10It is simply trying to go toward the
- 28:13star. All right. It is learning this
- 28:16through experience. The weights in its
- 28:18neural network are adjusting. It's
- 28:20learning this expected reward based on a
- 28:22series of its actions and states. And it
- 28:26now knows what to do.
- 28:29Okay. So that is this broader notion of
- 28:32reinforcement learning. After a series
- 28:33of iterations, right, it is now able to
- 28:36act on its own. It has the policy
- 28:38function and that policy function is
- 28:41going to dictate optimal actions for
- 28:43wherever it exists on this board in this
- 28:46environment as a function of the current
- 28:49state. Okay,
- 28:54understanding that let's go back to this
- 28:56idea now of
- 29:00deep hedging. We're going to apply this
- 29:02framework of reinforcement learning to
- 29:05the market making desk. We are going to
- 29:09be trying to train an agent now to hedge
- 29:14optimally. And this is the idea of deep
- 29:16hedging Hans Beller. All right. Instead
- 29:19of arbitrary frameworks, no trade
- 29:22regions, gamma bands, deterministic
- 29:24rehagging, whatever, the first thing
- 29:27that we're going to do is observe. Okay,
- 29:30there are more optimal ways in terms of
- 29:33risk that the desk is assuming in terms
- 29:35of optimizing the P&L at the end of the
- 29:38day to hedge and adjust our hedges over
- 29:43time.
- 29:45Number one, we are always going to have
- 29:47to assume model risk. It could be a
- 29:50classical model like stocastic
- 29:51volatility, jump diffusion. It could be
- 29:53data driven models like VAEEs, GANs,
- 29:55other generative models. I've used
- 29:57generative models in that capacity
- 29:59before to simulate tail risk events that
- 30:02may have not occurred yet empirically
- 30:05just as one example for you. But no
- 30:08matter what, even in this context, as
- 30:10we're training an artificial
- 30:12intelligence, a computational agent to
- 30:16make these hedging decisions for us, we
- 30:18have to assume a model. We have to give
- 30:20it information to learn from. And it's
- 30:23possible that we haven't experienced
- 30:25something that the agent has seen
- 30:28before. Just like us in the real world.
- 30:30All right, we've never experienced
- 30:32computational agents before. We could
- 30:34train on all the historic data we would
- 30:36like, right? Nobody knew that was going
- 30:38to happen. Then LLMs came and changed
- 30:42the world. And you can also have left
- 30:44tail world changing events that never
- 30:46occurred before either. So that's called
- 30:49model risk. There's no way around that.
- 30:52We have a data scarcity issue. These
- 30:54reinforcement learning agents are
- 30:55extremely difficult to train. They take
- 30:58a lot of data. They are very finicky to
- 31:01learn the policy function you saw in
- 31:03that little grid world example. It was
- 31:05reasonably easy. All right. But it still
- 31:08took a hundred episodes. 100. Okay. We
- 31:13can see right off the bat what to do.
- 31:16But it took a 100 episodes for that
- 31:19agent to learn what to do. So I want you
- 31:22to understand that in the context of
- 31:23learning for deep hedging, it is a
- 31:27tremendous tremendous issue from a
- 31:30computational tractability perspective.
- 31:33All right. So once we assume a model, we
- 31:35can select whatever one we would like in
- 31:36terms of a market simulator. We train a
- 31:39neural network agent. Okay. Does this
- 31:43look familiar? is exactly what we just
- 31:44talked about before in the grid world
- 31:46example. All right, the only difference
- 31:48is our state is not just the
- 31:49environment, right? The map effectively.
- 31:52It's inventory, current positions, the
- 31:55underlying assets, the current prices,
- 31:57relevant market observations,
- 31:59um path dependent features, realized
- 32:02volatility, running minax, time elapse,
- 32:04the list goes on. Whatever you think
- 32:06might be effective in explaining
- 32:08variation in the cross-section here. And
- 32:11again, this is for multiasset
- 32:12portfolios. you'll have implied
- 32:13volatility surfaces so on and so forth.
- 32:18The neural network is trying to
- 32:20approximate this map. Okay,
- 32:24neural networks and this is why it's
- 32:26called deep hedging are universal
- 32:27function approximators. The function
- 32:29we're approximating is what it is the
- 32:34policy function. It is an action given a
- 32:36current state and that state is just a
- 32:39vector. So everything is numerical.
- 32:41Everything is mathematical. If you have
- 32:42a news article that you want to feed to
- 32:44the neural network to approximate the
- 32:46policy, then you have to tokenize it
- 32:48first, interpret it and feed it into the
- 32:52state vector in a reasonable way. Okay,
- 32:55so that is literally how it functions.
- 32:56Then what do we do? Well, we have this
- 32:58reward function and then that reward
- 33:01function is what just expected P&L? No,
- 33:04we have entropic risk measures. We have
- 33:06expected shortfall. We have these risk
- 33:08measures that we are going to try to
- 33:10optimize for in the reinforcement
- 33:14learning context. And now here I can
- 33:17show you this beautiful dashboard. Yeah,
- 33:20I'm definitely not open sourcing this.
- 33:22This thing's awesome. Um, this
- 33:24is what it looks like to learn to hedge.
- 33:28I want you to see here that on the left
- 33:30we're observing one stochcastic path
- 33:32from the market. This is literally the
- 33:34exact same situation as we had before.
- 33:37Okay, it's the exact same situation.
- 33:39Meaning instead of a grid world, we're
- 33:41now observing a market and a market
- 33:43simulator. So we have this one
- 33:45stochcastic path, the underlying asset,
- 33:47we have whatever position we're taking
- 33:49in a European option contract. And here
- 33:52we have the hedge. I have the deep hedge
- 33:55that is the hedge that the neural
- 33:56network is learning and the classical
- 33:59delta hedge according to a rule set
- 34:01prior. So that delta hedge would be what
- 34:04you would need to hold to perfectly
- 34:06hedge the exposure. We're learning then
- 34:10a more optimal way to hedge as we
- 34:13continue to simulate more and more and
- 34:14more and more of these trades. And what
- 34:17you're going to observe is the tail risk
- 34:20changing relative to the arbitrary
- 34:23static delta hedge. Eg i.e. We are
- 34:28learning a more effective way to hedge
- 34:31the portfolio. And that's exactly what
- 34:34you are seeing here. You are seeing the
- 34:38tail risk during the training process
- 34:41fundamentally change relative to the
- 34:43fixed distribution of static delta
- 34:46hedging. This is the idea of deep
- 34:50hedging. In other words, it is more
- 34:54effective on average in the model sense
- 34:57to use an agent like this to go about
- 35:02constructing your hedges and rehedging a
- 35:05portfolio, a multiasset portfolio, a
- 35:07series of different exposures that you
- 35:10are going to be running in the market
- 35:13making sense. That's going to do it for
- 35:15this video on deep hedging for
- 35:18quantitative finance. I hope you
- 35:19enjoyed. But I hope you learned
- 35:21something. This video certainly took a
- 35:22tremendous effort to put together. So if
- 35:25you liked it and you want to see more
- 35:26like it in the future, please like,
- 35:28comment, subscribe, share. It helps me
- 35:30out tremendously. It is always greatly
- 35:32appreciated. Other than that, I'm going
- 35:35to thank you so much for watching and I
- 35:38will see you in the next
About this transcript
This page contains the full transcript of Deep Hedging for Quantitative Finance by Roman Paolucci, generated from the public captions YouTube serves with the video. The transcript has 5,691 words across 845 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.