Grok 4.5 + MCP finds a 1.9 Sharpe Strategy in 11 minutes! — Transcript
Full transcript
- 0:00Elon Musk's SpaceX
- 0:02Wait, that doesn't sound right. Just
- 0:04released Grok 4.5. And not only in some
- 0:07of the benchmarks, it is reaching to
- 0:09Claude Fable 5, but it is significantly
- 0:12cheaper. So, in other words, the quality
- 0:15to price ratio of this model is better
- 0:18than any other model out there. And that
- 0:20makes it a great candidate for us to do
- 0:23algo trading research with it. As
- 0:25always, what I'm curious about is
- 0:27whether this model can write me trading
- 0:29strategies, run back test optimization,
- 0:32and also Monte Carlo in order to make
- 0:34sure that the strategies that it is
- 0:36developing for me are not overfit. So,
- 0:38that's exactly what I'm going to do in
- 0:39this video. I'm going to test the model
- 0:41and ask it to write me real-world
- 0:42strategies. Before we continue, I got to
- 0:44mention I am not a financial advisor,
- 0:46and this video is for educational
- 0:48purposes only. So, with that out of the
- 0:50way, let's get right to the video.
- 0:53Now, here's the blog post released by
- 0:56the Grok team, and here are the
- 0:57benchmarks. So, you see in this one, it
- 1:00is right behind GPT-5. In this one, in
- 1:02this one, it is right before Opus 4.8.
- 1:06In this one, it is number one, so it's
- 1:08even beating Fable. And in this one, it
- 1:10is very close to Fable and GPT-5.5. But
- 1:14in this one, Fable is beating everything
- 1:16out of the water. Now, one thing we have
- 1:19to consider is that not only the price
- 1:21of this model is significantly cheaper,
- 1:24but it is also significantly faster in
- 1:26execution. And it is apparently also
- 1:29more efficient in token usage. So, that
- 1:31means that not only you'll you'll be
- 1:33getting a good enough quality, but you
- 1:35will also get it at a very much faster
- 1:38face. Now, a little bit of spoiling of
- 1:40the video, that's exactly what I got at
- 1:42the end of it. So, the result that I
- 1:44got, it was significantly faster than
- 1:47any other model that I have seen. But
- 1:48anyways, I don't want to expose just
- 1:50everything. So, let's continue. Here, it
- 1:52is claiming that this model is faster
- 1:54than flash models, and the speed is 80
- 1:57TPS or token per second, which is
- 1:59absolutely amazing. And you see it is
- 2:02also 4.2 times more efficient than Opus
- 2:064.8 in token usage, which is of course
- 2:09amazing. Now, for the algo trading side
- 2:12of things, as always, I'll be using the
- 2:13just framework, which also gives us an
- 2:16MCP, which I can just connect it to my
- 2:19agent
- 2:23the work. And for the editor, I'm going
- 2:26to be using the Zed editor, which is my
- 2:28new favorite code editor. It is written
- 2:30in Rust. It is extremely fast and just
- 2:32amazing, and it is also free to get
- 2:34started with. All right. Now, as for
- 2:36about the prompt, I've prepared this,
- 2:39which by the way, I will open source
- 2:41this and put it on GitHub and link to it
- 2:44in the description of the video in case
- 2:45you want to copy and paste it. But, I
- 2:47will read it very fast because this is
- 2:49the most important part of the video.
- 2:51Hey there. Do research and find me three
- 2:53trading strategies. Each one of them
- 2:55should have a sharp ratio above one
- 2:57since the beginning of the year until
- 2:59beginning of July. They should be trend
- 3:01following, include both and short
- 3:03positions filtered by a market regime
- 3:05indicator, such as the ADX, so trades
- 3:08only fire in trending conditions, and
- 3:10risk 3% of the account capital per each
- 3:13trade. However, since we are in a bear
- 3:15market, I want them to have more focus
- 3:18on short trades. And the entry rules for
- 3:21short and long positions should not
- 3:23necessarily be the same. In the first
- 3:25strategy, once the position is open,
- 3:27close it using a trailing stop based on
- 3:30some sort of indicator. In the second
- 3:32strategy, close the position using a
- 3:35fixed multiple of ATR from entry. You
- 3:38can test them on BTC, ETH, and SOL,
- 3:41Trump, Melania, and DOGE on the hourly
- 3:45and 4-hour time frames. The strategies
- 3:48should trade these symbols individually
- 3:50or altogether. Every time you develop a
- 3:52strategy, validate the results using a
- 3:55statistical significance test before
- 3:57writing the full strategy. Only proceed
- 3:59if the strategy's metrics demonstrate
- 4:01genuine statistical significance. Feel
- 4:04free to use optimization to improve the
- 4:06results. At the end, apply Monte Carlo
- 4:09simulations to ensure the results are
- 4:10not overfit. Continue until you find
- 4:13strategies that fully meet the criteria
- 4:15requested. Do not prompt me in the
- 4:17meanwhile. Good luck. And by the way,
- 4:19guys, I am using the Z editor, and the
- 4:21way that I am using Grok 4.5 is by
- 4:26picking it from here. So, I have already
- 4:28added my Open Router credentials, and
- 4:31now I I am able to pick whichever model
- 4:34that's being offered by them. And if I
- 4:36simply look it up, I can easily find
- 4:38Grok 4.5 and pick it. Now, the catch is
- 4:41that at the moment of recording the
- 4:42video, Grok 4.5 is not available to EU
- 4:46citizens, and the way I am accessing it
- 4:48is by simply using a VPN. Because if I
- 4:51were to use a an EU IP, this would have
- 4:54gave me an error. And by the way, here
- 4:56there is a typo, so let's fix it. All
- 4:58right. So, everything looks good. Let's
- 5:00just hit enter and wait for the agent to
- 5:03do the research. However, there is one
- 5:05thing that I would always like to do.
- 5:07So, I'm just going to copy this, delete
- 5:09it, and first I will say, "Hey, do you
- 5:11have access to Jesse MCP? Can you do
- 5:13research for me?" and hit enter. So, the
- 5:16reason I'm doing this because I want to
- 5:17ensure that my setup is working before I
- 5:20give the agent some sort of task that is
- 5:22really huge. Because if it doesn't have
- 5:25access to the correct tool, then it
- 5:27would be pointless for me. All right, so
- 5:28it turns out my VPN wasn't working, so
- 5:30I'm just going to fix that. All right,
- 5:32so it seems to be figured out, so let's
- 5:34retry this, and this time hopefully it
- 5:37should work. And there we go. So, yes,
- 5:39it has access to Jesse MCP. It says that
- 5:41everything is fine. All right, so let's
- 5:43begin the actual research process. So,
- 5:45I'm going to paste back to our prompt
- 5:48and hit enter. All right. So, as you can
- 5:50see, it is using the MC tool just fine.
- 5:54It is listing the indicators that it has
- 5:56access to. It's getting details because
- 5:58it is preparing to start writing
- 6:00strategies. And it is a thinking model,
- 6:02of course, so it's going to do lots of
- 6:04thinking. And here also, if I hover my
- 6:07mouse on it, I can see the context. So,
- 6:1012% has been already used and this model
- 6:13has 500 thousand context window, which
- 6:16is pretty big. So, it's not 1 million
- 6:18like Opus, that is how I would like it
- 6:21to be, but it's still 500 thousand is a
- 6:24lot. So, hopefully it's going to be more
- 6:26than enough for our use case. All right.
- 6:28So, I'm going to go back and come back
- 6:30later until this one is updated. All
- 6:32right. So, I faced one issue and that
- 6:36was because the limit for my API key had
- 6:39reached. So, this isn't really something
- 6:41you should worry about. Just ensure that
- 6:43the API key that you generate on
- 6:45OpenRouter has enough credit allowance.
- 6:47However, next I faced another issue and
- 6:50I cannot actually see it here to show it
- 6:52to you, but it was about the context
- 6:54window. And looking at here, we can see
- 6:56that 75% of the context window is full
- 6:59and apparently the request that it tries
- 7:01to send it is passing this number and as
- 7:04a result OpenRouter is returning an
- 7:06error. Now, the thing is the Z editor
- 7:08recently added an option for auto
- 7:10compacting, which is really helpful here
- 7:12because if it compacts the values here,
- 7:16then we won't have this issue. But I
- 7:17think the problem is this that the
- 7:19threshold is set to 90%. So, if we lower
- 7:22this to, let's say, 70% and hit retry,
- 7:26hopefully it should be resolved because
- 7:28now it tries to compact the context
- 7:30first and then continue everything by
- 7:32itself. So, apparently it's getting
- 7:33errors even for compacting stuff. So,
- 7:35I'm guessing that if we start over, it's
- 7:38going to work this time. So, I will just
- 7:40go up here and press enter again. So,
- 7:43now the context window is cleared again
- 7:45and this time we have the auto
- 7:47compacting option set to 70%. So,
- 7:50hopefully we won't face the same issue
- 7:51again. Now guys, I could have deleted
- 7:53this part from the video, but I didn't
- 7:55because I think this is some sort of
- 7:57issue that you might also face. So, I
- 7:59wanted to include that in the video. So,
- 8:01you guys know how to deal with it if you
- 8:03face the same issue. And you might be
- 8:05wondering, "Okay, so why didn't you face
- 8:07this issue before?" Well, in the
- 8:09previous videos I was using tools such
- 8:11as Cloud Code or OpenAI Codex and with
- 8:14those tools, they handle the auto
- 8:17compacting by themselves. So, I never
- 8:19had to worry about it. Sure, I was still
- 8:21using my Z editor, but those tools were
- 8:25the ones actually handling the auto
- 8:27compacting and context management stuff.
- 8:30But, now that I'm using the Z editor's
- 8:33built-in agent window, I have to worry
- 8:36about this. But, thankfully recently
- 8:38they added that option which allows us
- 8:40to auto compact. So, I just needed to
- 8:42change its default threshold from 90
- 8:45into 70. All right, so the results are
- 8:48in and it actually took only 11 minutes.
- 8:51Now, it is removed from here probably
- 8:53because I gave it another prompt after
- 8:56this. So, you're just going to have to
- 8:57trust my word, but it only took 11
- 9:00minutes for it to meet the criteria that
- 9:02I had set for the model. That means
- 9:05three strategies with a sharp ratio
- 9:07above one. And here's the output. It
- 9:09gave me couple of tables and URLs for me
- 9:12to check out the results on the
- 9:14dashboard with some pretty charts and
- 9:16stuff. However, I found it to be just a
- 9:18little bit all over the place because it
- 9:20gave me like multiple tables and that
- 9:22just wasn't fine. So, I said, "You know
- 9:24what? Give me just one table with all
- 9:27the results combined." And it gave me
- 9:29this which is perfect. But, before I
- 9:30open it, I want to show you this that it
- 9:32also gave me report files and these are
- 9:35really great. So, for example, let's
- 9:37open this one and you see it's in the
- 9:39markdown format. So, if I press command,
- 9:42shift, and V on my macOS machine, I can
- 9:45see the rendered result of the markdown
- 9:48format on my Z editor. And here you can
- 9:50see that it's telling me this was the
- 9:53objective, this is the target, the
- 9:55exchange that it used, how much to risk
- 9:57per each trade, and the style of the
- 9:59strategy, and things like that, and a
- 10:01pretty summary of it. Then we can see
- 10:03the entry validation or the RST, which
- 10:05is stands for rule significance test,
- 10:08which basically tells us if the entry
- 10:10rules of the strategy were just pure
- 10:12luck or if there was some sort of edge
- 10:14in it. And you can see the P value is
- 10:17really low, which is a great sign that
- 10:19the entry rules of the strategy were
- 10:21definitely not noise and there was some
- 10:23sort of edge in them. And it also give
- 10:25me the URL for the dashboard. So, if I
- 10:27wanted to open it in the dashboard, I
- 10:29could have just click on this. And then
- 10:31we can see that it's telling me about
- 10:33the iterations that it took. So, what
- 10:35was the first version of the strategy
- 10:37that it wrote, what was the second one,
- 10:38and what was the final one. And here's
- 10:41the results of the final strategy. So,
- 10:44the max drawdown is minus 26%, the net
- 10:46profit is 83%, the sharp is 1.93, which
- 10:50is absolutely great, the number of
- 10:51trades that it took, and things like
- 10:53that. Next, we have the result of the
- 10:55Monte Carlo simulation, which is
- 10:57basically a stress test, which tells us
- 11:00whether or not the strategy is going to
- 11:02be overfit or not. Or to put it better
- 11:05in a more statistical term, it will tell
- 11:07us how likely the strategy is to be
- 11:09overfit because you can never be 100%
- 11:12sure if the strategy is going to be
- 11:14overfit or not. It's just an estimation
- 11:16and this is exactly what gives us to us.
- 11:19And this is also the result. It just
- 11:21doesn't have the chart. Like otherwise,
- 11:22I didn't even need to open the Monte
- 11:25Carlo page. So, anyways, next we can see
- 11:27that it's telling me, "Okay, did it
- 11:29reach the target?" Yes, it did. And then
- 11:31the recommended next step. And we can do
- 11:33this for all the other ones. So, we have
- 11:35two more here. Next, let's take a look
- 11:37at the table that it gave me, so I can
- 11:40just open the results on the dashboard
- 11:42so we can see it. So, the backtest of
- 11:44the first one, the RSC, and the Monte
- 11:46Carlo. Now, I want to begin with the
- 11:48significance rule test because this one
- 11:50is the first priority when you want to
- 11:53develop a strategy because if the entry
- 11:55rule of the strategy is noise, and it
- 11:58doesn't have an actual edge in it, then
- 12:00all the other results that you're going
- 12:01to get for it will be meaningless. This
- 12:03is really important to always begin
- 12:05with. And in our case, you can see that
- 12:07yes, this would have been the return of
- 12:09the strategy in the simulation, and
- 12:11these are the simulations that were pure
- 12:14noise. And the fact that the strategy's
- 12:16returns are significantly higher than
- 12:18these is a great great sign. And here
- 12:20you can also see the results translation
- 12:23from the dashboard. It's telling us that
- 12:24it is exceptionally strong evidence of
- 12:28genuine edge. And here we can see the DP
- 12:30value. Here is the annualized return,
- 12:33the observed mean, and then we can see
- 12:35the number of simulations that were
- 12:37executed, which is 2,000. All right, so
- 12:39this looks awesome. Let's close it. Now,
- 12:41while that's going, I want to quickly
- 12:42remind you guys about our Telegram. It's
- 12:44the fastest way to get notified about my
- 12:46future work, whether it's a new tutorial
- 12:49or a tool that I create. Also, don't
- 12:50forget to check out our free Discord
- 12:52where more than 5,000 members like you
- 12:54and I are hanging out there and helping
- 12:56out each other with algo trading so we
- 12:57can all succeed together. The links for
- 12:59both are down in the description. Next,
- 13:01we have the result of the backtest
- 13:03itself. Now, here's the equity curve,
- 13:05and here's a log version of it. So, we
- 13:07can see while all the other assets, so
- 13:09BTC, ETH, and SOL, and DOGE, while they
- 13:12were all going down, we were making
- 13:14money here. However, one thing I noticed
- 13:16is that the strategy only makes money if
- 13:19there's a huge crash. So, for example,
- 13:21in this case. You see, but if the price
- 13:24is in a range, it will slightly bleed.
- 13:26So, this is something to be aware of,
- 13:28but then again here we had a huge crash
- 13:30and so the the strategy's equity curve
- 13:32just started going up again. So, if
- 13:34you're going to trade a strategy like
- 13:36this, you have to be prepared for a low
- 13:39win rate and the fact that for months it
- 13:42wasn't doing great because the market
- 13:44was in a range. So, this shouldn't be
- 13:46the only strategy that you're trading or
- 13:48otherwise it's going to be super hard
- 13:49for you mentally to just wait and let
- 13:52the strategy do its thing. Next, we have
- 13:54the five worst max drawdown periods.
- 13:56This one is the obvious one of course.
- 13:58Next, we can see the monthly returns.
- 14:00So, we have three months in green and
- 14:03three months in negative, but the months
- 14:05in green are significantly higher, which
- 14:07is why the annual return is also really
- 14:09high. Next, we can also take a look at
- 14:11the metrics of the strategy. So, the P&L
- 14:14was 82% and by the way, guys, this is
- 14:17not for the whole year. So, we need to
- 14:19also remember that. The max drawdown is
- 14:21minus 26%. The average win to loss ratio
- 14:24is three, which is a really important
- 14:26metric for us. The annual return is
- 14:29240%.
- 14:30The max underwater period is 116 days.
- 14:34Again, what I just explained with the
- 14:36psychology side of the strategy. So,
- 14:37this one's going to be really hard to
- 14:39execute mentally. The win rate is 41%,
- 14:43which basically means only four out of
- 14:45every 10 trades that the strategy was
- 14:47taking was in profit and the sharp ratio
- 14:49is 1.93. And the average trades that we
- 14:52took per month was 8. 58. So, this is
- 14:56also pretty important. So, do not expect
- 14:58a strategy like this to take trades like
- 15:00every day. Next, we can see the result
- 15:03of the Monte Carlo simulation, which
- 15:05will basically tell us how likely the
- 15:07result of the strategy is to be overfit.
- 15:09So, you see the original sharp ratio was
- 15:12two, the median was 1. almost four, and
- 15:15the best 5% of the simulations was 3.66.
- 15:19So, the strategy is not likely to be
- 15:21overfit. It is significantly lower than
- 15:24the best 5% and pretty close to the
- 15:26median. So, this is a really good result
- 15:29and the number of scenarios that we took
- 15:30in this simulation was 200. So, some
- 15:33people say increase this to like 500 to
- 15:361,000. You could do that, but in my
- 15:38personal test, even 100 was enough for
- 15:40most cases, but 200 is definitely more
- 15:43than enough for me. Next, let's take a
- 15:45quick look at the backtesting and the
- 15:47Monte Carlo of the other two strategies.
- 15:50So, this is the result of the
- 15:51backtesting of this one. Well, sure, it
- 15:54is making more money during the
- 15:56downturns, but in the ranging markets,
- 15:59this one is bleeding significantly
- 16:00higher than the first strategy. So, this
- 16:03is not really a good sign and I don't
- 16:05really think I would want to trade this
- 16:07one at all. So, let's just jump into the
- 16:10third strategy. Here's the result of its
- 16:12backtest. So, again, I would say it is
- 16:14bleeding more during the ranging market
- 16:17and this is not really a good sign, but
- 16:19overall, the result isn't that bad and
- 16:21if you take a look at the Monte Carlo
- 16:23result of it, you can see the original
- 16:25backtest had a sharp ratio of 1.28
- 16:29while the median was 1.56 and the best
- 16:325% was 3.40. So, the strategy is not
- 16:35likely to be overfit or in better terms,
- 16:38it is less likely to be overfit and I
- 16:40like these results, but again, I don't
- 16:42think this one is holding a candle
- 16:45compared to the first result that we
- 16:46got. So, all in all, I would say this is
- 16:48really good result. However, before you
- 16:51go and jump and start trading this
- 16:53strategy or something like this,
- 16:55remember that you should backtest it on
- 16:57out of sample. So, in this case, we were
- 16:59using only this year's result, but
- 17:01before actually starting to use it, you
- 17:03need to run it on previous year's data,
- 17:05but that's also going to be some sort of
- 17:07personal judgement because the thing is
- 17:09this year's market was very different
- 17:12than previous years that I've seen. It
- 17:14was definitely bear market, but not an
- 17:16obvious one. It's not like we had a real
- 17:18bear market followed by a bear market.
- 17:20So, that's not what we got. So, it was
- 17:22very different. And the reason that I
- 17:25asked the agent to generate strategies
- 17:27that were more focused on shorting was
- 17:30because I believe that we are still in a
- 17:31bear market. But if my belief is wrong,
- 17:34there's a good chance that these
- 17:35strategies are going to stop working.
- 17:37So, I might also run a back test for it
- 17:39for, let's say, 2022 because it was a
- 17:42similar year. But overall, in algo
- 17:44trading you want to back test your
- 17:46strategy on as much data as possible to
- 17:49gain more confidence in it before
- 17:51actually risking your hard-earned money.
- 17:54Now, in case you are curious to know how
- 17:56much tokens I just burned for this
- 17:58experiment using the Grok 4.5 model,
- 18:01well, here's the result. I spent about
- 18:04$18 on it, more than 13 million tokens
- 18:07in total with a cash hit rate of 81%.
- 18:10Now, for $18, which is almost 20, you
- 18:13could definitely get a subscription plan
- 18:15of a provider such as Anthropic, OpenAI,
- 18:19GLM, you name it. So, there are so many
- 18:21options out there. And if you're going
- 18:22to spend this much, I don't think it's
- 18:24feasible to spend it on API tokens. Now,
- 18:28the Grok team or xAI is offering their
- 18:31own coding plans, apparently, and they
- 18:33begin from $30. But the CLI, which I'm
- 18:36guessing is the one that we need, is
- 18:38called Super Grok Heavy. So, the light
- 18:41or this plan is not going to be enough
- 18:43for us for coding. And the Grok Heavy
- 18:46begins from $100 per month, and this is
- 18:49just for 3 months. So, after that, it's
- 18:50going to be $300 per month. Now, for the
- 18:53amount of tokens that they are offering,
- 18:56this might be good for you, but I
- 18:57personally do not find this a good
- 18:59enough deal because on Anthropic, for
- 19:01example, you can get the starting max
- 19:04plan for just $100, and it will be more
- 19:08than enough. So, I don't think the Grok
- 19:10team is offering a good deal at the
- 19:12moment. So, even though that the
- 19:13performance of the model was really
- 19:15good, this was just over the API. And
- 19:17sure, apple-to-apple comparison, like if
- 19:19you were to compare the Grok API to
- 19:22another API such as, let's say, Claude
- 19:25Opus or Fable, yes, this this Grok model
- 19:28would have been a better deal. But now
- 19:30that we have subscription plans on other
- 19:31providers, I don't think it is exactly
- 19:34what I would want to use. So, just be
- 19:36careful about that. And as always, I
- 19:38will be submitting these strategies that
- 19:40I just showed you guys to our strategy
- 19:42index page, where you can browse
- 19:44previously submitted strategies by me or
- 19:46others. And you can also click on a
- 19:48certain strategy and get to see its
- 19:51results for different trading periods,
- 19:53symbols, or time frames, which is pretty
- 19:54helpful. You can also get the source
- 19:56code of the strategies by clicking here,
- 19:58of course. I hope you guys enjoyed the
- 19:59video. If you did, please like the video
- 20:01and let me know what you think in the
- 20:03comment sections. And make sure to
- 20:05subscribe to the channel because I
- 20:06release tutorials just like this one all
- 20:08the time. Thank you so much for
- 20:10watching. I'll see you in the next one.
About this transcript
This page contains the full transcript of Grok 4.5 + MCP finds a 1.9 Sharpe Strategy in 11 minutes! by Algo-trading with Saleh, generated from the public captions YouTube serves with the video. The transcript has 3,912 words across 528 segments, with the original timestamps preserved so you can click any line to jump to that moment in the embedded player.
What you can do with it
Use the transcript to take notes, quote the speaker, build a study guide, generate a summary with ChatGPT or Claude via the YouTube Summary tool, or export it as a timed subtitle file with YouTube to SRT. You can also re-open it in the transcriber to translate the transcript into 100+ languages.
Free YouTube transcript tool
YouTube2Text is a free YouTube transcript generator — no signup, no daily limit. Paste any YouTube link and get the full transcript instantly, with timestamps, click-to-jump, translation to 100+ languages, AI prompts for ChatGPT, Claude, and Gemini, and exports to TXT, SRT, VTT, or Markdown.