[00:00] In this video, I'll show the thought process and four steps I always use when developing a new trading strategy. The approach is generic and compatible with almost every strategy. I've shown bits of this process in my past videos, but I wanted to make a standalone dedicated video I can reference back to in the future. [00:14] The four steps are In-Sample Excellent, In-Sample Monte Carlo Permutation Test, Walk Forward Test, and Walk Forward Monte Carlo Permutation Test. I'll show an example going through all four steps and generalize the concepts so you can try it with your own strategies. [00:27] First, I'll show how I assess a trading strategy. I'll use the moving average crossover as an example. We load in candlestick data and compute a fast and slow moving average. At each bar, we check if the fast moving average is above the slow moving average. [00:40] We create a signal that denotes the position of the strategy at each bar. 1 means we have a long position following that bar, and 0 means we have no position. Regardless of what the strategy is, one should be able to create a similar signal, a value denoting the position of the strategy after each bar. [00:55] Now if we compute close to close returns and shift them forward by one bar, we can multiply the position signal by the shifted returns to get a return for each bar that is attributable to the strategy. These strategy returns at the same granularity of the bars are what I use to compute objective functions, [01:10] such as the profit factor or the Sharpe ratio. By having a return for each bar instead of each trade, objective functions are passed much more data, and the calculations and results are much more stable. The book Testing and Tuning Market Trading Systems provides convincing arguments for why these higher granularity returns are superior. [01:26] That's where I got the idea and I'm thoroughly convinced, both from what the book says and my own experience. I'll not scrap the moving average crossover because it's lame, and instead I'll use the slightly more interesting Donchian Channel Breakout strategy I showed in my previous [01:39] video. Briefly, the strategy goes long when the current close is the highest over a given lookback, and it goes short when the current close is the lowest over a given lookback. The idea of this is to trade in the direction of range breakouts, so we can ride the trend [01:52] when extended trends occur. The codes for the Donchian breakout are short, and the output is the signal vector. It will have a value on each bar denoting the strategy's position, 1 or negative 1. This strategy needs a look back so we can touch several and pick the best. [02:05] Here is a grid search to do that. We look through a wide variety of look back values and find the one with the best profit factor. I ran the optimization on hourly Bitcoin data from 2016 through 2019. Over those four years, the best look back was 19 with a profit factor of 1.08. [02:21] I first like to look at the in-sample performance. Since these are log returns, the cumulative sum of the strategy returns will give us a cruise back test. Cool, so what are we actually doing here? Well, we're optimizing a mediocre trend follower, but more generally, we have an idea for a [02:35] trading strategy, a set of development or in-sample data, and a way to fit, optimize, or select the best version of the strategy for the data. These are the generic components for essentially every trading strategy, whether we're optimizing [02:48] or look back for a trend follower, selecting the best chart patterns, or training a fancy neural network. This is how trading strategies are optimized. So with that in mind, let's look at our in-sample performance again. Currently, we're in the development stage. Here, I ask myself [03:02] two questions. Is this excellent, and is it obviously overfit? These are in-sample results, so they should be pretty damn good. Maybe you could come up with some thresholds for an objective function to decide if it is excellent, but really I think it depends on the nature of the strategy. [03:15] I mainly look if the strategy has periods of inconsistency. I would drill down and really look at what is happening when the strategy is working poorly versus working well. Maybe there are ways one could improve this strategy. Maybe some YouTuber made a video about improving the very strategy you're looking at. [03:30] But this is the development stage. This is the time to study your strategy, test potential improvements, and refine your optimization process. If this was the strategy I was working on, I'd study it and try to improve it further. I would not say this is excellent. [03:43] I expect a bit more from example performance than this. But to keep the video going, I'll say it looks good enough here. Now the other question. Is this obviously overfit? This one is a bit harder to answer. I suppose one gets a feel for it eventually, [03:55] but if your results are suspiciously good, like a 100% win rate, you're probably overfitting. Or maybe you just accidentally allowed a future leak, which obviously needs to be fixed. But if you suspect overfitting, you may want to dial back the complexity of the strategy. [04:09] Ultimately, the answer to this question in the development stage should either be yes or not obviously. I'll show an obviously overfit strategy later. Once we're satisfied with our in-sample results, the question becomes, was this excellent in-sample [04:21] performance bound due to patterns intrinsic in the data, or was the good in-sample performance bound just because our optimization process was powerful enough to find something in noise? In other words, is data mining bias the main contributor to our excellent in-sample performance? [04:35] The problem with optimization is it works. If we compare multiple configurations of a strategy, one will be the best, but it will always have a data mining or selection bias. Of course, if we optimize even a great strategy, there will be some data mining bias, but a good strategy's [04:51] in-sample performance will mostly be from patterns in the data. But if our strategy is trash, then its in-sample performance will be entirely due to data mining bias. Our null hypothesis is that our strategy is garbage. We will use the in-sample Monte Carlo permutation test to disprove [05:06] our null hypothesis. So how does this test work? We optimize the Donchian breakout on these four years of price data and our optimized strategy has a profit factor of 1.08. But if we create a random permutation of this data, we get [05:19] something like this. In this random permutation, any legitimate patterns that existed in the real data are no longer present The permutation is just noise with nearly identical statistical properties If we optimize the Donchian breakout on this permeated data we get a profit factor of 1 The optimized strategy did better [05:37] on real data than it did on permeated data. This gives a small amount of evidence that our null hypothesis is false, because if the optimized strategy did just as good or better on random data, we could presume the main [05:49] contributor to the good and simple performance we saw was from data mining bias, but this was only one permutation. If we created many more and on each permutation we optimized the strategy, we could get an idea of how powerful the data mining bias induced by our [06:02] optimization is. If the optimized strategy's profit factor of 1.08 that we found on real data is better than what we found on the vast majority of permutations, we can disprove our null hypothesis. I don't think there's any value in looking at the equity curves of permutations, as we can simply [06:18] compute the objective function on each of these, but I think this helps visualize what we're actually doing, and it looks cool. Now I'll show the algorithm I'm using to generate these permutations of price. This function has a few parameters. The first is the data. It takes either [06:30] a data frame of open high low close prices or a list of such data frames. The option for the list of data frames is for permuting multiple markets. I'll talk about this option later. Start index is where the permutation starts. When set to the default zero, the function will permute all the [06:46] data given. We will use this parameter later when we talk about the walk forward permutation test. Since we cost data as either a list of data frames or a single data frame, we handle that first. If it's a single data frame, I put it into a list and set end markets to one. [06:59] If we have multiple markets, we ensure that their indexes are identical. Then we allocate space to store relative prices for each market and bar, and the first bar of each market. The first bar will be unchanged. It has a size of 4 to handle the open, high, low, and close price. [07:15] Now we compute prices on each bar relative to that bar's open. We loop through each of the markets and get the logarithmic prices. I copy the first bar at the start index here. The open is subtracted from the high, low, and close prices. [07:28] Since we're dealing with log prices, we're essentially recording the percentage off the open of each of these prices. The relative open is the current open minus the prior close, the gap. Then we copy over these prices into the arrays we made earlier. [07:41] Now we get the indices of the real data. We'll use these to shuffle the relative prices. We shuffle the indices once for the intrabar quantities, and again for the gaps. The gaps have little effect on crypto data. The crypto market never actually closes, so the open of one bar will usually only be at most a few ticks away from the prior bar's close. [07:59] But for daily stock data, the open can be quite far away from the prior close. After shuffling, we can now string together a permutation. We loop through each market and allocate space to store the permuted bars. We get the log prices of the real data. [08:12] We copy it into the permuted data before the start index. If the start index is set to zero, nothing happens here. Then we copy the start bar, the first bar of the permutation. We loop from the start bar to the end of the data. [08:24] We first set the open price. The zero is the index of the open and the three is the index of the close. To get the permuted bars open, we add the relative open value to the prior bar's close. Then we add the relative high, low, and close to that open value to get the rest of the [08:39] permuted bar's prices. After the loop, we exponentiate the prices to get them to the normal scale and add the bars to a data frame. The function will return either a single data frame or a list of data frames the same as what was passed to the function. Now we can pass in a data frame of real data and [08:54] return a data frame of permuted data. The first open and last close are exactly the same on both the real and permitted data, so the overall trend of the data is preserved, but the path the price takes between those two prices is completely different. The goal of the permutation algorithm [09:09] is to create bars that have the same statistical properties as the original. To compute some close the close returns we can see that the mean, standard deviation, skew, and kurtosis are all nearly identical. Now let's load in Ethereum data for the same time period. Here is a plot of [09:23] Bitcoin and Ethereum in 2018 and 2019. We can see that they're obviously correlated, and if we permute them together, the correlation between the two markets stays the same. I won't cover the multi-market case here beyond this, but if your strategy involves two or more markets, the [09:38] permutation tests can still be applied. While the algorithm produces permutations with many similar statistical properties to the original, it is not without its flaws, as price is not a random walk. Real prices have volatility clustering and long memory, both of which could be a topic for a [09:53] different time, but the permutation algorithm will destroy both of these properties. If your strategy is heavily focused on one of these properties or some other property that the permutation algorithm doesn't preserve, the Monte Carlo permutation tests can be optimistically biased, [10:08] but this really isn't a forward problem as if your strategy cannot pass the permutation test even with a potential optimistic bias then you know your strategy is probably overfitting. Now that we've gone over the bar permutation algorithm we can return to where we were. [10:21] We optimized the donchian breakout on hourly bitcoin data from 2016 through 2019 and the best look back gave a profit factor of 1.08. When we optimized the strategy on many different price permutations we found that the results were worse than what we got on real prices. We've essentially [10:37] We already done the in-sample permutation test, but now I'll show you the code and how to apply it. First we load our data into the four years we're using to train. We call the optimized notion function and this gives us our real profit factor. [10:49] Now we can do the permutation test. We set the number of permutations then lose that many times We get a permutation of the bars and optimize our strategy on them to get a permuted profit factor If the permutation profit factor is just as good or better than [11:03] the real profit factor, we increment the permutation was better count. After the loop, we can calculate a quasi-p value, the number of times the permutation was better divided by the total number of permutations. This value is roughly the probability that our real profit factor was found mainly [11:18] due to data mining bias. This next part isn't really necessary, but I like to plot a histogram of the profit factors from the permutations, then add a line showing where in the distribution the real profit factor fell. I ran this test with 1000 permutations [11:34] and got this. Only a couple permutations did better than the original, so the p-value is very low, 0.3%. If a sufficient amount of permutations are done, the permutation distribution should be roughly bell shaped. If the distribution looks really weird there's probably an issue with [11:49] your code. I like to see the p-value below 1% so I would call this a pass. Now I'll quickly show you a strategy that is over fit. This function fits a decision tree. We compute three indicators, just basic price differences, then create a classification target whether the next 24 hours [12:05] go up or down. Then we create a decision tree, train it with our indicators and target, and return the model. Notice that I've set the minimum samples per leaf very low. This is one of the key regularization parameters and is pretty much guaranteed to overfit. To test the model we can [12:20] use this function. We compute the same indicators and use them to predict the model. Then we can use the model's predictions to create a position vector. We'll go long when the tree predicts the price will go up and we'll go short when the tree predicts the price will go down. Then finally we [12:33] compute the problem factor of that signal. Here are the in-sample results for our decision tree. Again, this is when I ask myself, is this obviously overfit? The answer is yes. Generally speaking, if your backtest ever looks like this, you have a future leak or you're horribly overfit. [12:47] But if we didn't know any better, we can use the in-sample permutation test to crush our dreams. And the test does the job. The model performs just as good or better on the permutations. When you see this, it's time to throw your strategy idea in the trash. [13:01] Ideally, you should use the test with as many permutations as possible. I think 1000 is a reasonable minimum. This of course means we have to optimize our strategy 1000 times, and it will probably take some time. If optimizing your strategy 1000 times is simply not feasible, you probably have a very complex strategy or a very poorly coded strategy. [13:21] In which case I suppose 100 would be sufficient, but I would say that's a hard minimum. This test provides a quasi p-value that roughly indicates the probability that your in-sample results were primarily found from data mining bias. [13:33] I generally don't continue if it is over 1%, but don't treat that like a target. This is a measure. If a measure becomes a target, it is no longer a good measure. Basically, if you fiddle with your strategy enough, you could probably make this test pass on anything, so don't overuse it. [13:49] We've not seen the ensample permutation test pass the Donchian channel breakout and reject the decision tree nonsense. But why even do this? Couldn't we just try the strategy on 2020 data? If it worked on data that wasn't used for the optimization, then the strategy is probably not overfit. [14:05] Well, sure we could, but once out-of-sample data is used even once, it is no longer truly out-of-sample. Suppose we optimize strategy A on 2016 through 2019, then test it on 2020 data, and we find our out-of-sample results to be decent. [14:19] But then we come up with another idea, strategy B. We optimize strategy B just the same, on 2016 through 2019, then we also test it on 2020. We find that strategy B did better than strategy A. [14:31] So one might think the new idea was better, but now there is a selection bias. Strategy B did better compared to strategy A. The results of B on 2020 data are inflated by selection bias. [14:43] Now realistically, I've already tested many things on 2020 data, and it definitely isn't out of sample for me. Rather, it is a validation set. It is a good idea to walk forward and optimize strategy, testing it on data it did not use to optimize. [14:56] That is how the strategy will have to trade in reality after all. And if we test a strategy on data it did not use to optimize, the results will not benefit from any data mining bias. However, if we walk forward test 100 different strategies on 2020 data and select the best [15:11] one, there will be a massive selection bias. Selection bias can allow us to effectively overfit the validation data, despite it not being used for strategy optimization. Every time we reuse out of sample data, or rather validation data, the selection bias [15:25] is adding up. This is why we use the in-sample permutation test. We can detect that our idea is bad before we waste the out-of-sample data or stack up even more selection bias on the validation data. Anyways, with all that in mind, our optimization of the Donchian breakout lookback [15:41] passed the in-sample permutation test. So now let's walk forward the Donchian breakout. This function will return the walk-forward signal. One of the parameters is the train lookback, how much data to optimize on. I have it set to optimize on the last four years by default, [15:56] assuming hourly data is used. The train step is how often we re-optimize. I set it to 30 days. Ideally, you should retrain strategies as often as is feasible, but to make the code run fast enough to accommodate my brain rot, I used 30 days. We set the index of the next optimization [16:12] and loop through all the data. Every time the index is equal to the next train variable, we re compute the new signal and increment the optimization index by the train step This is pretty inefficient code but it simple and it works Here are the results of the walk signal on 2020 data It had a profit factor of 1 [16:31] which is worse than what we saw in sample. Generally, that is to be expected, as these results do not benefit from any data mining bias. The only bias in play here is the potential selection bias. If we had already walk-forward tested other strategies using this data. [16:45] At this stage, I ask myself, is this worth trading? The answer is subjective. It depends on your standards. Maybe you have higher or lower standards than me. For me, I wouldn't bother with this. It kind of sucks. But the line did go up, and to keep the video going, I'll say this is good enough. To get these results, we optimized the Donchian breakout on these four years of data. [17:04] These four years are the first training fold of the walk-forward. After the first training fold, the walk-forward function can output a signal that we can test. Since our walk-forward results were satisfactory, we're assuming that whatever patterns the [17:16] strategy learned or optimized on from past data are also present in this future unseen data. But what if our optimized strategy is actually worthless? What is the chance a worthless strategy could have achieved walk-forward results just as good as what we found? If we generate a [17:31] permutation of the data after the first training fold, any legitimate patterns in this data will no longer be present. There are no legitimate patterns in this permutation. If we walk forward the same strategy on this permutation and compute its profit factor, we get an estimate of a profit [17:45] factor that a worthless strategy could produce, and if we generate many permutations, we get a distribution of what worthless strategies can produce. If our real walk-forward results are to be attributed to patterns learned from past data reoccurring in this future data, then our real [17:59] walk-forward profit factor should be better than the vast majority of profit factors produced by worthless strategies. This is the walk-forward permutation test. You will notice the code is very similar to the in-sample permutation test. We load in our data and set the train window to four years, [18:14] Then we compute the walk forward signal. With the signal we can compute our real walk forward profit factor. Then we set the number of permutations and loop through them. We call the same get permutation function, but we set the start index to the train window to only [18:28] permute data after the first training fold. To compute the profit factor of the walk forward signal in the same way. Then we compare the profit factor found on the permutation to the profit factor we found on real data. To compute our quasi p-value and make a histogram of the permutation's [18:43] profit factors. I ran the test with 200 permutations and got this. The p-value is 22%, roughly meaning there is a 22% chance the walk-forward profit factor of 1.04 could have [18:55] been achieved by a worthless strategy. In other words, a 22% chance our walk-forward results were just dumb luck. Ideally, this probability is very low, and a great strategy will have a very low p-value. I tend to be slightly more lenient with this test. Don't get me wrong, these results are [19:10] not good, but they're only from 2020, just one year. Generally, I'm willing to accept around a 5% p-value on just one year of data, but if we have done the walk-forward permutation test on two or more years of data, then I won't accept a p-value above 1%. I also used just 200 permutations. [19:28] This is because the walk-forward permutation test can take forever to run, even on an extremely simple strategy. This is the test to start before going to sleep. Overall, I don't think these results are good enough, and I would not trade the Donchuan channel with an optimized look [19:42] back going forward. If you want, you can test the strategy in 2021 and beyond to verify it sucks for yourself. I learned about these two permutation tests from the book Permutation and Randomization Tests for Trading System Development by Timothy Masters. The author [19:56] has a PhD in statistics and is the goat of algorithmic trading. This video only covers two of the many tests covered in the book, and the book provides much more detail than I did here. My copy is very beaten up and I've spilled coffee on it several times as I always [20:10] have it on my desk. If you're serious about algorithmic trading, this is a must read. I'll leave a link to it below. The Don Tune channel with an optimized lookback does not fare well against the walk-forward permutation test. In my experience, optimizing lookbacks [20:23] of indicators rarely generalizes well. Rather, when dealing with strategies that require a lookback, I find a stable lookback value, meaning a large variety of lookbacks have decent performance. Then I pick a reasonable look back value and stick with it, then look to improve the strategy [20:38] with the chosen look back. There are many tests and tools to help validate trading strategies, and I only covered two of them in this video. Some tests or tools are more useful for certain types of strategies, but I always use these two tests regardless of the strategy. [20:52] I will not use a trading strategy if it did not have very low p-values for both the in-sample and walk-forward permutation tests. This video has been a rough outline of how I develop and validate a trading strategy. But the process and steps should be applicable to most strategies, and certainly all price-based strategies. [21:09] No process is bulletproof. If you are irresponsible with your development process, no amount of fancy tests can save you. But this is what I do, and I'll keep doing it until I find a better way. Here we compared the profit-back-servant trading strategy between the real price and price permutation. [21:25] You can use permutation tests in many different ways. In my previous video, I compared the percentage of times price bounced off a moving average between the real price and price permutations. You can use permutation tests to help verify or disprove any assumptions or theories you [21:40] may have about the markets.