Genetic Algorithm vs. Human Mining: Speed Showdown
45sHighlights the dramatic time savings of AI-driven model mining, sparking curiosity about the technology.
▶ Play Clip"Delivers a solid, technical walkthrough of genetic programming for index timing, though the title likely promises more than the brief overview provides."
This video demonstrates the application of genetic programming (GP) to develop quantitative timing models for the Hushen 300 index. It covers the entire workflow, from initial population setup and fitness function design to selection, crossover, mutation, and final factor pool formation, culminating in impressive backtest results.
Genetic programming is used to mine models for financial scenarios, significantly reducing time compared to manual mining. The example uses the Hushen 300 index data.
Preparation includes using the DEAP genetic algorithm toolbox, selecting terminal datasets, defining backtest rules, and choosing appropriate crossover and mutation methods.
The initial population is formed randomly from operators, constants, and variables. Operator selection is a core issue; operators include custom ones and those from the PayLab package, designed to smooth short and long-term fluctuations and extract trend information.
Constants are chosen from multiples of 5 (5, 10, 15, 20) to prevent overfitting. Population size is critical: too low leads to premature convergence, too high reduces efficiency. For this case, population size is set to 50.
The fitness function is the core of GP. It is multi-dimensional, incorporating annualized return, extreme tail risk, win rate, annual average trade count, and average long holding days.
The fitness function is: Calmar ratio * sqrt(average annual trade count) * win rate * sqrt(average long holding days). Code is shown.
Custom selection keeps the top 50% by fitness, with additional filters: Calmar ratio > 0.3, annual trades > 2, win rate > 50%, average long holding days > 6. Single-point crossover is used with probability 0.9. Uniform integer mutation is used with probability 0.9.
Factors must maintain Calmar ratio > 0.7 in out-of-sample backtests and significantly outperform the benchmark to enter the final factor pool. The process generates many expressions and screens them iteratively.
A combination of factors yields excellent timing results on the Hushen 300: annualized return 61.52%, max drawdown 14.79%, average annual trades 26.38, Calmar ratio 2.439, win rate 61.14%, profit/loss ratio 2.941.
Genetic programming effectively discovers timing factors for the Hushen 300 index, achieving strong risk-adjusted returns. The methodology emphasizes careful design of fitness functions and multi-stage screening to ensure robustness.
Multi-dimensional fitness function
Emphasizes that fitness must balance multiple performance metrics to avoid overfitting to a single criterion.
02:29Auxiliary selection filters
Shows a practical method to ensure solution convergence without over-optimizing one metric.
03:29Double-layer screening for robustness
Out-of-sample validation with a Calmar threshold ensures factors generalize beyond training data.
04:43Impressive backtest performance
Demonstrates the potential of GP-generated factors to achieve high returns with controlled drawdown.
06:06[00:00] 请不吝点赞 订阅 转发 打赏支持明镜与点点栏目
[00:30] 等金融场景进行量化分析以及预测时, 经常使用短期模型、长期模型以及二者结合形成的综合模型。 使用机器挖掘相比使用人工挖掘,
[00:44] 消耗时间将大大缩短。 本节将使用遗传规划算法挖掘模型, 并利用护身300指数数据进行实际研究。
[00:57] 我们的准备工作包括, 借助遗传算法工具箱DEAP 选取终端数据集 制定回测规则
[01:09] 选取合适的交叉变异方法 遗传规划主要流程包括 首先是构建初始种群 在本案例中
[01:21] 初始化种群由算子、强数 以及变量随机形成 算子的选择是遗传规划的核心问题之一
[01:34] 本案例中的算子包括自定义算子和部分Personal中PayLab包中的算子算子包括自定义算子和部分Person中PayLab包中的算子算子包括了针对长短周期的各种绿波
[01:48] 可以有效平滑长短周期的波动 提取趋势信息 常数 为了防止因子收链值局部间点
[02:00] 降低过拟合程度 公式数常数 将从5的倍数5 10 15 20中进行选择 种群数量过低会造成快速收敛之局部最优解的情况
[02:17] 种群数量过大会降低计算的效率 因此选择合适的初始种群数量对于解决问题至关重要
[02:29] 针对本案例中的所用数据 我们选取种群数量为50 接下来是构建适应度函数 适应度函数的构建是遗传规划的最核心问题
[02:44] 因此需从则是策略最为重视的年化收益率 极端尾部风险 胜率 年平均交易次数
[02:56] 和平均多头持仓天数等方面出发 另一个满足多维度的符合适应度函数 遗传规划生成因子表达式
[03:08] 回归测试结果的卡玛比率乘以年平均交易次数开方 再乘以胜率 再乘以平均多头持仓天数开方代码如下接下来是执行选择交叉和定义
[03:29] 在进化选择方式上 采用自定义的选择最佳模式 除了选择适应度为最佳的前50%以外 还需要从以下辅助条件进行再次筛选
[03:44] Cama比率大于0.3 年平均交易次数大于2 胜率大于50% 平均多头持仓天数大于6 这些辅助条件可以保证解的收敛方向
[03:59] 不会过度趋向某一个或两个条件 在交叉方式上 本节选择单点交叉 交叉概率为0.9 大家可以看下角
[04:13] 在这里regist函数是给tobox类添加新函数 regist函数的第一个参数是要添加的新函数名称
[04:26] 第二个是函数 第三个是新函数的参数 在辨语方式上 本节选取均匀整数突变 辨语概率为0.9
[04:43] 接下来是双层筛选形成最终因子池 被选因子在样本外回测中保持CAM比率大于0.7 而且能大幅跑赢基准
[04:57] 则为通过了样本外检测才能进入最终的因子池使用终端数据和常数以及上文中算子生成大量的表达式
[05:09] 按照依此进行筛选 进化选择 交叉和变异 形成最终的因子池 下表给出了因子池中的部分因子
[05:22] 接下来是结果测试 其中部分因子的回归测试结果如下 这里是关于因子1的测试结果
[05:38] 这是关于因子2的测试结果 接下来我们给出遗传规划择时的一个组合模型
[05:54] 对因子进行简单的组合 得到一个效果 较好的组合因子 最终在互生三白指数 则时获得了年化收益
[06:06] 61.52% 最大回撤 14.79% 年平均交易次数 26.38次 下谱比例 2.439
[06:18] 胜率为61.14% 盈亏比 2.941的优秀则时回溯结果 大家可以看以下表格和图片 好
[06:31] 我们本节就到此结束 谢谢大家
⚡ Saved you 0h 06m reading this? Transcribe any YouTube video for free — no signup needed.