Genetic Programming Factor Mining — Step-by-Step Guide & Transcript

机器学习与智能金融: 宽基指数择时因子的挖掘

0h 06m video Published May 15, 2026 Transcribed Sep 13, 2026 赚钱逻辑 赚钱逻辑
3 views Recent velocity 0.0 views/hour View full performance history →
AI Trust Score 65/100
⚠️ Average / Some Fluff

"Delivers a solid, technical walkthrough of genetic programming for index timing, though the title likely promises more than the brief overview provides."

AI Summary

This video demonstrates the application of genetic programming (GP) to develop quantitative timing models for the Hushen 300 index. It covers the entire workflow, from initial population setup and fitness function design to selection, crossover, mutation, and final factor pool formation, culminating in impressive backtest results.

[00:30]
GP for Quantitative Analysis

Genetic programming is used to mine models for financial scenarios, significantly reducing time compared to manual mining. The example uses the Hushen 300 index data.

[00:57]
Preparation Steps

Preparation includes using the DEAP genetic algorithm toolbox, selecting terminal datasets, defining backtest rules, and choosing appropriate crossover and mutation methods.

[01:09]
Initial Population Construction

The initial population is formed randomly from operators, constants, and variables. Operator selection is a core issue; operators include custom ones and those from the PayLab package, designed to smooth short and long-term fluctuations and extract trend information.

[02:00]
Constants and Population Size

Constants are chosen from multiples of 5 (5, 10, 15, 20) to prevent overfitting. Population size is critical: too low leads to premature convergence, too high reduces efficiency. For this case, population size is set to 50.

[02:29]
Fitness Function Design

The fitness function is the core of GP. It is multi-dimensional, incorporating annualized return, extreme tail risk, win rate, annual average trade count, and average long holding days.

[03:08]
Fitness Function Formula

The fitness function is: Calmar ratio * sqrt(average annual trade count) * win rate * sqrt(average long holding days). Code is shown.

[03:29]
Selection, Crossover, and Mutation

Custom selection keeps the top 50% by fitness, with additional filters: Calmar ratio > 0.3, annual trades > 2, win rate > 50%, average long holding days > 6. Single-point crossover is used with probability 0.9. Uniform integer mutation is used with probability 0.9.

[04:43]
Double-Layer Screening for Factor Pool

Factors must maintain Calmar ratio > 0.7 in out-of-sample backtests and significantly outperform the benchmark to enter the final factor pool. The process generates many expressions and screens them iteratively.

[05:54]
Combined Model Results

A combination of factors yields excellent timing results on the Hushen 300: annualized return 61.52%, max drawdown 14.79%, average annual trades 26.38, Calmar ratio 2.439, win rate 61.14%, profit/loss ratio 2.941.

Genetic programming effectively discovers timing factors for the Hushen 300 index, achieving strong risk-adjusted returns. The methodology emphasizes careful design of fitness functions and multi-stage screening to ensure robustness.

Mentioned in this Video

Tutorial Checklist

1 00:57 Prepare the environment: use DEAP toolbox, select terminal dataset (Hushen 300), define backtest rules, and choose crossover/mutation methods.
2 01:09 Build the initial population: randomly combine operators, constants, and variables. Set population size to 50.
3 02:29 Define the fitness function: Calmar ratio * sqrt(annual trades) * win rate * sqrt(avg long holding days).
4 03:29 Apply selection: keep top 50% by fitness, then filter by Calmar > 0.3, trades > 2, win rate > 50%, holding days > 6.
5 04:13 Perform crossover (single-point, probability 0.9) and mutation (uniform integer, probability 0.9).
6 04:43 Screen factors: out-of-sample Calmar ratio > 0.7 and outperform benchmark to enter final factor pool.
7 05:54 Combine factors into a composite model and backtest on Hushen 300 index.

💡 Key Takeaways

⚖️

Multi-dimensional fitness function

Emphasizes that fitness must balance multiple performance metrics to avoid overfitting to a single criterion.

02:29
🔧

Auxiliary selection filters

Shows a practical method to ensure solution convergence without over-optimizing one metric.

03:29
🔧

Double-layer screening for robustness

Out-of-sample validation with a Calmar threshold ensures factors generalize beyond training data.

04:43
📊

Impressive backtest performance

Demonstrates the potential of GP-generated factors to achieve high returns with controlled drawdown.

06:06

[00:00] 请不吝点赞 订阅 转发 打赏支持明镜与点点栏目

[00:30] 等金融场景进行量化分析以及预测时, 经常使用短期模型、长期模型以及二者结合形成的综合模型。 使用机器挖掘相比使用人工挖掘,

[00:44] 消耗时间将大大缩短。 本节将使用遗传规划算法挖掘模型, 并利用护身300指数数据进行实际研究。

[00:57] 我们的准备工作包括, 借助遗传算法工具箱DEAP 选取终端数据集 制定回测规则

[01:09] 选取合适的交叉变异方法 遗传规划主要流程包括 首先是构建初始种群 在本案例中

[01:21] 初始化种群由算子、强数 以及变量随机形成 算子的选择是遗传规划的核心问题之一

[01:34] 本案例中的算子包括自定义算子和部分Personal中PayLab包中的算子算子包括自定义算子和部分Person中PayLab包中的算子算子包括了针对长短周期的各种绿波

[01:48] 可以有效平滑长短周期的波动 提取趋势信息 常数 为了防止因子收链值局部间点

[02:00] 降低过拟合程度 公式数常数 将从5的倍数5 10 15 20中进行选择 种群数量过低会造成快速收敛之局部最优解的情况

[02:17] 种群数量过大会降低计算的效率 因此选择合适的初始种群数量对于解决问题至关重要

[02:29] 针对本案例中的所用数据 我们选取种群数量为50 接下来是构建适应度函数 适应度函数的构建是遗传规划的最核心问题

[02:44] 因此需从则是策略最为重视的年化收益率 极端尾部风险 胜率 年平均交易次数

[02:56] 和平均多头持仓天数等方面出发 另一个满足多维度的符合适应度函数 遗传规划生成因子表达式

[03:08] 回归测试结果的卡玛比率乘以年平均交易次数开方 再乘以胜率 再乘以平均多头持仓天数开方代码如下接下来是执行选择交叉和定义

[03:29] 在进化选择方式上 采用自定义的选择最佳模式 除了选择适应度为最佳的前50%以外 还需要从以下辅助条件进行再次筛选

[03:44] Cama比率大于0.3 年平均交易次数大于2 胜率大于50% 平均多头持仓天数大于6 这些辅助条件可以保证解的收敛方向

[03:59] 不会过度趋向某一个或两个条件 在交叉方式上 本节选择单点交叉 交叉概率为0.9 大家可以看下角

[04:13] 在这里regist函数是给tobox类添加新函数 regist函数的第一个参数是要添加的新函数名称

[04:26] 第二个是函数 第三个是新函数的参数 在辨语方式上 本节选取均匀整数突变 辨语概率为0.9

[04:43] 接下来是双层筛选形成最终因子池 被选因子在样本外回测中保持CAM比率大于0.7 而且能大幅跑赢基准

[04:57] 则为通过了样本外检测才能进入最终的因子池使用终端数据和常数以及上文中算子生成大量的表达式

[05:09] 按照依此进行筛选 进化选择 交叉和变异 形成最终的因子池 下表给出了因子池中的部分因子

[05:22] 接下来是结果测试 其中部分因子的回归测试结果如下 这里是关于因子1的测试结果

[05:38] 这是关于因子2的测试结果 接下来我们给出遗传规划择时的一个组合模型

[05:54] 对因子进行简单的组合 得到一个效果 较好的组合因子 最终在互生三白指数 则时获得了年化收益

[06:06] 61.52% 最大回撤 14.79% 年平均交易次数 26.38次 下谱比例 2.439

[06:18] 胜率为61.14% 盈亏比 2.941的优秀则时回溯结果 大家可以看以下表格和图片 好

[06:31] 我们本节就到此结束 谢谢大家

⚡ Saved you 0h 06m reading this? Transcribe any YouTube video for free — no signup needed.