Files
tac-exp-dev/book/data/chat_mining/regulation-signed-diff.txt

134 lines
6.2 KiB
Plaintext

[user] investigate https://tradeac-dev.h.lizhao.net/rd/result?expId=21&run=28c7fa08e7e946b689893171957b997f training curves. the rankIC on training data set keep disperse instead of converge, while valid dataset rankIC wasn't hurt with round progress. the performance wasn't hurt as well. is this a phenomena that we can use? how about add more metrics to evaluate, like MSE and sign accuracy. Or we can do some structural regularization? think loud and see how to deep dive to gain trading edge
[assistant] Let me pull the run details, model info, and results to understand the training dynamics.
[assistant] Let me look at the ensemble model source to understand the architecture better.
[assistant] Here's my deep analysis of this run.
## Training Dynamics Summary
**Model**: `RankICEnsembleLGBModel` (5 seeds, MSE loss, early-stops on RankIC feval)
**Universe**: 50 ETFs, label = 5-day forward return
**Segments**: Train 2016-2025 (9.7y), Valid 2025-09 to 2026-01 (4mo), Test 2026-01 to 2026-08 (7mo)
### The Divergence Pattern
| Metric | Train | Valid | Test |
|---|---|---|---|
| RankIC | **0.318** (seed42, step 360+) | **0.084-0.106** | **0.059** |
| ICIR | - | - | 0.227 |
Training RankIC climbs monotonically from 0.107 to 0.318+ (3x), while validation stays flat at 0.08-0.11. The 0.0 entries in the training curve are from the feval logging alternating with valid set — the non-zero envelope is the actual train RankIC progression.
### Performance
- **Excess return with cost**: 7.77% ann., IR=0.79, max DD=-7.93%
- **Excess return without cost**: 17.02% ann., IR=1.72
- **Cost drag**: 17% → 7.77% — turnover is eating 55% of alpha
- 590 trades over 150 days = ~4 trades/day avg
---
## Analysis: What's Happening
This is **not** harmful overfitting — it's the expected behavior of MSE loss with RankIC early-stopping. Here's why:
1. **MSE optimizes point estimates** — the model learns to make increasingly "confident" (wider-spread) predictions on training data, inflating in-sample RankIC
2. **RankIC early-stopping selects for ordinal quality** — the best iteration maximizes rank correlation on validation, not MSE
3. **The ranking structure generalizes** — validation RankIC doesn't degrade because the model is learning genuine cross-sectional signal, just with overconfident magnitudes in-sample
4. **The train/valid gap (~3x) is a measure of prediction confidence inflation**, not signal degradation
The ensemble seeds show good agreement on validation (0.084-0.11) despite different training paths — this is healthy model stability.
---
## Is This Exploitable? Yes, Multiple Angles
### 1. Confidence Gap as a Diagnostic
The train/valid RankIC ratio (0.318/0.10 ≈ 3.2x) could be monitored as a **regime detector**:
- Ratio < 2x: underfitting, model hasn't converged
- Ratio 2-4x: healthy range
- Ratio > 5x: potential harmful overfitting approaching
### 2. Sign Accuracy — The Metric You Should Add
RankIC measures ordinal correlation but doesn't tell you **direction accuracy**. For a topk=10 strategy, what matters is: "when the model picks a name, does it go up?"
```python
# Sign accuracy: fraction of predictions with correct sign
sign_acc = (np.sign(pred) == np.sign(label)).mean()
# Per-day version
sign_acc_daily = df.groupby(level='datetime').apply(
lambda g: (np.sign(g['pred']) == np.sign(g['label'])).mean()
)
```
This directly translates to trading edge: if sign accuracy > 50% on the top-k names, the strategy makes money.
### 3. Quantile Spread Analysis
Instead of just RankIC, measure the **return spread** between top and bottom quintiles:
- Long-short spread = mean return(top 20%) - mean return(bottom 20%)
- This is what the TopkDropout strategy actually captures
### 4. Rolling IC Stability
The mean RankIC of 0.059 masks the distribution. Compute:
- **IC t-stat**: is it consistently positive or driven by a few big days?
- **IC hit rate**: % of days with positive IC
- **IC half-life**: how quickly does IC decay?
---
## Structural Regularization Approaches
### A. Label Smoothing (Easiest win)
Add noise to training labels to prevent overconfident predictions:
```python
# In the dataset handler or reweighter
label_smoothed = label * (1 - epsilon) + noise * epsilon # epsilon=0.1-0.2
```
This directly attacks the train/valid RankIC gap without changing the model architecture.
### B. Multi-Task Loss: MSE + Rank Regularizer
Add a differentiable rank-aware penalty:
```python
loss = mse_loss + lambda * rank_penalty
# Where rank_penalty penalizes pairwise inversions
```
LightGBM doesn't natively support this, but you could:
1. Use `Huber loss` instead of MSE (reduces outlier sensitivity)
2. Post-process with a rank-aware calibrator
### C. Feature Stochasticity
The current `colsample_bytree=0.8` is good. Additional ideas:
- **Temporal feature dropout**: randomly zero out the most recent N days of features
- **Feature importance ceiling**: cap any single feature's importance at X%
### D. Temperature Scaling (Post-hoc)
Divide predictions by a learned temperature parameter T > 1:
```python
calibrated_pred = pred / T # T learned on validation set
```
This shrinks the prediction spread without changing rankings — exactly what we need.
### E. Stochastic Weight Averaging
Average model weights across the last K boosting rounds instead of using the single best iteration. This smooths the training trajectory.
---
## Recommended Next Steps
1. **Add sign accuracy tracking** to the workflow (easy, high-value)
2. **Add rolling IC + IC hit rate** to understand stability
3. **Run an ablation**: same config with `Huber loss` instead of MSE — does the train/valid gap shrink without hurting validation RankIC?
4. **Try label smoothing**: inject 10% Gaussian noise into labels, retrain, compare
5. **Temperature scaling**: learn T on validation, see if test RankIC improves
6. **The cost drag (55%) is the real problem** — consider:
- Reducing n_drop from 2 to 1 (fewer trades)
- Adding a turnover penalty to the strategy
- Using `notional` orders instead of `qty` to limit position sizes
The training divergence pattern itself is a **feature, not a bug** — it tells you the model has signal capacity. The question is whether you can extract more of it out-of-sample through better regularization and calibration.