[user] use /skills to get configuration of experiment id 11 or name `tac-rd-rank-ablate` [assistant] Experiment 11 is `tac-rd-rank-ablate` with 2 runs. Fetching its config: [assistant] Now fetching the saved `config` artifacts (canonical input record) for both runs: [assistant] Experiment **11 `tac-rd-rank-ablate`** — 2 runs, both share identical config except `feature_fields`. ## Shared config - **Model**: `RankICLGBModel` (`tac_qlib.contrib.model.rank_gbdt`), loss=mse, lr=0.02, num_leaves=31, num_boost_round=3000, early_stop=200, min_data_in_leaf=20, lambda_l2=0.5, colsample_bytree=0.8, subsample=0.8, reg_alpha=0.1, reg_lambda=1.0, seed=42 - **Dataset**: `DatasetH`/`TACHandler`, freq=day, lake_root=/home/data/lake, market=US, 50 ETFs (SPY,QQQ,TLT,XL*,ARKK,SMH,...REM) - label `Ref($close,-6)/Ref($close,-1)-1` - processors: DropAllNaN → ProcessInf → CSRankNorm → ZScoreNorm → Fillna (fit 2015-01-03..2025-09-01) - segments: train `2015-01-03..2025-09-01`, valid `2025-09-03..2026-01-03`, test `2026-01-04..2026-08-10` - **Records**: SignalRecord, SigAnaRecord (ana_long_short, ann 252), PortAnaRecord — TopkDropoutStrategy (topk=10, n_drop=2, risk_degree 0.95), backtest benchmark=SPY, $1M, deal_price=$close, costs 0.05% open / 0.15% close / $5 min ## Runs | | Run A (baseline, all 24) | Run B (generic-only, 19) | |---|---|---| | run_id | `5cf2c249…90f596` | `7b1e7972…ced74f` | | feature_fields | all sp_* incl. **ou** (zscore, half_life, revert) + **hmm** (p_regime1, state) | drop ou + hmm; keep jump/har/trend/hurst/signature + ret, max_move | | RankIC / RankICIR | 0.0301 / 0.1457 | **0.0635 / 0.2760** | | IC / ICIR | -0.0153 / -0.0662 | 0.0217 / 0.0849 | | L-S ann ret / Sharpe | -0.133 / -0.83 | 0.428 / 2.55 | | net-of-cost excess | -9.4% (IR -1.22) | +3.1% (IR 0.28) | Notes: A = baseline reference; B = hypothesis that generic families dominate and ou/hmm add noise — confirmed, B beats A on every metric. [user] Is rankICIR 0.2760 consider good in this context [assistant] No — 0.276 is modest at best in this context. Per the skill's guidance, |ICIR| ≥ 1 is "decent", ≥ 2 "strong"; 0.276 sits well below the bar. Rank ICIR is the Spearman analog of ICIR (mean RankIC / std × √252). Mitigating factors here: - Short test window (2026-01-04..08-10, ~7 months) and a small 50-ETF daily panel → few cross-sections, so noise is high and the metric is fragile. - RankIC 0.0635 is a decent signal *level* (the daily mean), but the low RankICIR reflects high day-to-day variance in that signal's strength. - The economic evidence is consistent: net-of-cost excess return +3.1% ann with IR 0.28 — real but weak, and it flips to +12.4% (IR 1.11) without costs. So the ablation is a win relative to baseline (0.1457 → 0.2760), but 0.276 alone would not clear a typical go-live bar (≥1) without more evidence. [user] how to improve [assistant] RankICIR is low because RankIC's day-to-day std is high relative to its mean. Improve the *consistency* of the ranking, not just its level. Concrete levers, in order of expected impact: **1. Real training data (biggest lever, cheapest)** The lake bars only start `2021-08-02`, so your 2015 train start was truncated to ~4y. Backfill full history (`tac-engine` lake tools) → more cross-sections → stabler RankIC and less valid-period oscillation (you saw valid RankIC bounce 0.045–0.06). **2. More/richer features** You use only bars + `sp_*` model-free families. For a 50-ETF daily panel the alpha is mostly cross-sectional and inter-asset: - TA: momentum (roc/mom), RSI, MACD, bollinger position, ATR, adx via `get_lake_ta` (covers all of these) — persist alongside `sp_*`. - Inter-asset/relative: term-structure (TLT vs IEF, TLT slope), credit (HYG–LQD spread), sector momentum, gold/USD, rolling 22d cross-sectional rank-momentum of each ETF vs the panel. - Risk-adjusted: returns / ATR, drawdown from 60d high, vol regime. **3. Stabilize the model** (directly targets RankICIR variance) - Multi-seed ensemble: average 5–10 seeds of RankICLGBModel → large std reduction. - Walk-forward/purged CV for hyperparameter selection instead of single train/valid split (your early-stop picked a noisy point). - The 3000 rounds @ lr 0.02 is heavy; consider lr 0.05 + 800 rounds or stricter `min_data_in_leaf`/`lambda_l2` to reduce overfit (train RankIC ~0.45 vs valid ~0.05 is a huge gap). **4. Label / horizon** Test 1d, 10d, 22d labels and blend signals. RankIC at 5d may be dominated by short-horizon noise; a smoothed multi-horizon signal usually has stabler daily IC. **5. Strategy side (doesn't raise RankICIR but raises realized value)** Net-of-cost excess is only +3.1% (IR 0.28) — costs eat the edge. Reduce turnover (larger topk, longer holding, `n_drop` tuning) or lower cost via IEX/day orders. A 0.0635 RankIC that doesn't survive costs isn't monetizable yet. Want me to backfill the lake to 2015 and/or run a feature-expansion (TA + inter-asset) ablation in a new experiment? [user] go ahead to prepare data and features, and then use qlib workflow config to trigger and record the experiment you suggested [assistant] This is a multi-step task. Let me set up a plan and check the current state first. [assistant] Current state: all 72 symbols have 1d bars `2021-08-02..2026-08-12` (IEX), features exist but the ablate run's `sp_*` came from an earlier prep. I need to (a) backfill earlier history via SIP, (b) recompute/persist SP + TA features over the full range, (c) add inter-asset/relative features, (d) wire an ensemble workflow. Let me inspect the existing code structures first. [assistant] State confirmed: lake has all 72 symbols at 1d `2021-08-02..2026-08-12` (IEX), and `rank_gbdt.py` in repo == installed copy. Now let me backfill earlier history. SIP reaches back to 2016-01-04, so I'll extend the calendar and bars for the 50-ETF universe. [assistant] Calendar seeded to 2016. Now backfill bars for the full 50-ETF universe back to 2016-01-04 (SIP), extending the existing IEX-loaded window. [assistant] The lazy backfill timed out at the MCP layer. Let me check whether it actually completed in the background by inspecting coverage. [assistant] The backfill ran in the background despite the timeout — SPY/TLT now cover `2016-01-04..2026-08-13` (2668 bars). Let me verify all 50 ETFs got extended. [assistant] Partial backfill — the timeout cut off after the first 16 symbols. The remaining 34 ETFs are still at 2021. Let me backfill them in smaller batches to stay under the MCP timeout. [assistant] The batch returned data (persisted). Let me check remaining coverage gaps rather than trust the truncated output. [assistant] 32 of 50 done. Backfilling the remaining 18 in two batches. [assistant] That batch succeeded (persisted). Let me confirm remaining gaps and finish the last batch. [assistant] Both batches persisted. Let me verify all 50 now extend back to 2016. [assistant] All 50 ETFs now cover `2016-01-04..2026-08-13`. Data prep done. Now recompute SP features over the full range (fit_end = train cutoff 2025-09-01, no lookahead) and add TA features. [assistant] Now computing SP features over the full backfilled range (fit_end = 2025-09-01 train cutoff for no-lookahead HMM). Let me first inspect the handler/provider and an existing feature file to confirm column routing, then compute. [assistant] Features file currently holds TA + sp_* from the old (2021+) window. I need to recompute over the backfilled 2016+ range. Starting with SP features for all 50 symbols (winning generic families: jump,har,trend,hurst,signature — matching ablation B, fit_end=2025-09-01 for no-lookahead). [assistant] SP features for SPY are computed+persisted over 2016–2026. The remaining 49 symbols would flood context, so I'll delegate that to a subagent while I build the inter-asset feature script and the ensemble workflow in parallel. [assistant] All 49 symbols' SP features persisted (2016→2026). Now I'll build the two custom pieces: (1) inter-asset/relative features script, (2) the multi-seed ensemble model. Let me check the LGBModel init to subclass correctly. [assistant] All 50 symbols' SP features now cover 2016–2026. Let me verify the feature file state after the SP recompute, then compute a curated TA subset. [assistant] Let me check whether TA columns are populated in the 2016–2021 range (they may only cover the old 2021+ window after the SP merge). [assistant] All 50 done. TA columns only cover 2021+ (rsi_14 has 0 pre-2021 rows). Let me test whether a `get_lake_ta` persist merges (keeps sp_*) or drops columns, using one symbol first. [assistant] TA persist ran for SPY (2016+). I need to confirm it merged (kept `sp_*`) rather than overwrote. Let me verify and then delegate TA computation for the remaining 49 symbols. [assistant] SP columns survived the TA merge and TA now covers pre-2021 too. The persist merges columns. Now delegate TA for the other 49 symbols. [assistant] SP + TA features now cover 2016–2026 for all 50 ETFs. Next: compute inter-asset/relative features (cross-sectional rank momentum, betas/correlations to market factors, dispersion) and persist them per-symbol. [assistant] The append logic is buggy (`append_column` with a Table). Let me fix it to append each column individually. [assistant] Inter-asset features persisted for all 50. Let me verify a sample and confirm column presence across files. [assistant] All 50 experiment ETFs now have SP + TA + xr inter-asset features over 2016–2026 (the 21 "missing" are non-universe symbols — fine). Data/feature prep is complete. Now the ensemble model. I'll create a multi-seed `RankICEnsembleLGBModel` that averages predictions across seeds (stabilizes RankICIR), subclassing `RankICLGBModel`. [assistant] Now the multi-seed ensemble model. Creating `RankICEnsembleLGBModel` in the repo and syncing to the installed venv copy. [assistant] Now sync to the installed venv copy (per the skill's "installed package copy" gotcha): [assistant] Model imports cleanly. Now build the workflow YAML. I'll place it in the repo workflows dir and check the reference rankic workflow for exact structure, then write the ensemble workflow with the new feature set. [assistant] Now run the workflow. Note the skill says the MCP call may time out for long runs but execution continues — I'll poll after. [assistant] As expected — the MCP call timed out but the run continues. Let me poll for the new experiment. [assistant] Experiment 15 `tac-rd-rank-ensemble` created, run `b981620825c848d282555b5f0cb87b48` is RUNNING (5×3000-round LightGBM on 9.5y of data will take a while). Let me wait and poll for completion. [assistant] Still training (5 boosters × 3000 rounds on 9.5y). Let me wait more and poll again. [assistant] Still running. Let me check elapsed time and keep polling. [assistant] ~9 min elapsed, still training (5 seeds × 3000 rounds). Let me wait longer. [assistant] Still training. The config artifact appears only when the run finishes. Let me continue polling. [assistant] Let me check the training process is actually alive (CPU) while waiting. [assistant] Two rd_server processes alive — training is running. Let me continue waiting. [assistant] Artifacts are appearing — records are being written. Let me check the run status. [assistant] Run finished. Now let me evaluate the result and compare against the ablate baseline. First, the full result details and input config. [assistant] Run complete. Recording hypothesis + evaluation notes on the run: [assistant] The experiment is recorded. Let me update the todos. [assistant] Done. Full pipeline executed and recorded as **experiment 15 `tac-rd-rank-ensemble`** (run `b9816208…`). ## Data & feature prep - **Backfilled** 1d bars for all 50 ETFs `2021-08-02 → 2016-01-04` (SIP) + calendar seeded to 2016; train now has ~9.5y of real cross-sections. - **Persisted features** over 2016–2026: `sp_*` generic set (jump/har/trend/hurst/signature), a curated TA subset (rsi_14, roc_10, macd_hist, willr_14, atr_14), and **new inter-asset features** `xr_csrank_mom20/60`, `xr_disp_20`, `xr_beta_spy_60`, `xr_corr_tlt_60`, `xr_corr_hyg_60` (via `build_interasset.py`). - **New model** `RankICEnsembleLGBModel` (`tac_qlib/contrib/model/rank_ensemble.py`, synced to venv) — 5 seeds, each RankIC early-stopped, predictions averaged. ## Result — regression vs the ablate-B baseline | | ablate-B (exp 11) | ensemble run (exp 15) | |---|---|---| | RankIC / RankICIR | 0.0635 / **0.2760** | 0.0474 / 0.1470 | | IC / ICIR | 0.0217 / 0.0849 | 0.0306 / 0.1176 | | Long-Short ann (Sharpe) | 0.428 (2.55) | 0.236 (1.24) | | net-of-cost excess | +3.1% | **−24.1%** | IC improved but RankIC fell and the backtest turned deeply negative. Notes recorded on the run. The likely culprit is the **TA subset** — the repo's own SP+TA ablation showed TA dilutes CSRankNorm+LGBM. The ensemble + full-history effects are confounded by the feature change. Want me to run an **isolation experiment** — exact ablate-B feature set (generic sp_* only) with the 5-seed ensemble on the full 2016+ history — to measure the ensemble alone? [user] record git lineage, and then run a isolation experiment with qlib workflow yml [assistant] I'll set up the git lineage tracking via the skill's trace tooling, then run the isolation experiment. [assistant] Env vars are set, experiments submodule exists. Let me look at the trace tooling to understand how to run it and check the submodule git state. [assistant] Let me init the trace tooling and check the current experiment state in the DB. [assistant] Trace tooling is ready and the DB has the lineage chain (id 8 baseline → 9 ablate → 11 moments). Now let me start the git-lineage record for the isolation experiment, evolving from the ablate winner (id 9). [assistant] The exp-9 branch isn't resolvable locally. Let me check what branches exist in the experiments clone. [assistant] The fork failed because `exp/9-...` only exists as `origin/exp/9-...`. Let me check whether row 12 was already inserted, then work around the fork. [assistant] Row 12 created but branch fork failed (base only exists as `origin/exp/9-...`). Let me fork manually from the origin ref and update the DB row's branch. [assistant] Branch forked from exp-9's branch. Now update row 12's git_branch and seed-commit it, replicating what `trace.sh start` would have done. [assistant] Branch `exp/12-...` created and pushed, row 12 recorded. Now write the isolation workflow YAML inside the experiments submodule and commit it. [assistant] Let me check what the ablate-B workflow yaml used, to replicate the exact feature set. Let me look at the exp-9 branch content. [assistant] The isolation config: ablate-B feature set (generic sp_* only) + full 2016 history + 5-seed ensemble — differing from exp-15 only by dropping TA/xr features. Writing it into the experiments submodule. [assistant] Committed. Now run the isolation experiment via `rd_run_workflow`. [assistant] Expected — the run continues in the background. Polling for the new experiment. [assistant] Experiment 16 `tac-rd-rank-ensemble-isolated` running (`0cea66d9...`). Training takes ~18 min (5×3000 rounds). Let me wait and poll. [assistant] Still training. Let me continue polling at intervals. [assistant] ~9 min elapsed. Let me check the run status directly. [assistant] Config is correct (ablate-B features, no TA/xr). Still training. Continue waiting. [assistant] ~17 min elapsed, close to exp-15's ~18 min training time. Checking if it finished. [assistant] ~20 min elapsed. Let me verify the training process is alive and keep polling. [assistant] ~25 min. The exp-15 run finished at ~18 min, but this one may take longer. Let me check the run status again. [assistant] Let me check whether the training is actually consuming CPU (not stuck). [assistant] PID 4328 has ~3500s of CPU time — actively training (LightGBM multithreaded). Still running at ~33 min. Continuing to wait. [assistant] Artifacts count jumped to 11 — records are being written. The run should be near completion. Let me get the final status. [assistant] Run finished — this is the best result so far. Recording the trace finish. [assistant] Isolation run is FINISHED and it's the best net-of-cost result so far. Let me record the trace finish with evaluation + metrics. [assistant] Experiment 12 finished and pushed. The guardrail warning shows only pre-existing untracked files in the parent repo (`.venv`, `tac-app/`, etc. — these were already untracked, not modified by this experiment; my changes were confined to `experiments/`). Let me verify the branch state. [assistant] The intermediate workflow-YAML commit got reset by `finish`'s `fork` (which does `reset --hard origin/`). Let me check whether the YAML survives on disk and re-push it. [assistant] The `finish` step's `reset --hard` wiped my intermediate workflow-YAML commit (the YAML file is gone from disk too). Let me recreate it and push it to the branch. [assistant] The workflow YAML and finish record are now committed and pushed on `exp/12-...`. Let me update todos and check the git lineage / DB final state. [assistant] Done. Git lineage recorded and isolation experiment run. ## Git lineage (experiment 12) - Row inserted in `rd_experiments` (id **12**, `tac-rd-rank-ensemble-isolated`), **evolved_from = 9** (ablate-B winner). - Branch `exp/12-isolate-the-multiseed-rankic-ensemble-ef` forked from `exp/9-...`, workflow YAML committed and pushed, and `trace.sh finish` recorded evaluation + metrics (mlflow ref `0cea66d9...`). - Note: `trace.sh start`'s fork failed (base branch only exists as `origin/exp/9-...`), so I forked manually and updated the row — lineage is intact. - Guardrail surfaced only pre-existing untracked parent files (`.venv`, `tac-app/`, etc.) — no experiment side effects. ## Isolation experiment (mlflow exp 16, full 2016+ history) Same as ablate-B features (generic `sp_*` only) + 5-seed `RankICEnsembleLGBModel`, dropping the TA/xr that caused the exp-15 regression. | Metric | ablate-B (exp 11) | exp-15 ensemble+TA/xr | **exp-16 isolation** | |---|---|---|---| | RankIC / RankICIR | 0.0635 / 0.276 | 0.0474 / 0.147 | 0.0586 / 0.224 | | IC / ICIR | 0.0217 / 0.0849 | 0.0306 / 0.118 | 0.0354 / 0.150 | | Long-Short ann (Sharpe) | 0.428 (2.55) | 0.236 (1.24) | 0.603 (3.71) | | net-of-cost excess | +3.1% (IR 0.28) | −24.1% | **+7.8% (IR 0.79)** | **Conclusion**: exp-15's regression was the TA/inter-asset features, not the ensemble. Dropping them + full history gives the best net-of-cost result so far (+7.8%, IR 0.79, L-S Sharpe 3.71) — though RankICIR (0.224) still trails ablate-B's 0.276, the economic metrics are the strongest observed. [user] what attribute to the slight lower IC/ICIR variants and higher return? [assistant] Quick correction on the direction: **IC/ICIR went up** in the isolation run (0.0217→0.0354, 0.085→0.150); it's **RankIC/RankICIR that dipped** (0.0635→0.0586, 0.276→0.224) while returns rose. The two runs differ in two ways at once — full 2016+ history *and* 5-seed ensemble vs single-seed on 2021+ — so the attribution is inferred, not cleanly measured. **Why RankIC slightly lower but L-S return/Sharpe higher:** 1. **Ensemble averaging narrows day-to-day variance.** Averaging 5 seeds removes per-day outlier scores, so the *top-k* basket (what the backtest holds) is consistently good. That directly raises ICIR (0.150 vs 0.085) and L-S Sharpe (3.71 vs 2.55). But averaging also truncates extreme scores, and RankIC is computed over *all 50 names* — most of which sit in the noisy middle of the distribution where the signal is weakest. A slightly more conservative score distribution costs a little broad RankIC while improving top-decile precision. 2. **Full history de-overfits the recent regime.** ablate-B trained on only ~4y (2021+), with a valid window immediately before test — early-stopping picked an iteration flattered by the recent regime (hence its 0.276 RankICIR). The 9.5y-trained model generalizes better: slightly lower point RankIC but stabler across the 7-month test, converting to higher net-of-cost return (+7.8% vs +3.1%). The truncated model's high RankICIR was partly regime-overfit. 3. **Backtest only cares about the tail.** L-S ann +0.603 (vs +0.428) is driven by top/bottom decile separation, which improved even though full-panel rank correlation dipped — consistent with "better top-k, slightly diluted mid-rank". To attribute precisely, the clean follow-ups are: (a) single-seed RankICLGBModel on full 2016+ history (isolate history effect), or (b) 5-seed ensemble on the old 2021+ window (isolate ensemble effect). Want me to run one? [user] rerun tac-rd-rank-ensemble-isolated with OptimalStopControl strategy [assistant] I'll rerun the isolation experiment with `OptimalStopControl` strategy. Let me check the strategy's config surface and the existing optstop workflow. [assistant] Strategy understood. Let me verify the strategy module is in the venv copy, then set up the workflow + git lineage. [assistant] Strategy is in the venv. Now set up the todo list and start the git lineage. [assistant] Same fork issue — base branch only exists as `origin/exp/12-...`. Let me check if row 13 was inserted, then fork manually. [assistant] Row 13 created. Forking the branch from origin/exp/12 and updating the DB, as before. [assistant] Branch forked and pushed. Now write the optstop isolation workflow YAML inside the experiments submodule. [assistant] Now run the optstop isolation experiment via `rd_run_workflow`. [assistant] The installed copy may have been lost. Let me check and re-sync it. [assistant] The rd_server restarted and the venv copy of `rank_ensemble.py` was wiped. Re-syncing it. [assistant] The repo file is gone too — the environment appears to have reset. Let me check what survived. [assistant] The environment reset wiped the model file and my ensemble workflow YAML. Let me check what survived in the experiments submodule and lake. [assistant] The experiments submodule YAMLs survived (committed), but the `rank_ensemble.py` model was wiped from repo + venv. Let me recreate it and re-sync. First check the trace DB state. [assistant] The optstop run failed because the model module was wiped. Let me recreate `rank_ensemble.py` in the repo and sync to the venv. [assistant] Model restored. The failed run `6c89efd...` in exp 17 is a dead attempt; I'll retry the workflow run (fresh run in exp 17). Re-running: [assistant] Timed out as expected but the run started. Let me verify a new RUNNING run exists in exp 17. [assistant] New run `4e1f77b4...` is RUNNING. Training takes ~40 min (OptimalStopControl backtest runs too). Let me wait and poll. [assistant] ~12 min elapsed (5×3000-round training). Continuing to poll. [assistant] ~24 min. Let me check the process is alive and the run status. [assistant] PID 122 has ~1976s CPU — training actively. Continue waiting. [assistant] ~33 min. Let me check the run status directly. [assistant] Still training (~33 min). The exp-16 isolation took ~40 min including records. Continuing to wait. [assistant] Artifacts count jumped to 11 — records writing, near completion. Checking final status. [assistant] Finished. Let me get the full results. [assistant] Run finished. The signal metrics are identical (same ensemble predictions); the strategy change is what's isolated. OptimalStopControl underperformed TopkDropout. Recording the trace finish. [assistant] The finish wiped the workflow YAML commit again (same `reset --hard` behavior). Let me verify and re-push it to the branch. [assistant] The optstop YAML was wiped from disk and the branch by the `finish` reset. Recreating it and pushing. [assistant] Workflow YAML re-committed and pushed on `exp/13-...`. Let me verify the final DB record and branch state. [assistant] Done. **Experiment 13 `tac-rd-rank-ensemble-optstop`** (mlflow exp 17, run `4e1f77b4...`) — traced on branch `exp/13-portfolioconstruction-variant-of-the-iso`, evolved from exp 12, finished + pushed. ## Result: OptimalStopControl underperforms TopkDropout Same ensemble signal (IC/RankIC identical to exp-16 — clean strategy isolation): | Backtest (test 2026-01-04..08-10) | exp-16 TopkDropout | **exp-13 OptimalStopControl** | |---|---|---| | net-of-cost excess | **+7.8%** (IR 0.79) | −2.7% (IR −0.31) | | without-cost excess | +17.0% (IR 1.72) | +8.6% (IR 0.95) | | max drawdown (with cost) | −7.9% | −11.0% | | cost drag | ~9.2pp | **−11.3pp** | **Why it's worse**: OptimalStopControl's percentile-gated exits (`exit_pct 0.7` + 10-day time stop) churn positions on this signal — the 5d-forward signal re-ranks fast, so exits trigger frequently and bleed the 0.15% close cost (11.3pp drag vs 9.2pp). The stop-loss/entry-gating also concentrated risk into fewer, longer-held names → deeper MDD. TopkDropout topk=10 n_drop=2 remains the better construction. Notes were recorded via `trace.sh finish`; the workflow YAML was re-committed to the branch after the finish-step reset (same behavior as exp-12). One environment note: the container reset mid-task wiped `rank_ensemble.py` from the repo+venv; I recreated it and the second `rd_run_workflow` attempt succeeded (the first attempt's run `6c89efd...` shows FAILED — dead attempt, ignore it). [user] where is RankICEnsembleLGBModel saved [assistant] `RankICEnsembleLGBModel` lives in `tac_qlib/contrib/model/rank_ensemble.py`, saved in two places: - Repo source: `/app/tac-qlib/tac_qlib/contrib/model/rank_ensemble.py` - Installed venv copy (what the MCP `rd_server` imports): `/app/.venv/lib/python3.12/site-packages/tac_qlib/contrib/model/rank_ensemble.py` They must stay in sync — the env reset wiped both earlier and I recreated them. Also note a trained ensemble's boosters are pickled per-run at `/mlruns///artifacts/params.pkl` (e.g. `/home/data/lake/mlruns/16/0cea66d9.../artifacts/params.pkl`, 4.4MB for the 5-seed model).