tensorlink-dev/tsfm-toto2-4m-400k-native
TSFM tokenizer ablation — Toto-2-4m-recipe clone (v10)
VQ-VAE OPTIMIZATION ablation (codebook health + next-code objective vs raw patches) on a ~3.6M-param decoder-only patched transformer trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long pool uses the cheap families only) with contiguous patch masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only the CONTEXT tokenizer varies:
Each subfolder is one (arm, seed) run: model.pt (final), ckpt_15000.pt (rank-stability snapshot), config.json, results.json (dev-GIFT CRPS + long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset — NOT the full leaderboard; treat numbers as ablation-internal, not comparable to published GIFT scores. Generated by the v10 experiment notebook.
Results
=== transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) ===
gm@30000 gm@100000 gm@200000 gm@300000 gm@final gm_final_std long_season params_m perp probe_amp n
arm
V0_champion_400k 0.2373 0.205 0.2126 0.2246 0.1991 NaN 0.2693 3.878 552.512 0.307 1
--- tripwire checks ---
