CoolFace
Modelpublic

tensorlink-dev/tsfm-toto2-4m-learned-tokenizer-tempopfn

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes
Model Card

TSFM tokenizer ablation — Toto-2-4m-recipe clone (v10)

LEARNED-tokenizer ablation (discrete bottlenecks vs raw patches at equal token count) on a ~3.6M-param decoder-only patched transformer trained on the OFFICIAL TempoPFN synthetic prior generators (cloned from automl/TempoPFN; GP/KernelSynth families capped at 6,144 steps, so the long pool uses the cheap families only) with contiguous patch masking, a 9-level quantile head, NorMuon+AdamW, and index RoPE at time-scaled positions. Horizon decoding is fixed patch-32 in every arm; only the CONTEXT tokenizer varies:

armtokenizerhistoryctx tokens
T0fixed-32 (control)4,096128
T1pyramid, iso-context4,09644
T2pyramid, iso-token16,384128
T3adaptive equal-surprise16,384128

Each subfolder is one (arm, seed) run: model.pt (final), ckpt_15000.pt (rank-stability snapshot), config.json, results.json (dev-GIFT CRPS + long-horizon probe). Dev metrics use a fixed 14-task GIFT-Eval subset — NOT the full leaderboard; treat numbers as ablation-internal, not comparable to published GIFT scores. Generated by the v10 experiment notebook.

Results

=== transfer curve: GM-CRPS by checkpoint (dev-14, mean over seeds) ===
               gm@7500  gm@15000  gm@22500  gm@final  gm_final_std  long_season  params_m      perp  probe_amp  n
arm                                                                                                              
L2_vq_tpfn      0.2219    0.2324    0.2314    0.2089        0.0032       0.2594    3.7130  422.0938     0.2360  3
L0_fixed_tpfn   0.2166    0.2258    0.2101    0.2106        0.0085       0.2632    3.5802       NaN     0.1910  3
L1_fsq_tpfn     0.2400    0.2494    0.2500    0.2333        0.0021       0.2733    3.6827    5.4675     0.3427  3

--- tripwire checks ---
  L2_vq_tpfn vs L0_fixed_tpfn: final gap=0.0017 -> UNRESOLVED (< 2*sigma)
  L0_fixed_tpfn vs L1_fsq_tpfn: final gap=0.0226 -> RESOLVED
  L0_fixed_tpfn: TRANSFER INVERSION — best at ckpt #3 (0.2101) vs final (0.2106); synthetic overfit persists