CoolFace
Datasetpublic

LisaMegaWatts/sshfighter-17-head-router-delay-replication-v1

SSH Fighter 17-Head Frozen-Expert Router v1 — Delay Stress Seventeen character-specific heads route among nine immutable Agent Gym policies. The package contains every three-seed checkpoint, averaged inference weights, the grouped training tensor cache, and fresh-seed exact-engine evaluation rows. This package is the predeclared six-scenario role-delay stress matrix. Fresh exact-engine evaluation condition points rate wins losses draws switches/match… See the full description on the dataset page: https://huggingface.co/datasets/LisaMegaWatts/sshfighter-17-head-router-delay-replication-v1.

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes54downloads
Dataset Card

SSH Fighter 17-Head Frozen-Expert Router v1 — Delay Stress

Seventeen character-specific heads route among nine immutable Agent Gym policies. The package contains every three-seed checkpoint, averaged inference weights, the grouped training tensor cache, and fresh-seed exact-engine evaluation rows. This package is the predeclared six-scenario role-delay stress matrix.

Fresh exact-engine evaluation

conditionpoints ratewinslossesdrawsswitches/match
fixed-jumper0.5704517093853333940.00
static0.5936539883645431940.00
linear0.612456732356841220138.14
chebyshev0.607456190360841362123.66
  • Matches: 374,544
  • Engine snapshot: f478c1320f86f428c882e391c11e97ca4cf7b47b
  • Technical failures: 0
  • Evaluation seeds: [1608202801, 1608202819, 1608202847]
  • Passive samples: 8,946,534 with all nine proposal codes, literal HP-kick oscillator state, disagreement, agreement-CUSUM, and momentum labels. These signals have no control authority.

Training scope

  • Train/dev/retrospective-holdout samples: {'dev': 117504, 'holdout': 117504, 'train': 352512}
  • Model seeds: [42, 123, 456]
  • Experts remain frozen; only routing parameters are learned.
  • The exported ensembles carry Python golden logits which the TypeScript evaluator must reproduce before executing a match.

Limitations

Training labels are retrospective contextual-bandit observations, not full counterfactual rewards for all experts. Fresh evaluation uses known synthetic opponent families and direct exact-engine inputs—not SSH transport or human adaptation. This is the pre-balance-intervention f478c13 stratum and must not be pooled with later mechanics fingerprints.