datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-30b-0617-6skill-regretonly-spare-games-envs
qwen3-30B-A3B-Instruct-0617-6skill-regretonly — generated environments
Environments generated by the SPARE proposer during training run
a1s4s63z (qwen3-30B-A3B-Instruct-0617-6skill-regretonly), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
158
Steps covered
6 (step 0–138)
With recovered skill
128
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0617-6skill-regretonly-spare-games-envs.repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-dynamic-regret-via-discounted-to-dynamic-reduction-with-applications-to-curved-l-traces
Agent traces
Agent sessions published from a Trackio Logbook.
chess-stockfish-regret
Chess RLVR Stockfish Regret 1400/100 Snapshot
This snapshot dataset stores chess positions for reinforcement learning with verifiable rewards.
Each row contains:
{
"id": "chess_rlvr_000001",
"fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1",
"legal_moves": "{\"Nf3\": -0.015, \"e4\": 0.0}"
}
legal_moves is a JSON object encoded as a string. The object maps each legal SAN move to a Stockfish-derived negative regret score for the player to move.
The RLVR… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-stockfish-regret.repro-tighter-regret-lower-bound-for-gaussian-process-bandits-with-squared-exponential-traces
Agent traces
Agent sessions published from a Trackio Logbook.
chess-rlvr-stockfish-regret
Chess RLVR Stockfish WDL
This dataset stores chess positions for reinforcement learning with verifiable rewards.
Each row contains:
{
"id": "chess_rlvr_000001",
"fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1",
"legal_moves": {
"Nf3": -0.015,
"e4": 0.0
}
}
legal_moves maps each legal SAN move to a Stockfish-derived negative regret score for the player to move.
The RLVR reward is negative expected-score regret:
reward =… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-rlvr-stockfish-regret.qwen3-30b-0612-self-hint-regret-mem-spare-games-envs
qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem — generated environments
Environments generated by the SPARE proposer during training run
x5mpt53a (qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
93
Steps covered
4 (step 0–12)
With recovered skill
93
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0612-self-hint-regret-mem-spare-games-envs.
