datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qwen3-30b-0617-6skill-regretonly-spare-games-envs
qwen3-30B-A3B-Instruct-0617-6skill-regretonly — generated environments
Environments generated by the SPARE proposer during training run
a1s4s63z (qwen3-30B-A3B-Instruct-0617-6skill-regretonly), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
158
Steps covered
6 (step 0–138)
With recovered skill
128
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0617-6skill-regretonly-spare-games-envs.llm-serving-selector-regret
LLM-Serving Selector Regret
LLM-Serving Selector Regret is a metrics-only research dataset for studying learned policy selection in LLM-serving schedulers. It contains derived selector/oracle/regret objects generated by Soroush Vahidi's research workflow, not raw request traces.
Creator / Provider
Dataset creator/provider: Soroush Vahidi.
The released selector/regret and policy-suitability metrics were generated by Soroush Vahidi's research workflow. Underlying… See the full description on the dataset page: https://huggingface.co/datasets/SoroushVahidi/llm-serving-selector-regret.repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-dynamic-regret-via-discounted-to-dynamic-reduction-with-applications-to-curved-l-traces
Agent traces
Agent sessions published from a Trackio Logbook.
chess-stockfish-regret
Chess RLVR Stockfish Regret 1400/100 Snapshot
This snapshot dataset stores chess positions for reinforcement learning with verifiable rewards.
Each row contains:
{
"id": "chess_rlvr_000001",
"fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1",
"legal_moves": "{\"Nf3\": -0.015, \"e4\": 0.0}"
}
legal_moves is a JSON object encoded as a string. The object maps each legal SAN move to a Stockfish-derived negative regret score for the player to move.
The RLVR… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-stockfish-regret.repro-tighter-regret-lower-bound-for-gaussian-process-bandits-with-squared-exponential-traces
Agent traces
Agent sessions published from a Trackio Logbook.
chess-rlvr-stockfish-regret
Chess RLVR Stockfish WDL
This dataset stores chess positions for reinforcement learning with verifiable rewards.
Each row contains:
{
"id": "chess_rlvr_000001",
"fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1",
"legal_moves": {
"Nf3": -0.015,
"e4": 0.0
}
}
legal_moves maps each legal SAN move to a Stockfish-derived negative regret score for the player to move.
The RLVR reward is negative expected-score regret:
reward =… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-rlvr-stockfish-regret.LIGHT-REGRETSrepro-no-swap-regret-fpa-bundlelora-without-regretsLIGHT-REGRETARegRet-Datasetyoutube_regrets
Dataset Card for Mozilla RegretsReporter Public Data
Dataset Summary
RegretsReporter Data
This data set card describes the public data sets made available based on Mozilla’s RegretsReporter
research as well as the Viu Política research
from the University of Exeter and Vero Instituto.
This data was collected from participants in Mozilla’s RegretsReporter studies.
Participants installed a web extension to participate in each study. In the case of the
first… See the full description on the dataset page: https://huggingface.co/datasets/mozilla-foundation/youtube_regrets.qwen3-30b-0612-self-hint-regret-mem-spare-games-envs
qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem — generated environments
Environments generated by the SPARE proposer during training run
x5mpt53a (qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem), recovered from the spare-viz durable cache.
The run's scratch directory no longer exists; this dataset is the surviving copy.
Games
93
Steps covered
4 (step 0–12)
With recovered skill
93
With hint
0
Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0612-self-hint-regret-mem-spare-games-envs.
