CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01msr-spare-1 /qwen3-30b-0617-6skill-regretonly-spare-games-envs qwen3-30B-A3B-Instruct-0617-6skill-regretonly — generated environments Environments generated by the SPARE proposer during training run a1s4s63z (qwen3-30B-A3B-Instruct-0617-6skill-regretonly), recovered from the spare-viz durable cache. The run's scratch directory no longer exists; this dataset is the surviving copy. Games 158 Steps covered 6 (step 0–138) With recovered skill 128 With hint 0 Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0617-6skill-regretonly-spare-games-envs.textn<1K0 likes132 downloads1mo agoHugging Face02tomyimkc /repro-on-regret-bounds-of-thompson-sampling-for-bayesian-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes55 downloads2mo agoHugging Face03tomyimkc /repro-optimal-regret-for-policy-optimization-in-contextual-bandits-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes52 downloads2mo agoHugging Face04tomyimkc /repro-finite-and-corruption-robust-regret-bounds-in-online-inverse-linear-optimization-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes46 downloads2mo agoHugging Face05tomyimkc /repro-dynamic-regret-via-discounted-to-dynamic-reduction-with-applications-to-curved-l-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes41 downloads2mo agoHugging Face06albertklorer /chess-stockfish-regret Chess RLVR Stockfish Regret 1400/100 Snapshot This snapshot dataset stores chess positions for reinforcement learning with verifiable rewards. Each row contains: { "id": "chess_rlvr_000001", "fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1", "legal_moves": "{\"Nf3\": -0.015, \"e4\": 0.0}" } legal_moves is a JSON object encoded as a string. The object maps each legal SAN move to a Stockfish-derived negative regret score for the player to move. The RLVR… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-stockfish-regret.textreinforcement-learning1K<n<10K0 likes39 downloads3mo agoHugging Face07tomyimkc /repro-tighter-regret-lower-bound-for-gaussian-process-bandits-with-squared-exponential-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes39 downloads2mo agoHugging Face08albertklorer /chess-rlvr-stockfish-regret Chess RLVR Stockfish WDL This dataset stores chess positions for reinforcement learning with verifiable rewards. Each row contains: { "id": "chess_rlvr_000001", "fen": "rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1", "legal_moves": { "Nf3": -0.015, "e4": 0.0 } } legal_moves maps each legal SAN move to a Stockfish-derived negative regret score for the player to move. The RLVR reward is negative expected-score regret: reward =… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/chess-rlvr-stockfish-regret.text1K<n<10K0 likes37 downloads3mo agoHugging Face09msr-spare-1 /qwen3-30b-0612-self-hint-regret-mem-spare-games-envs qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem — generated environments Environments generated by the SPARE proposer during training run x5mpt53a (qwen3-30B-A3B-Instruct-0612-self-hint-regret-mem), recovered from the spare-viz durable cache. The run's scratch directory no longer exists; this dataset is the surviving copy. Games 93 Steps covered 4 (step 0–12) With recovered skill 93 With hint 0 Actor / proposer model… See the full description on the dataset page: https://huggingface.co/datasets/msr-spare-1/qwen3-30b-0612-self-hint-regret-mem-spare-games-envs.textn<1K0 likes7 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.