datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rhan-checkpointsfinetuning-checkpointsofficeqa-checkpoint-eval-data
Checkpoint evaluation plot data
Snapshot: 2026-09-14T16:26:45.684890+00:00. Aggregate inputs to notes/Sept-2-2026.md performance figures.
No model execution, grading, publication, or source-result changes were performed to make this export.
Contents
checkpoint_evaluations: 454 checkpoint rows, one evaluation per run/iteration/protocol; score, mean output tokens, mean steps, and the existing two-sided 95% confidence bounds.
pareto_points: current mean-token/USD… See the full description on the dataset page: https://huggingface.co/datasets/YWZBrandon/officeqa-checkpoint-eval-data.rec_rl_checkpoints
rec_rl_checkpoints
Checkpoints for Semantic-ID reasoning generative recommendation, trained with
the SIDReasoner three-stage recipe:
Stage-1 SFT — direct Semantic-ID (SID) prediction, no reasoning.
Stage-2 reasoning activation — cold-start "think" format activation.
Stage-3 GRPO RL — reinforcement learning over SID reasoning traces.
Base model: Qwen3-1.7B. Domains: Office_Products, Video_Games,
Industrial_and_Scientific (Amazon Reviews).
Layout
reproduced/… See the full description on the dataset page: https://huggingface.co/datasets/yufan/rec_rl_checkpoints.rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.translation-checkpointsvpswap-checkpoint-scores
VP-Swap checkpoint scores
Per-item correctness on the VP-Swap benchmark for nine models at twenty
points in training. This is the raw material behind Figures 4-6 of
Augustinian BabyLM (paper, code).
Layout
<model>/<revision>.jsonl, one line per benchmark item:
{"property": "color", "line": 6, "which": 1,
"pll_orig": -21.4213, "pll_swap": -23.9077, "correct": true}
property and line identify the source line in
eval/vpswap_bb24/vp_swap_<property>_pairs.txt; which… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/vpswap-checkpoint-scores.checkpoints-Vllm-agentic-swiss-legal-checkpoints
LLM Agentic Legal Information Retrieval — Checkpoints
Public artifacts from competing in the Kaggle competition.
Best public LB
Submission
LB
v6 LightGBM baseline
0.0709
v9 DeepSeek paragraph injection
0.13167
v11 court sibling expansion
0.13665
v12 Qwen2.5-14B LoRA
0.13204
Structure
submissions/ — final submission CSVs per version
picks/ — per-query LLM output caches (V4-Pro picks, LoRA picks, profiles)
training/ — LEXam fine-tuning data… See the full description on the dataset page: https://huggingface.co/datasets/Dharun72/llm-agentic-swiss-legal-checkpoints.agentboard-babyai-v1-v071-four-run-checkpoints
AgentBoard BabyAI v1 Prime-RL Four-Run Checkpoints
This dataset preserves four runs from the v071 experiment series using
Qwen/Qwen3.5-9B, Prime-RL 0.7.0 at commit
cb74ae2bebe9710b11551970a9661d10116c7179, and the packaged
agentboard-babyai-v1-context 0.2.2 Verifiers environment.
Runs
Directory
Algorithm
ECHO user alpha
Exact resume step
Final trainer step
Retained adapters
rlonly/
GRPO
0
95
100
21
echo005/
ECHO
0.05
100
100
23
echo050/
ECHO
0.5
95… See the full description on the dataset page: https://huggingface.co/datasets/bhoy/agentboard-babyai-v1-v071-four-run-checkpoints.LeroyDyer__CheckPoint_R1-details
Dataset Card for Evaluation run of LeroyDyer/CheckPoint_R1
Dataset automatically created during the evaluation run of model LeroyDyer/CheckPoint_R1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__CheckPoint_R1-details.LeroyDyer__CheckPoint_A-details
Dataset Card for Evaluation run of LeroyDyer/CheckPoint_A
Dataset automatically created during the evaluation run of model LeroyDyer/CheckPoint_A
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__CheckPoint_A-details.LeroyDyer__CheckPoint_C-details
Dataset Card for Evaluation run of LeroyDyer/CheckPoint_C
Dataset automatically created during the evaluation run of model LeroyDyer/CheckPoint_C
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__CheckPoint_C-details.LeroyDyer__CheckPoint_B-details
Dataset Card for Evaluation run of LeroyDyer/CheckPoint_B
Dataset automatically created during the evaluation run of model LeroyDyer/CheckPoint_B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/LeroyDyer__CheckPoint_B-details.tinker-rl-bench-checkpoints
TinkerRL-Bench Checkpoint Manifest
A catalogue of every Tinker training run and checkpoint referenced by our
NeurIPS paper "A Unified Benchmark for RL Post-Training of Language
Models" (repo).
Because Tinker stores weights behind an authenticated tinker://... URI
(only the account that ran the training can materialise them), this
dataset does not contain the raw .safetensors/archive blobs — it
contains the canonical pointer table and full training metadata so
anyone with a Tinker… See the full description on the dataset page: https://huggingface.co/datasets/arvindcr4/tinker-rl-bench-checkpoints.droid-action-checkpointheretic-checkpointscarbon-tax-checkpointsmeditsolutions__Llama-3.2-SUN-2.4B-checkpoint-34800-details
Dataset Card for Evaluation run of meditsolutions/Llama-3.2-SUN-2.4B-checkpoint-34800
Dataset automatically created during the evaluation run of model meditsolutions/Llama-3.2-SUN-2.4B-checkpoint-34800
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__Llama-3.2-SUN-2.4B-checkpoint-34800-details.meditsolutions__Llama-3.2-SUN-2.4B-checkpoint-26000-details
Dataset Card for Evaluation run of meditsolutions/Llama-3.2-SUN-2.4B-checkpoint-26000
Dataset automatically created during the evaluation run of model meditsolutions/Llama-3.2-SUN-2.4B-checkpoint-26000
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/meditsolutions__Llama-3.2-SUN-2.4B-checkpoint-26000-details.feedback-to-lora-step0-checkpointcirr-augmented-captions-checkpointdataset-checkpointasr-eval-checkpointttllamaA0-checkpointsllamaA7-checkpointspixeldit-xl-three-1p6m-checkpoints-20260922
Three full PixelDiT-XL checkpoints at optimizer step 1,600,000
These are distinct experimental branches, not three copies of one model.
Each file is a full training checkpoint (online weights, EMA, optimizer and
resumption state), not an inference-only weights export. For evaluation, use
the EMA weights and keep each branch label attached to its results.
Directory
Branch
h200_lr2e5_strict_rng_from1m
H200 strict-RNG continuation from the 1M LR2e-5 checkpoint… See the full description on the dataset page: https://huggingface.co/datasets/ziqiaow/pixeldit-xl-three-1p6m-checkpoints-20260922.
