datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-mid-checkpoint-transcripts
RLVR reward-hacking mid-checkpoint full trajectories
This release contains 600 full held-out trajectories from intermediate RLVR
checkpoints selected to yield substantially more balanced reward-hacking datasets: 300
from Qwen3.5-9B at optimizer update 110 and 300 from GPT-OSS-120B at update 180.
Each row preserves the task and tests, complete prompts, native reasoning, final answer,
rendered and sampled token IDs, token log-probabilities, sampling metadata, extracted
files… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-mid-checkpoint-transcripts.tinker-rl-bench-checkpoints
TinkerRL-Bench Checkpoint Manifest
A catalogue of every Tinker training run and checkpoint referenced by our
NeurIPS paper "A Unified Benchmark for RL Post-Training of Language
Models" (repo).
Because Tinker stores weights behind an authenticated tinker://... URI
(only the account that ran the training can materialise them), this
dataset does not contain the raw .safetensors/archive blobs — it
contains the canonical pointer table and full training metadata so
anyone with a Tinker… See the full description on the dataset page: https://huggingface.co/datasets/arvindcr4/tinker-rl-bench-checkpoints.
