datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
self-reward-collapse-terse
self-reward-collapse-terse
Per-round trajectory and held-out samples from an iterative self-training loop on GSM8K (Qwen2.5-7B-Instruct, LoRA DPO, 6 rounds). Same loop every arm; only the preference-label source differs. This dataset is the terse arm: pairs labelled by a deliberately gameable brevity reward (the shortest of K samples wins).
Study question: does a model training on its own judgment collapse? Short answer on verifiable math: the reward can be hacked, but… See the full description on the dataset page: https://huggingface.co/datasets/yavuz-ai/self-reward-collapse-terse.danish-tool-procv4-abl6-tersehelpsteer2_binarized_tersenuextract3-cards-terse
NLS Advocates Library index cards → structured JSON (NuExtract3)
Demo output: scanned manuscript index cards from the National Library of Scotland's
Advocates Library, run through NuExtract3
(4B, Apache-2.0) for schema-guided structured extraction. Each row pairs the card
image with extraction — the JSON the model returned for that card.
How it was made
Input: NationalLibraryOfScotland/nls-index-cards-object-detection, filtered to the 49 pages that contain a card… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/nuextract3-cards-terse.
