datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gsm8k_only_answerThe data is exactly like the original GSM8k (https://huggingface.co/datasets/gsm8k ), but with the label consisting of the correct answer(one number) only.
@misc{krishna2024gsmansweronly,
title={GSM8k (Answer only)},
author={Satyapriya Krishna},
year={2023},
url={skrishna/gsm8k_only_answer},
}
2026-09-16-da-7-answer-only-mix
DA supervision answer; all 752 DA and 9284 identical replay rows
field
value
experiment
DA supervision answer; all 752 DA and 9284 identical replay rows
date_generated
2026-09-16
constitution
constitutions/claude_distilled_09_principles/constitution.md
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @ 4648153af4b834b70bd2e5374f639aaad219c83c
models
Tokenizer Qwen/Qwen3.6-27B@6a9e13bd6fc8f0983b9b99948120bc37f49c13e9; replay… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-16-da-7-answer-only-mix.2026-09-01-answer-only-supervision-chunk-only-702
Answer-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: train the 702 principle-scoped difficult-advice rows on their VISIBLE ANSWER ONLY -- the reasoning trace stays in the token stream as unsupervised context (no truncation, full forward pass) and simply earns no loss, while the 9,284 Table2 rows train exactly as in the control. The EXACT COMPLEMENT of the CoT-only arm on the same base: on every one of the 702… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702.Eklav-Reranker-AnswerOnly-Data
Eklav-Reranker-AnswerOnly-Data
Training data for the Eklav paper.
Task: passage reranking (BRIGHT / NevIR benchmarks)
Method: Answer-only (no reasoning of any kind -- the no-CoT floor)
Examples: 381,934 train / 3,857 held-out val
Format: ShareGPT (system + conversations: [{from, value}]), used for LoRA SFT via LLaMA-Factory.
Single-turn ShareGPT conversations. Each row: a query+passage relevance-judgment prompt (human turn) and a bare true/false judgment (gpt turn) -- no hint… See the full description on the dataset page: https://huggingface.co/datasets/AdarshSingh7647/Eklav-Reranker-AnswerOnly-Data.
