datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-07-31-toolcalling-tulu-20-80-mixture
Tool-calling + TULU3 replay SFT mixture (20/80) for Qwen3.6-27B
The training mixture behind
LASR-Callum/2026-07-31-wrongly-trained-qwen36-toolcalling-tulu-lora-20-80: 1,492,442 Qwen3.6
tokens across 2,002 pre-rendered conversations, split
19.96% agentic tool-use / 80.04% TULU3 replay.
Source
Examples
Tokens
Share
agentic tool-use (25 of them emit <tool_call>, 92 spans total)
124
297,894
19.96%
TULU3 replay
1,878
1,194,548
80.04%
Total
2,002
1,492,442… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-31-toolcalling-tulu-20-80-mixture.OpenVetQA-sft-v1
OpenVetQA v1.0
A veterinary multiple-choice SFT dataset where every item is traceable to a CC-BY source article. Questions were generated from — and verified against — single passages from PMC open-access reviews and guidelines, then filtered by structural gates and blind cross-family LLM judging.
2,651 items · CC BY-NC 4.0 · Trust report · Gemma adapter · Qwen adapter
Results
Two 4B models were fine-tuned on the train split and evaluated on two held-out tiers.… See the full description on the dataset page: https://huggingface.co/datasets/Douxgen/OpenVetQA-sft-v1.2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think
Qwen3.6-27B SFT mixture — 500k maths-weighted, empty-think markers
499,595 tokens across 1,001 conversations, weighted toward maths, with Qwen3.6's empty
think marker on the non-maths rows. md5 c433f31eba2b5b4919fb166043caccb5.
Source
Examples
Tokens
Share
Marker
NuminaMath-CoT
611
333,351
66.9%
no
No Robots
271
82,239
16.5%
yes
TULU3
119
82,445
16.5%
yes
Total
1,001
499,595
390 marked
Derived from
qwen3.6-27b-mixture-500k-numina-heavy
by adding the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-02-qwen36-mixture-500k-numina-heavy-empty-think.2026-08-31-difficult-advice-716-seeds-bundle
da716 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates of the da716 arm (Table2 9,284 filtered + difficult-advice-v2 716, 7.16%) so the arm carries training-seed variance like its siblings. da716 was the last arm on a single seed and is the comparison baseline for the… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-716-seeds-bundle.2026-08-17-table2-9284-peer-critique-good-716-train-mixture
Qwen3.6-27B SFT mixture: 9,284 Table2 + 716 peer_critique GOOD ARM (10,000 rows)
The one-variable twin of LASR-Callum/2026-08-16-table2-9284-peer-critique-716-train, whose
716 peer-critique rows are 358 good / 358 flawed. Here all 716 are drawn from the good arm.
field
value
experiment
Arm ablation: does the peer-critique FLAWED arm contribute anything? Train on good-arm-only critiques and compare against the 358/358 arm.
date_generated
2026-08-17
constitution… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-17-table2-9284-peer-critique-good-716-train-mixture.2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle
chunk-only 702 seed replicates — training bundle (seeds 42 and 69)
code.tar.gz (trainer + src/ + the two seed configs) beside seed 0's mixture,
byte-identical. scripts/gpu/runpod_train.py up reads both from this one repo.
field
value
experiment
Seed replicates so this arm carries training-seed variance. Table2 9,284 filtered + chunk-only difficult advice 702 (7.03%). The rewrite stages never saw the constitution, only their one target principle. Between-seed spread on… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-principle-scoped-702-seeds-bundle.fvapps
Proving the Coding Interview: Formally Verified APPS
Paper
by Ronak Mehta and Quinn Dougherty. Based on APPS by Hendrycks et al.
We introduce the Formally Verified Automated Programming Progress Standards, or FVAPPS, a benchmark of 4715 samples for writing programs and proving their correctness, the largest formal verification benchmark, including 1083 curated and quality controlled samples. Previously, APPS provided a benchmark and dataset for programming puzzles to be completed… See the full description on the dataset page: https://huggingface.co/datasets/quinn-dougherty/fvapps.doundobench
Do-Undo Bench Annotations
Dataset Description
This dataset contains action annotations for egocentric video clips. Each example includes the original action narration, temporal boundaries, verb and noun labels, and paired natural-language prompts describing the forward action and its reverse or undo action.
Dataset Structure
The repository contains two annotation JSON files and one Croissant metadata file:
File
Purpose
Examples… See the full description on the dataset page: https://huggingface.co/datasets/doundo/doundobench.
