datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-07-29-msm-philosophy-spec-surf-audit
SURF audit: harmful-omission rubric against the MSM+AFT+CoT checkpoint
experiment: SURF (Surfacing Unintended Response Failures) EM-loop search over a generic instruction-following prompt pool, scoring responses against a harmful-omission rubric, against the primary MSM target checkpoint. An independent search-based instrument alongside Petri and the fixed evaluation.
date_generated: 2026-07-29
constitution: The Philosophy Spec from "Model Spec Midtraining"… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-07-29-msm-philosophy-spec-surf-audit.2026-08-07-surf-synthdoc-difficult-advice-attributes-full
SURF Attributes (Full)
Complete dataset for SURF research and extension.
Paper: Chunky Post-Training (link pending)
Quick Start
For running SURF, use the minimal dataset: LASR-Callum/2026-08-07-surf-synthdoc-difficult-advice-attributes
uv run -m surf.cli.main sweep \
--attributes LASR-Callum/2026-08-07-surf-synthdoc-difficult-advice-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Dataset Fields
prompt: The query text… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-07-surf-synthdoc-difficult-advice-attributes-full.2026-08-07-surf-synthdoc-difficult-advice-attributes
SURF Attributes
Minimal dataset for running SURF (Surfacing Unintended Response Failures).
Paper: Chunky Post-Training (link pending)
Usage
uv run -m surf.cli.main sweep \
--attributes LASR-Callum/2026-08-07-surf-synthdoc-difficult-advice-attributes \
--rubric rubrics/rebuttal.yaml \
-o results/
Fields
prompt: The query text
sae_attributes: List of semantic attribute cluster summaries
How it works
Each prompt was… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-07-surf-synthdoc-difficult-advice-attributes.atlas-21-neutral-text-on-the-tool-surface
ATLAS report 21: is it the words or the shape? The neutral selector's text on the orchestrator's conversation
Complete raw products of ATLAS rl-training report 21 (GitHub issue #45).
Two new selector
surfaces over every question of the canonical LiveCodeBench (175) and GPQA
(198) validation sets, each question with all eight of its cached candidates
revealed. Both carry the neutral selector's text, and only that text, on the
orchestrator's conversation shape: an empty system… See the full description on the dataset page: https://huggingface.co/datasets/t2ance/atlas-21-neutral-text-on-the-tool-surface.task1428_country_surface_area
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1428_country_surface_area
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1428_country_surface_area.surfer
🏄 Surfer Voice Dataset
A conversational dataset designed to fine-tune language models to speak like a surfer dude!
Each example contains a user question and a response written in authentic surf culture style,
with beach slang, wave metaphors, and a totally chill laid-back vibe.🤙
This dataset was used to train the Quill Voice Surfer voice model:
https://huggingface.co/quill-voice/surfer
📊 Dataset Details
Property
Details
Size
566 examples
Format… See the full description on the dataset page: https://huggingface.co/datasets/quill-voice/surfer.cve-decision-seeds
CVE Decision Seeds (500 Clean Verified Seeds)
This dataset contains 500 high-fidelity, verified C/C++ vulnerability seeds generated using the GEPA-First (Generative Explanation of Program Anomalies) framework for the BARRED synthetic debate pipeline.
Overview
Source Corpus: Extracted from CVEFixes.
Clean-Room Anti-Leakage Partitioning: Excluded against all 5,000 held-out evaluation scenarios in cve-decision using exact, normalized, and 5-gram fuzzy shingling ($J… See the full description on the dataset page: https://huggingface.co/datasets/surfiniaburger/cve-decision-seeds.
