datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
20241230_icl_output_supervisionsupervision-tradeoff
The Supervision Tradeoff — Reproducibility Bundle
Format Scaffolds, Judgment Pleasing, and Anti-Calibration in Post-Training
Paper DOI: 10.5281/zenodo.19748277 · Concept DOI: 10.5281/zenodo.19748276 · Code repo: github.com/codex-curator/supervision-tradeoff
Author: Tad MacPherson, Metavolve Labs · ORCID: 0009-0002-8659-7479
What this is and why it might help your research
This repository ships everything we used to falsify our own headline finding, in a form you can… See the full description on the dataset page: https://huggingface.co/datasets/Metavolve-Labs/supervision-tradeoff.rsi-synthetic-world-supervision-public2026-08-31-cot-only-supervision-chunk-only-702
CoT-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-chunk-only-702.2026-09-01-empty-cot-supervision-chunk-only-702
Empty-CoT supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: REPLACE each of the 702 principle-scoped difficult-advice rows' reasoning traces with the empty think marker, leaving prompt and answer byte-identical. The marker is masked whole by the existing generation-boundary rule, so the model is supervised on the visible answer and never learns to emit an empty close. Tests whether the reasoning was doing work AS… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-empty-cot-supervision-chunk-only-702.2026-08-31-cot-only-supervision-t2-9284-synthdoc-716
CoT-only supervision mixture (Table2 9,284 + difficult-advice 716)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated
2026-08-31… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-cot-only-supervision-t2-9284-synthdoc-716.2026-08-31-odcv-cot-only-supervision-716-1x65
ODCV-Bench eval of LASR-Callum/2026-08-31-qwen36-lora-table2-9284-synthdoc-716-cotonly-rank-64 (mode=think) - the CoT-only supervision arm, whose 716 difficult-advice rows trained on their REASONING ONLY (each row truncated at its reasoning close, answer removed from both the loss and the forward pass). 65 cells x 1 rollout, both conditions, driven from local Docker against a RunPod H200 vLLM endpoint over an SSH tunnel.
field
value
experiment
ODCV-Bench eval of… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-odcv-cot-only-supervision-716-1x65.2026-08-31-difficult-advice-716-cot-only-supervision
CoT-only supervision mixture, difficult-advice-v2 (Table2 9,284 + DA-v2 716)
field
value
experiment
Arm: train the 716 difficult-advice rows on their REASONING ONLY — each row is truncated at its </think> close, so the visible answer leaves both the loss and the forward pass — while the 9,284 Table2 rows train exactly as in the control. Tests whether the difficult-advice effect on agentic misalignment is carried by the reasoning or by the answer.
date_generated… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-08-31-difficult-advice-716-cot-only-supervision.2026-09-01-answer-only-supervision-chunk-only-702
Answer-only supervision mixture, principle-scoped (Table2 9,284 + chunk-only 702)
field
value
experiment
Arm: train the 702 principle-scoped difficult-advice rows on their VISIBLE ANSWER ONLY -- the reasoning trace stays in the token stream as unsupervised context (no truncation, full forward pass) and simply earns no loss, while the 9,284 Table2 rows train exactly as in the control. The EXACT COMPLEMENT of the CoT-only arm on the same base: on every one of the 702… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-01-answer-only-supervision-chunk-only-702.africa-synth-chw-performance-supervision-all
CHW Performance & Supervision Dataset (iCCM, Stockouts, Motivation) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-chw-performance-supervision-all.toolcomp_process_supervision_eval
Dataset Card for "toolcomp_process_supervision_eval"
More Information needed
Evaluation code can be found at https://github.com/vaskar-open-source-research/toolcomp and paper can be found at https://arxiv.org/abs/2501.01290.
eeg-self-supervisioncybersecurity-supervision-feedbacksupervision
Dataset Card for supervision
This dataset has been created with Argilla. As shown in the sections below, this dataset can be loaded into your Argilla server as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Using this dataset with Argilla
To load with Argilla, you'll just need to install Argilla as pip install argilla --upgrade and then use the following code:
import argilla as rg
ds =… See the full description on the dataset page: https://huggingface.co/datasets/allezallezallez/supervision.pharmacy-technician-supervision-ratios
State pharmacy technician-to-pharmacist supervision ratios
Canonical, always-current version: https://referencesource.org/pharmacy-technician-supervision-ratios/
Machine-readable: https://referencesource.org/pharmacy-technician-supervision-ratios/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-16
Stale after: 2027-08-16 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 21
The maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/pharmacy-technician-supervision-ratios.Supervision-PLF-PLR-Prevision-Execution-2019levir-yolov8n-p2-deep-supervision-gap-ftal-seed42cape-train-supervision
CAPE Train Supervision Maps (colpali_train_set, 20,000 samples × 7 versions)
Evidence-localization supervision maps for the first 20,000 samples of
vidore/colpali_train_set.
For each sample we provide 7 map variants, all on the same ColQwen2.5 patch grid.
Versions (7)
version
method
block_only_K1/K2/K3
MinerU2.5 block layout → block-level leave-one-out (blur block, re-encode, MaxSim drop); the drop of each of the top-K blocks is copied to all its… See the full description on the dataset page: https://huggingface.co/datasets/wm07070/cape-train-supervision.cape-train-supervision-loo2x2k1-full
CAPE Train Supervision Maps — block_loo2x2_K1 (FULL colpali_train_set)
Evidence-localization supervision maps for the entire
vidore/colpali_train_set
(118,195 samples), on the ColQwen2.5 patch grid.
version
#samples
method
block_loo2x2_K1
118,195 (FULL)
top-1 MinerU2.5 block → 2×2 super-patch leave-one-out (ColQwen2.5 MaxSim drop)
Settings
ColQwen2.5-v0.2 MaxSim("{query} {answer}") · Gaussian blur r=15 · MinerU2.5 block layout · seed 42.… See the full description on the dataset page: https://huggingface.co/datasets/wm07070/cape-train-supervision-loo2x2k1-full.Supervision-Projet-Loi-Finance-Prevision-Realisationai-5node-sup-buf-lag-cpl-supervision-breakdown-v0.1
What this repo does
This dataset models supervision breakdown cascades in AI operations. It detects when supervision pressure rises, protective buffers weaken through expanding auto-approval, governance lag delays overrides, and tight coupling through shared supervisor pools propagates bottlenecks across products, crossing the five-node cascade threshold into an unrecoverable supervision breakdown cascade.
This dataset models a five-node cascade: four interacting instability drivers… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/ai-5node-sup-buf-lag-cpl-supervision-breakdown-v0.1.Supervisionado
