datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.zscreen-pilot
Z-Screen pilot: chemical recipes and cellular responses
Version 2.0.0
Z-Screen connects combinatorial chemistry to high-dimensional cellular measurements. This pilot data resource contains 190,699 compound-context profiles from 162,914 public chemical recipes across eight library–cell contexts. Each RNA response is available on a 6,000-gene panel and as coordinates on 32 shared transcriptional programs.
The package provides processed data, aligned identifiers, reference model… See the full description on the dataset page: https://huggingface.co/datasets/Zafrens/zscreen-pilot.vistr-process-verification-pilot
ViSTR Process-Verification Pilot (14 answer-correct trajectories, multimodal)
Agent trajectories for studying process false positives in multimodal agents:
cases where the answer is correct but the visual reasoning that produced it is
wrong. Ships the raw perception tool outputs so any claim in a trajectory can be
independently re-verified, plus human annotations and an unmodified XSkill
critique of the same trajectories.
Why this exists
Harness / skill… See the full description on the dataset page: https://huggingface.co/datasets/MihailSlutsky/vistr-process-verification-pilot.jev-stage2-image-beans-pilot
Beans: one natural question per image
Open the corrected preview.
natural_v4 is the recommended and default preview: 100 original images, 100 rows, one three-way condition-class Choice question per image. All targets come directly from the source labels column (34 angular leaf spot, 33 bean rust, 33 healthy). Original image bytes and source annotations are unchanged.
Example question: “Which source-defined condition class describes the bean leaf?” Options: angular_leaf_spot… See the full description on the dataset page: https://huggingface.co/datasets/FaroukMoc2/jev-stage2-image-beans-pilot.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.carbon-pilot-corpusproperty-pilot-tickets
🏢 PropertyPilot — Maintenance Tickets
A synthetic dataset of 13,725 residential-maintenance tickets written the way real tenants write them — polite, panicked, passive-aggressive, or confused — each paired with operational metadata (category, urgency, assigned contractor, cost, resolution time).
Built for an end-to-end NLP pipeline: triage classification, similar-case retrieval (embeddings + FAISS), and work-order / reply generation.
About this release. Earlier versions of… See the full description on the dataset page: https://huggingface.co/datasets/propertypilot/property-pilot-tickets.SWE-ZERO-V2PRs-1k-pilotfr-bj-speech-pilot
fr_bj, Beninese French read-speech pilot
Read speech in Beninese French, in the FLEURS format. The sentences were
written in Benin, about local realities, and read by Beninese speakers in their
own French. For evaluation, not for training.
See DATASHEET.md for provenance and intended use.
Segments
209
Sentences
105, all covered
Speakers
2 (1 female, 1 male)
Duration
25.0 min
Words
2921
Audio
WAV PCM 16-bit, 16 kHz, mono
Split
single test split… See the full description on the dataset page: https://huggingface.co/datasets/labari-voice/fr-bj-speech-pilot.fr-sn-speech-pilot
fr_sn, Senegalese French read-speech pilot
Read speech in Senegalese French, in the FLEURS format. The sentences were
written in Senegal, about local realities, and read by Senegalese speakers in
their own French. For evaluation, not for training.
See DATASHEET.md for provenance and intended use.
Segments
210
Sentences
105, all covered
Speakers
5 (3 female, 2 male)
Duration
19.3 min
Words
2832
Audio
WAV PCM 16-bit, 16 kHz, mono
Split
single test split… See the full description on the dataset page: https://huggingface.co/datasets/labari-voice/fr-sn-speech-pilot.mixed-language-detection-pilot-complete-sentences
Mixed-Language Speech Detection Pilot — Complete Sentences
This is the complete-sentence revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed
Earlier generation forced source transcripts into arbitrary… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-complete-sentences.mixed-language-detection-pilot-fleurs-voices
Mixed-Language Speech Detection Pilot — Native FLEURS Voices
This is the native-reference revision of a 6,000-clip binary
audio-classification pilot. label = 0 denotes one intended language and
label = 1 denotes more than one intended language. The covered languages are
Turkish (tur), Northern Kurdish/Kurmanji (kmr), Central Kurdish/Sorani
(ckb), Arabic (ara), Persian (fas), and English (eng).
What changed in this revision
Synthetic speech is cloned from 36 real… See the full description on the dataset page: https://huggingface.co/datasets/TartarusXXX/mixed-language-detection-pilot-fleurs-voices.vesuvius-physical-fusion-pilot
Vesuvius finite-thickness fusion pilot
This is the preregistered eight-cell handoff requested by Jinho Jeong in
ScrollPrize/villa issue #191.
It is designed to measure whether a frozen Vesuvius surface checkpoint keeps
neighbouring finite-thickness sheets separated as their true air gap closes.
Fixed design
painter: 2c483dd
checkpoint: scrollprize/surface_recto_059_redo, Model_epoch499.pth
papyrus level 90; noise sigma 6
fixed 150 µm sheets; 30 µm voxels; 12… See the full description on the dataset page: https://huggingface.co/datasets/AviadCoh/vesuvius-physical-fusion-pilot.r1-d002-number-pointing-pilot-20260908
R1 D002 number-pointing pilot
Private engineering pilot converted to LeRobot Dataset v3.0 from the accepted
episodes of run d002_20260908T020005Z.
This upload is for validating the conversion, Hub viewer, download, and smoke
training workflow. It is not a production training dataset and makes no
hardware-readiness claim.
Contents
7 episodes, 733 frames, 8 FPS
one 640×480 simulated head-camera stream
Unitree R1 A5 arm state (10,) and arm action (10,)
per-episode… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/r1-d002-number-pointing-pilot-20260908.loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.qwen3-4b-perfectblend-deepspec-rollout
Qwen3-4B PerfectBlend DeepSpec Rollout
This dataset contains the complete DeepSpec-aligned Qwen3-4B
self-distillation rollout over the filtered PerfectBlend corpus. The seeded
95/5 split is published as separate train and eval splits.
Splits
Split
Conversations
Shards
Path
train
1,349,860
128
data/*.jsonl
eval
71,046
64
eval/*.jsonl
total
1,420,906
192
Data construction
Canonical filtered corpus: 1,420,906 conversations.
Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.HyPoradise-pilot
Dataset Name: Pilot dataset for Multi-domain ASR corrections
Description
This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains.
It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0
Structure
Data Split
The dataset is divided into training and test splits:
Training Data: 281,082 entries
Approximately… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HyPoradise-pilot.chinchilla-m5.1-predictive-logo-pilot
MarinDNA m5.1 chinchilla predictive sequence logo
This dataset contains canonical A/C/G/T log-probability and derived glyph-height
BigWigs from the MarinDNA m5.1 two-pass predictive sequence-logo approximation on
the UCSC/NCBI RefSeq chinchilla assembly GCF_000276665.1. It was produced
with marin-dna/marin-dna-exp135-m5.1 at immutable
revision c0676b2012b8b9c526deb26ff517f6b92b6d375d by the commit-pinned chinchilla-logo pipeline.
This is a predictive next-token logo, not an LLR… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/chinchilla-m5.1-predictive-logo-pilot.docclass-pilot
Docclass Pilot Sample
Deterministic stratified pilot slice of
Lucius-Morningstar/docclass-merged v5
(KANBAN-084, built 2026-08-24): every doc type and every subclass stratum
present in the parent contributes at least one row; each stratum contributes
at most 3 rows, selected by ascending sha256(filename) within the
stratum (content-addressed order — decorrelated from the md5-keyed family
split rule; rebuilds are byte-identical).
Coverage
doc_type
rows… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/docclass-pilot.muse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.instruction-pilot-outputs-filteredswe-agent-cpu-dynamorio-pilot-sympy-15599
One-agent CPU trace pilot
Preliminary research data. Validation is incomplete; this is not a confirmed dead-state result.
One live mini-SWE-agent 2.4.6 execution of sympy__sympy-15599, using a separate Qwen3-Coder-30B-A3B-Instruct AWQ server. The collector finished normally in 907.94 seconds. The agent made 57 model calls and submitted a patch; benchmark evaluation was not run. This is mini-SWE-agent, not the original full SWE-agent implementation.
What is included… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-pilot-sympy-15599.korean-embedding-performance-v1-pilot-50k
Korean Embedding Performance v1 — Pilot 50K
주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다.
사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다.
파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은
ablation-200k이다.
Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용
contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개,
hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다.
사용 조건과 공개 범위
이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.fineweb-legal-pilot
⚖️ FineWeb-Legal-Pilot
66.8M words of the finest legal domain data the 🌐 web has to offer.
Repo: GitHub | Report: Technical Report
What is it?
FineWeb-Legal-Pilot is a pilot dataset consiting of 52k high-quality legal documents filtered from the 10-billion-token sample of 🍷 FineWeb.
To enhance FineWeb's utility for legal AI domain adaptation, we draw inspiration from the FineWeb-Edu methodology: creating a legal quality classifier using annotations… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/fineweb-legal-pilot.bratota_pilot_training_datazveb-pilot-results
ZVEB Pilot Results (zambodotdev)
The ZVEB pilot benchmark: 8 agent-tool tasks run against Zambo's live MCP endpoint (zambo.dev), scored 9.38/10 on pilot run v0 (strict first-attempt 8.75/10). Supplement re-runs are disclosed on 3 of 8 tasks: catalog-qrcode, credits-ai, universal-multistep. Each task directory holds the full execution trace with verifiable AI agent execution receipts (AER-1, an open draft/RFC).
Contents
runs/run-20260921-145712/ — full pilot run… See the full description on the dataset page: https://huggingface.co/datasets/zambodotdev/zveb-pilot-results.uldr-v0.1-pilot
Ukrainian Language Decolonization & Reasoning (ULDR) — Pilot Canary Release v0.1
[!IMPORTANT]
Exploratory Pilot / Canary Release (v0.1): This dataset represents an early exploratory pilot canary release (v0.1-pilot) establishing our baseline data pipeline, schema contracts, and directional validation. It is not the final production release (v1.0). The full production release is scheduled for Phase 5.5 following complete evaluation suite assembly, dialect & historical protection… See the full description on the dataset page: https://huggingface.co/datasets/krisztiankoos/uldr-v0.1-pilot.pdecert-pilot
PDECert Natural-Candidate Pilot
This is a provenance-bearing pilot benchmark for checking symbolic candidate
solutions to partial differential equations. Each row contains the unedited
generator output, a fully instantiated verification case, content digest,
producer metadata, and completed human annotation.
Dataset summary
Records: 20
Symbolic-solver outputs: 10
Open-model outputs: 10
Valid: 10
Invalid: 10
Unclear: 0
Corpus SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/oroikono/pdecert-pilot.amazon-coevolve-midtrain-pilot-sft-preview
Amazon Coevolve Mid-training SFT Preview
This is an accepted-only point-in-time snapshot of amazon-mt-pilot-native-v7 for inspecting and launching initial SFT experiments. Collection is still active, so this repository is deliberately marked incomplete.
Configurations and splits
Configuration
Train
Validation
single
540
11
structured
1,239
24
diff
1,240
23
merged
2,561
50
The 50-row merged validation set holds out one complete reviewer… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/amazon-coevolve-midtrain-pilot-sft-preview.
