datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
2026-09-14-dataset-refresh-revised-pilot-audit
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_230322
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo
https://github.com/Matthew-Bozoukov/Lessons_from_constituitional_AFT @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-revised-pilot-audit.2026-09-14-dataset-refresh-pilot-audit
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
field
value
experiment
Failed first pilots for moral low-stakes and nonmoral craft advice refresh; audit evidence only
date_generated
20260914_224408
constitution
constitutions/claude_distilled_09_principles/constitution.md; low-stakes principle generation, nonmoral compatibility review only
source_repo… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-14-dataset-refresh-pilot-audit.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.vesuvius-physical-fusion-pilot
Vesuvius finite-thickness fusion pilot
This is the preregistered eight-cell handoff requested by Jinho Jeong in
ScrollPrize/villa issue #191.
It is designed to measure whether a frozen Vesuvius surface checkpoint keeps
neighbouring finite-thickness sheets separated as their true air gap closes.
Fixed design
painter: 2c483dd
checkpoint: scrollprize/surface_recto_059_redo, Model_epoch499.pth
papyrus level 90; noise sigma 6
fixed 150 µm sheets; 30 µm voxels; 12… See the full description on the dataset page: https://huggingface.co/datasets/AviadCoh/vesuvius-physical-fusion-pilot.r1-d002-number-pointing-pilot-20260908
R1 D002 number-pointing pilot
Private engineering pilot converted to LeRobot Dataset v3.0 from the accepted
episodes of run d002_20260908T020005Z.
This upload is for validating the conversion, Hub viewer, download, and smoke
training workflow. It is not a production training dataset and makes no
hardware-readiness claim.
Contents
7 episodes, 733 frames, 8 FPS
one 640×480 simulated head-camera stream
Unitree R1 A5 arm state (10,) and arm action (10,)
per-episode… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/r1-d002-number-pointing-pilot-20260908.muse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.qwen3-4b-perfectblend-deepspec-rollout
Qwen3-4B PerfectBlend DeepSpec Rollout
This dataset contains the complete DeepSpec-aligned Qwen3-4B
self-distillation rollout over the filtered PerfectBlend corpus. The seeded
95/5 split is published as separate train and eval splits.
Splits
Split
Conversations
Shards
Path
train
1,349,860
128
data/*.jsonl
eval
71,046
64
eval/*.jsonl
total
1,420,906
192
Data construction
Canonical filtered corpus: 1,420,906 conversations.
Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.swe-agent-cpu-dynamorio-pilot-sympy-15599
One-agent CPU trace pilot
Preliminary research data. Validation is incomplete; this is not a confirmed dead-state result.
One live mini-SWE-agent 2.4.6 execution of sympy__sympy-15599, using a separate Qwen3-Coder-30B-A3B-Instruct AWQ server. The collector finished normally in 907.94 seconds. The agent made 57 model calls and submitted a patch; benchmark evaluation was not run. This is mini-SWE-agent, not the original full SWE-agent implementation.
What is included… See the full description on the dataset page: https://huggingface.co/datasets/harry1332/swe-agent-cpu-dynamorio-pilot-sympy-15599.instruction-pilot-outputs-filteredkorean-embedding-performance-v1-pilot-50k
Korean Embedding Performance v1 — Pilot 50K
주의: 이 revision은 공개 benchmark 성능 후보 학습에 사용하면 안 된다.
사후 15-task exact text-hash 감사에서 평가 query 고유 hash 4개가 확인됐다.
파이프라인·최적화 진단과 contamination ablation에만 남기며, 교체본은
ablation-200k이다.
Qwen3-Embedding 계열의 한국어 retrieval 성능 실험을 위한 50,000-row 연구용
contrastive dataset이다. 각 row는 instruction-aware query, positive passage 1개,
hard/easy negative passage 1–7개를 ms-swift embedding message schema로 저장한다.
사용 조건과 공개 범위
이 저장소의 통합 라이선스는 other다.… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-performance-v1-pilot-50k.zveb-pilot-results
ZVEB Pilot Results (zambodotdev)
The ZVEB pilot benchmark: 8 agent-tool tasks run against Zambo's live MCP endpoint (zambo.dev), scored 9.38/10 on pilot run v0 (strict first-attempt 8.75/10). Supplement re-runs are disclosed on 3 of 8 tasks: catalog-qrcode, credits-ai, universal-multistep. Each task directory holds the full execution trace with verifiable AI agent execution receipts (AER-1, an open draft/RFC).
Contents
runs/run-20260921-145712/ — full pilot run… See the full description on the dataset page: https://huggingface.co/datasets/zambodotdev/zveb-pilot-results.uldr-v0.1-pilot
Ukrainian Language Decolonization & Reasoning (ULDR) — Pilot Canary Release v0.1
[!IMPORTANT]
Exploratory Pilot / Canary Release (v0.1): This dataset represents an early exploratory pilot canary release (v0.1-pilot) establishing our baseline data pipeline, schema contracts, and directional validation. It is not the final production release (v1.0). The full production release is scheduled for Phase 5.5 following complete evaluation suite assembly, dialect & historical protection… See the full description on the dataset page: https://huggingface.co/datasets/krisztiankoos/uldr-v0.1-pilot.2026-09-10-nonmoral-grounded-revision-pilot-audit
Grounded nonmoral full-response revision: two bounded pilots; candidate stopped before production
field
value
experiment
Grounded nonmoral full-response revision: two bounded pilots; candidate stopped before production
date_generated
2026-09-10
constitution
none applied to model prompts; nonmoral preferences file required as a SynthDoc container only
source_repo
https://github.com/Matthew-Bozoukov/teaching_claude_why_replication.git @… See the full description on the dataset page: https://huggingface.co/datasets/dougalldeepmind/2026-09-10-nonmoral-grounded-revision-pilot-audit.cannabis-fda-extractive-pilot
FDA Cannabis Extractive Experimental Pilot
Experimental, automatically screened, unreviewed draft dataset. This dataset is not medical advice, is not production-ready, and must not be represented as clinician-reviewed, legally cleared, or suitable for patient-facing systems.
This small English conversational dataset was created to test an auditable Gemma 4 fine-tuning pipeline. It contains exact answer passages from captured FDA pages about CBD/cannabis safety, paired with… See the full description on the dataset page: https://huggingface.co/datasets/aznatkoiny/cannabis-fda-extractive-pilot.pdecert-pilot
PDECert Natural-Candidate Pilot
This is a provenance-bearing pilot benchmark for checking symbolic candidate
solutions to partial differential equations. Each row contains the unedited
generator output, a fully instantiated verification case, content digest,
producer metadata, and completed human annotation.
Dataset summary
Records: 20
Symbolic-solver outputs: 10
Open-model outputs: 10
Valid: 10
Invalid: 10
Unclear: 0
Corpus SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/oroikono/pdecert-pilot.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.imagenet-variations-synth-pilot-v3
ImageNet Variations Synth Pilot v3.1 (diverse, 1K)
Huu + Claude multimodal instruction pipeline — quality-fixed re-run.
Pipeline
Flux.1-schnell (Huu style phrasings + visual style boosts + aspect ratios)
→ Seed2 → Florence-2 grounding/caption → tracks A–F → USER / ASSISTANT.
Tracks
A programmatic QA (counts/style/spatial/absence) from detector
B LLaVA-style (conversation / detailed description / complex reasoning)
C Evol-Instruct (seed atypical Q →… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot-v3.veriaudit-pilot-v0.1
VeriAudit Pilot v0.1
Part of VeriAudit — see the full
technical report
in the source repository for complete methodology, statistics, and
limitations. This dataset card summarizes it.
Dataset Summary
VeriAudit Pilot v0.1 is a 240-evaluation pilot benchmark measuring whether
two open-weight language models (Qwen3-8B, Aya Expanse 8B) apply consistent
safety behavior when the same harmful intent is expressed in English, Urdu,
Roman Urdu, or code-switched Roman… See the full description on the dataset page: https://huggingface.co/datasets/abeeranajam31/veriaudit-pilot-v0.1.pilot-licence-minimum-requirements-by-authority
Pilot licence minimum hours, age and prerequisites by civil aviation authority
Canonical, always-current version: https://referencesource.org/pilot-licence-minimum-requirements-by-authority/
Machine-readable: https://referencesource.org/pilot-licence-minimum-requirements-by-authority/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-10
Stale after: 2027-08-10 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/pilot-licence-minimum-requirements-by-authority.ch-pilot-rollouts-qwen3.5-9b
C&H Pilot Rollouts — Qwen/Qwen3.5-9B
20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm
corpus (the open-sourced world from harvey-labs
tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM.
Part of an actor-selection pilot for a world-internalization research project: the goal is to
mine agent trajectories into verified fact stores and rewritten likelihood-training targets.
Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.router-pilot-tasks
BenchGen Router Pilot: task pool
The prompt pool a BenchGen router head trains on: 1110 normalized questions drawn from
public benchmark sources, balanced across domain and difficulty with a fixed seed, so results
reproduce. This is the task half of a paired release — the reward half (how well each pool
agent actually scored on a subset of these) is published separately at
benchgen/router-pilot.
Only task_id is the join key between the two datasets — load both and match on it… See the full description on the dataset page: https://huggingface.co/datasets/benchgen/router-pilot-tasks.lemonseed-rl-tasks-cogen-pilot
lemonseed-rl-tasks-cogen-pilot
LemonSeed — RL co-gen pilot tasks.
Contents
rl_tasks_cogen_pilot.jsonl (61 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
Synthetic, generated programmatically for the LemonSeed 1.5B project (by Geramy L. Loveless). Data authored by Michael Anthony Falabella.
korean-embedding-ko-triplet-hn-pilot-10k
Korean Embedding Ko-Triplet Hard-Negative Pilot 10K
nlpai-lab/ko-triplet-v1.0에서 결정론적으로 뽑은 한국어 retrieval train 10,000행과
validation 512행에 Qwen3-Embedding-8B dense hard negative 4개씩을 붙인 연구용
ms-swift embedding dataset이다.
사용 조건
원 source 카드에 명시적 라이선스가 없어 통합 라이선스는 other, manifest의
release_eligible은 false다. 연구·비상업 성능 실험용이며 이 카드가 원 source의
권리를 재허가하지 않는다.
출처와 sampling
source: nlpai-lab/ko-triplet-v1.0
pinned revision: 1f5d72d21ae8309b5221a588b13930b423385bff… See the full description on the dataset page: https://huggingface.co/datasets/LLM-OS-Models/korean-embedding-ko-triplet-hn-pilot-10k.amalia-pilot-honesty-v2
AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix)
Training data from the first two iterations of a verifier-gated fine-tuning
pilot on AMALIA-9B-0626-DPO,
targeting identity/fact confabulation (the model's weakest measured behavior:
43.3% on our honesty harness). Full methodology, harness, and reports:
github.com/teex-pt/pt-amalia.
These are research pilot artifacts — small by design (the pilot validates
the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.router-pilot
BenchGen Router Pilot: per-question agent rewards
A routing dataset: for each question, how every agent in a fixed pool actually performed.
It is the input a model-selection policy trains on — not a question-answering dataset.
Each row is one task. mean_reward[i] is the accuracy of agent agent_order[i] over
3 independent attempts. Ties are preserved explicitly rather than collapsed by argmax,
because on easy questions several agents are genuinely equal and pretending otherwise… See the full description on the dataset page: https://huggingface.co/datasets/benchgen/router-pilot.gen-games-v9-feedback-pilot1toolcall-tr-pilot
ToolCall-TR — Pilot v0.1
Türkçe, execution-verified tool-calling veri seti. Her tool çağrısı gerçek
bir implementasyon tarafından çalıştırıldı ve çıktısı doğrulandı; hiçbir tool
observation'ı bir LLM tarafından yazılmadı.
⚠️ Önce şunu bilin: bu küçük bir ilk sürüm
İçinde 200 örnek var. Bu sayı bir modeli eğitmek için yeterli değildir.
Ne için uygun:
Yöntemi ve veri biçimini incelemek
Örneklere tek tek bakmak
Az sayıda örnekle deneme yapmak (few-shot)
Kendi ölçüm… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/toolcall-tr-pilot.imagenet-variations-synth-pilot-v4
imagenet-variations-synth-pilot-v4
Synthetic multimodal instruction data. Prompts come from laion/imagenet_variations;
images are generated with FLUX.1-schnell, encoded to 32 SEED-2 tokens per image,
grounded with Florence-2, and turned into USER / ASSISTANT records.
This is v4, rebuilt from v3.1 to cover the review feedback of 2026-08-15.
What is new vs v3.1
Axis
v3.1
v4
Reasoning
none
<think> traces on ~50% of records
Prompts per image
1
1-4 turns… See the full description on the dataset page: https://huggingface.co/datasets/EmpathicRobotics/imagenet-variations-synth-pilot-v4.r1-hw-h2-pilot01-active-trim-10fps
