datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace
TCS Qwen3.6-27B Direction Beam Full-Trace Pilot
Verified compact export for tcs_qwen36_27b_direction_beam_pilot50_full_logging_20260814.
Source dataset: TCS train-00000-of-00001.parquet
Problems: 50
Displayed trajectories: 200
Chunk probes: 1,600
Terminal answers and rubric grades: 6,400
Policy: Qwen/Qwen3.6-27B
Judge: openai/gpt-oss-20b (low reasoning)
At each displayed chunk, the policy proposes four directions plus a null
continuation. A width-four stochastic beam reaches… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/tcs-qwen36-27b-direction-rollouts-pilot-50-full-trace.qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3
Qwen3.6-27B GGUF quantization on a bounded BFCL V4 pilot
Q4_K_M matched Q8_0 on both tested categories: each scored 94 of 100 selected cases correct. Q5_K_M also scored 94/100; Q3_K_M scored 92/100.
Read the results page · Inspect all 400 scored rows
This is a post-result-corrected exploratory analysis of two selected non-live BFCL V4 categories, not a full leaderboard result.
Inspect the scored rows without cloning
The Hub Dataset Viewer does not render this… See the full description on the dataset page: https://huggingface.co/datasets/CyberNative-AI/qwen36-27b-gguf-bfcl-v4-quantization-pilot-corrected-v3.SWE-universe-repaired-bug-pilot-trajectories
SWE-universe repaired BugPilot trajectories
Combined trajectory artifacts for the Qwen3.6 + mini-swe-agent evaluation of VmaxRL/SWEUniverse-Repaired-Bugpilot.
This dataset contains one row per evaluated task in metadata.jsonl, plus per-task files under trajectories//. The combined set uses the main full eval and replaces the two original infra-failure rows with the clean infra rerun trajectories.
Summary:
rows: 804
effective attempts: 804
passes: 629
pass rate: 0.782338
infra… See the full description on the dataset page: https://huggingface.co/datasets/VmaxRL/SWE-universe-repaired-bug-pilot-trajectories.property-pilot-tickets
🏢 PropertyPilot — Maintenance Tickets
A synthetic dataset of 13,725 residential-maintenance tickets written the way real tenants write them — polite, panicked, passive-aggressive, or confused — each paired with operational metadata (category, urgency, assigned contractor, cost, resolution time).
Built for an end-to-end NLP pipeline: triage classification, similar-case retrieval (embeddings + FAISS), and work-order / reply generation.
About this release. Earlier versions of… See the full description on the dataset page: https://huggingface.co/datasets/propertypilot/property-pilot-tickets.qwen3-4b-perfectblend-deepspec-rollout
Qwen3-4B PerfectBlend DeepSpec Rollout
This dataset contains the complete DeepSpec-aligned Qwen3-4B
self-distillation rollout over the filtered PerfectBlend corpus. The seeded
95/5 split is published as separate train and eval splits.
Splits
Split
Conversations
Shards
Path
train
1,349,860
128
data/*.jsonl
eval
71,046
64
eval/*.jsonl
total
1,420,906
192
Data construction
Canonical filtered corpus: 1,420,906 conversations.
Split:… See the full description on the dataset page: https://huggingface.co/datasets/TIE-Pilot/qwen3-4b-perfectblend-deepspec-rollout.fineweb-legal-pilot
⚖️ FineWeb-Legal-Pilot
66.8M words of the finest legal domain data the 🌐 web has to offer.
Repo: GitHub | Report: Technical Report
What is it?
FineWeb-Legal-Pilot is a pilot dataset consiting of 52k high-quality legal documents filtered from the 10-billion-token sample of 🍷 FineWeb.
To enhance FineWeb's utility for legal AI domain adaptation, we draw inspiration from the FineWeb-Edu methodology: creating a legal quality classifier using annotations… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/fineweb-legal-pilot.uldr-v0.1-pilot
Ukrainian Language Decolonization & Reasoning (ULDR) — Pilot Canary Release v0.1
[!IMPORTANT]
Exploratory Pilot / Canary Release (v0.1): This dataset represents an early exploratory pilot canary release (v0.1-pilot) establishing our baseline data pipeline, schema contracts, and directional validation. It is not the final production release (v1.0). The full production release is scheduled for Phase 5.5 following complete evaluation suite assembly, dialect & historical protection… See the full description on the dataset page: https://huggingface.co/datasets/krisztiankoos/uldr-v0.1-pilot.cannabis-fda-extractive-pilot
FDA Cannabis Extractive Experimental Pilot
Experimental, automatically screened, unreviewed draft dataset. This dataset is not medical advice, is not production-ready, and must not be represented as clinician-reviewed, legally cleared, or suitable for patient-facing systems.
This small English conversational dataset was created to test an auditable Gemma 4 fine-tuning pipeline. It contains exact answer passages from captured FDA pages about CBD/cannabis safety, paired with… See the full description on the dataset page: https://huggingface.co/datasets/aznatkoiny/cannabis-fda-extractive-pilot.toolathlon-qwen35-4b-base-pass1-pilot15
Toolathlon Qwen3.5-4B Pass@1 Pilot
Readable trajectories from a 15-task Toolathlon-Verified Pass@1 pilot using
Qwen/Qwen3.5-4B in thinking mode. The model was served on 6 B200 GPUs with
data parallelism; Toolathlon's Docker environments ran on its remote public
evaluation service.
Contents
Repository: violetxi/toolathlon-qwen35-4b-base-pass1-pilot15
Rows: 15 (one row per task)
Reported result: 5/15 (33.33%)
messages: a Viewer-friendly list containing only user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/toolathlon-qwen35-4b-base-pass1-pilot15.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.ch-pilot-rollouts-qwen3.5-9b
C&H Pilot Rollouts — Qwen/Qwen3.5-9B
20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm
corpus (the open-sourced world from harvey-labs
tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM.
Part of an actor-selection pilot for a world-internalization research project: the goal is to
mine agent trajectories into verified fact stores and rewritten likelihood-training targets.
Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.japanese-math-empirical-difficulty-pilot-50k
Japanese Math Empirical Difficulty Pilot 50k
This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset.
It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale.
Current Status
This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.oellm-longctx-tokenized-natural-128k-256k-pilot-v1
OELLM Natural Long-Context Tokenized 128K/256K Expanded Pilot
This dataset contains Megatron-LM tokenized continuation-training artifacts for natural long-context extension experiments at 128K and 256K sequence scales.
This public revision expands the original pilot from 128 to 512 packed examples per source/tier, for 3,072 packed examples total. The raw pack manifest reports approximately 586M source-side estimated tokens across all six source/tier shards.
Accessible source… See the full description on the dataset page: https://huggingface.co/datasets/birgermoell/oellm-longctx-tokenized-natural-128k-256k-pilot-v1.amalia-pilot-honesty-v2
AMALIA pilot — honesty vector datasets (v1 refusals + v2 corrective mix)
Training data from the first two iterations of a verifier-gated fine-tuning
pilot on AMALIA-9B-0626-DPO,
targeting identity/fact confabulation (the model's weakest measured behavior:
43.3% on our honesty harness). Full methodology, harness, and reports:
github.com/teex-pt/pt-amalia.
These are research pilot artifacts — small by design (the pilot validates
the loop, not the scale). Every sample was produced… See the full description on the dataset page: https://huggingface.co/datasets/teex-pt/amalia-pilot-honesty-v2.toolcall-tr-pilot
ToolCall-TR — Pilot v0.1
Türkçe, execution-verified tool-calling veri seti. Her tool çağrısı gerçek
bir implementasyon tarafından çalıştırıldı ve çıktısı doğrulandı; hiçbir tool
observation'ı bir LLM tarafından yazılmadı.
⚠️ Önce şunu bilin: bu küçük bir ilk sürüm
İçinde 200 örnek var. Bu sayı bir modeli eğitmek için yeterli değildir.
Ne için uygun:
Yöntemi ve veri biçimini incelemek
Örneklere tek tek bakmak
Az sayıda örnekle deneme yapmak (few-shot)
Kendi ölçüm… See the full description on the dataset page: https://huggingface.co/datasets/bilalabic/toolcall-tr-pilot.IVR-pilot-benchmark
IntentSpec Benchmark — Data Supplement
This archive contains the benchmark data used to compute Intent Violation
Rate (IVR) in the paper: 49 tasks, each derived from a HumanEval problem and
extended with an ambiguous/gold prompt pair and a decomposed set of
executable constraints.
Files
spec_pairs.jsonl
The benchmark itself — one JSON object per line, one line per task. This is
the file consumed directly by the evaluation pipeline… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/IVR-pilot-benchmark.pilot-readback-corpus
Pilot Readback Corpus
A synthetic reference corpus of Air Traffic Control (ATC) pilot–controller exchanges in standard ICAO phraseology, intended as ground truth for evaluating the output of LLM-based pilot agents.
⚠️ Synthetic data. All exchanges are author-written following ICAO Doc 4444 / Annex 10 phraseology. They are not transcriptions of real ATC recordings and contain no real flight identifiers, clearances, or operational data.
Overview
The corpus covers… See the full description on the dataset page: https://huggingface.co/datasets/santiisoutoo/pilot-readback-corpus.r16-behavioral-metamerism-pilot
R16 Behavioral Metamerism Pilot
Brand Function x synthetic cohort interaction experiment from the Spectral Brand Theory research program.
Dataset Summary
675 API calls testing whether Brand Function specification differentially affects dimensional collapse across synthetic observer cohorts. Design: 5 cohorts x 5 brands x 3 conditions (no BF, structural BF, enriched BF) x 3 models x 3 repetitions.
Companion paper: AI-Native Brand Identity: From Visual Recognition… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/r16-behavioral-metamerism-pilot.oasst1-zh-pilot
oasst1 Chinese Translation Pilot (10 samples)
This is a pilot release of 10 parallel English→Chinese samples translated from
OpenAssistant/oasst1. It is
intended as a methodology demonstration and quality evaluation artifact, not as
a training-ready dataset.
Why this exists
We are evaluating whether LLM-assisted translation of open instruction-tuning datasets
into low-resource languages can be done at a quality bar that the ML community will
accept. Chinese is our first… See the full description on the dataset page: https://huggingface.co/datasets/AgenticCommons/oasst1-zh-pilot.self_repair_parsing_pilot_data
Self-Repair SBN Parsing Pilot Data
A small, static pilot set for studying how a semantic parser (NL → SBN) behaves when
the input natural language carries self-repair disfluencies. Each row is one gold
sentence from the Parallel Meaning Bank (PMB) together with synthetically generated
repair variants. The dataset holds inputs only — parser predictions are stored
elsewhere so this set stays fixed across modelling stages.
Source: PMB 5.1.0 (gold), English single-sentence items.… See the full description on the dataset page: https://huggingface.co/datasets/Shrikes/self_repair_parsing_pilot_data.
