datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zscreen-pilot
Z-Screen pilot: chemical recipes and cellular responses
Version 2.0.0
Z-Screen connects combinatorial chemistry to high-dimensional cellular measurements. This pilot data resource contains 190,699 compound-context profiles from 162,914 public chemical recipes across eight library–cell contexts. Each RNA response is available on a 6,000-gene panel and as coordinates on 32 shared transcriptional programs.
The package provides processed data, aligned identifiers, reference model… See the full description on the dataset page: https://huggingface.co/datasets/Zafrens/zscreen-pilot.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.carbon-pilot-corpusproperty-pilot-tickets
🏢 PropertyPilot — Maintenance Tickets
A synthetic dataset of 13,725 residential-maintenance tickets written the way real tenants write them — polite, panicked, passive-aggressive, or confused — each paired with operational metadata (category, urgency, assigned contractor, cost, resolution time).
Built for an end-to-end NLP pipeline: triage classification, similar-case retrieval (embeddings + FAISS), and work-order / reply generation.
About this release. Earlier versions of… See the full description on the dataset page: https://huggingface.co/datasets/propertypilot/property-pilot-tickets.SWE-ZERO-V2PRs-1k-pilotcarrot-to-plate-pilot3This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so_follower",
"total_episodes": 5,
"total_frames": 4256,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/choiwoong/carrot-to-plate-pilot3.r1-d002-number-pointing-pilot-20260908
R1 D002 number-pointing pilot
Private engineering pilot converted to LeRobot Dataset v3.0 from the accepted
episodes of run d002_20260908T020005Z.
This upload is for validating the conversion, Hub viewer, download, and smoke
training workflow. It is not a production training dataset and makes no
hardware-readiness claim.
Contents
7 episodes, 733 frames, 8 FPS
one 640×480 simulated head-camera stream
Unitree R1 A5 arm state (10,) and arm action (10,)
per-episode… See the full description on the dataset page: https://huggingface.co/datasets/vasco281204/r1-d002-number-pointing-pilot-20260908.loopwan-opensora-pilot-v1
LoopWan Open-Sora-Plan pilot
Status: completed bounded curation. Counts: {"long_audit": 22, "train": 2000, "val": 128}.
Fixed 320x480, timestamp sampling at 16 FPS; train/validation crops are real
contiguous 10-second shots, audit crops 20 seconds. Sources are disjoint and
captions are matched to pinned official annotations. See DATASET_REPORT.md for
filter thresholds, caption limitations and full provenance.
Official dataset revision: ab77293def393e6938f11a7bfd12163decfb9620.… See the full description on the dataset page: https://huggingface.co/datasets/Nicholas0228/loopwan-opensora-pilot-v1.docclass-pilot
Docclass Pilot Sample
Deterministic stratified pilot slice of
Lucius-Morningstar/docclass-merged v5
(KANBAN-084, built 2026-08-24): every doc type and every subclass stratum
present in the parent contributes at least one row; each stratum contributes
at most 3 rows, selected by ascending sha256(filename) within the
stratum (content-addressed order — decorrelated from the md5-keyed family
split rule; rebuilds are byte-identical).
Coverage
doc_type
rows… See the full description on the dataset page: https://huggingface.co/datasets/Lucius-Morningstar/docclass-pilot.muse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.instruction-pilot-outputs-filteredpilot-corpus-v0.1
Pilot Corpus v0.1
A frozen, pointer-based miniature of a robot-learning pre-training mixture:
7 tranches spanning teleop, UMI, egocentric dexterous manipulation, ego/exo
wild video, internet action video, and sim — with every label type
(vision / proprio / action / force / pose / lang) and every domain tag
represented, sized to run scratch-vs-init experiments on a single GPU.
What this repo is
Manifests + code, not media. A clip is a virtual (episode, t_start… See the full description on the dataset page: https://huggingface.co/datasets/atharva-pantheon/pilot-corpus-v0.1.fineweb-legal-pilot
⚖️ FineWeb-Legal-Pilot
66.8M words of the finest legal domain data the 🌐 web has to offer.
Repo: GitHub | Report: Technical Report
What is it?
FineWeb-Legal-Pilot is a pilot dataset consiting of 52k high-quality legal documents filtered from the 10-billion-token sample of 🍷 FineWeb.
To enhance FineWeb's utility for legal AI domain adaptation, we draw inspiration from the FineWeb-Edu methodology: creating a legal quality classifier using annotations… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/fineweb-legal-pilot.ctboost-tabarena-lite-pilot-0.1.59
CTBoost 0.1.59 validation-only pilot
No task family met the frozen gate; this pilot does not justify the proposed full HPO rerun.
All 88 predeclared fits are accounted for across 14 datasets. Status counts: ok: 88.
Only official outer-training rows were used. Two fixed inner folds supplied validation scores and early stopping. Outer-test scores and Elo were neither computed nor used for selection. This is an author-run development pilot on noncanonical hardware, not a TabArena… See the full description on the dataset page: https://huggingface.co/datasets/Maiernator/ctboost-tabarena-lite-pilot-0.1.59.omx_pick_place_pilot_alcan_open_loop_20260918_150040This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ericc430/omx_pick_place_pilot_alcan_open_loop_20260918_150040.omx_pick_place_pilot_plastic_bottle_open_loop_20260918_152143This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ericc430/omx_pick_place_pilot_plastic_bottle_open_loop_20260918_152143.omx_pick_place_pilot_plastic_bottle_open_loopThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ericc430/omx_pick_place_pilot_plastic_bottle_open_loop.toolathlon-qwen35-4b-base-pass1-pilot15
Toolathlon Qwen3.5-4B Pass@1 Pilot
Readable trajectories from a 15-task Toolathlon-Verified Pass@1 pilot using
Qwen/Qwen3.5-4B in thinking mode. The model was served on 6 B200 GPUs with
data parallelism; Toolathlon's Docker environments ran on its remote public
evaluation service.
Contents
Repository: violetxi/toolathlon-qwen35-4b-base-pass1-pilot15
Rows: 15 (one row per task)
Reported result: 5/15 (33.33%)
messages: a Viewer-friendly list containing only user… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/toolathlon-qwen35-4b-base-pass1-pilot15.stage3-real-expansion-agent-teacher-separated-pilot
Teacher-Separated Expansion Agent Pilot
A 10-task inspection batch generated by Qwen3-235B-A22B-Instruct-2507 from real
CLAPNQ, PubMedQA, MAUD, ContractNLI, and FinQA source tasks.
The teacher-only trajectory-generation system prompt is recorded in
metadata/generation-manifest.json for auditability, but is absent from every
saved training trajectory. Each final messages list begins with the real
memory-wrapped task user message, followed by native assistant expand calls,
exact… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-real-expansion-agent-teacher-separated-pilot.omx_pick_place_pilot_60_alcan_20260918_100719This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ericc430/omx_pick_place_pilot_60_alcan_20260918_100719.omx_pick_place_pilot_60_alcan_t50_20260918_115400This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 15,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/ericc430/omx_pick_place_pilot_60_alcan_t50_20260918_115400.fr3-leader-command-pilot
FR3 Leader-Command Pilot
Four FR3 + GELLO teleoperation episodes recorded with the corrected action convention:
action is the GELLO leader's commanded joint position, following ACT/ALOHA. This is a
pilot run verifying the fix before full-scale collection — not a training set.
Contents
Robot / teleop
Franka Research 3 + GELLO leader arm
Task
pick up the skyblue cup and place it on the yellow bowl
Episodes / frames
4 / 1092 (319, 208, 347, 218)… See the full description on the dataset page: https://huggingface.co/datasets/knu-physical-ai/fr3-leader-command-pilot.ch-pilot-rollouts-qwen3.5-9b
C&H Pilot Rollouts — Qwen/Qwen3.5-9B
20 agentic exploration rollouts over the Calderwood & Harkness (C&H) synthetic law-firm
corpus (the open-sourced world from harvey-labs
tasks/firm-knowledge/, MIT), generated by Qwen/Qwen3.5-9B served with vLLM.
Part of an actor-selection pilot for a world-internalization research project: the goal is to
mine agent trajectories into verified fact stores and rewritten likelihood-training targets.
Companion dataset (same seeds/tasks, different… See the full description on the dataset page: https://huggingface.co/datasets/violetxi/ch-pilot-rollouts-qwen3.5-9b.fr3-scene-pilot-lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
8
],
"names": [
"joint1.pos",
"joint2.pos",
"joint3.pos",
"joint4.pos",
"joint5.pos",
"joint6.pos"… See the full description on the dataset page: https://huggingface.co/datasets/knu-physical-ai/fr3-scene-pilot-lerobot.japanese-math-empirical-difficulty-pilot-50k
Japanese Math Empirical Difficulty Pilot 50k
This dataset is a 50,000-problem empirical difficulty pilot, not a full empirical labeling of the original 5.66M-row source dataset.
It was created for LLM-jp experiment 0399, Team Victory SFT, to validate empirical difficulty label distribution, downstream split behavior, and the rollout/scoring pipeline before attempting labeling at the full 5.6M scale.
Current Status
This upload uses the v3 scorer with assistant-only… See the full description on the dataset page: https://huggingface.co/datasets/argo11/japanese-math-empirical-difficulty-pilot-50k.lsec-sicilian-lonp2-pilot-results-v1
lsec-sicilian-lonp2-pilot-results-v1
Pilot results (5 samples) from re-detecting splice junctions with SICILIAN
(Dehghannasiri, Olivieri, Salzman -- Genome Biology 2021) on raw FASTQ, to test whether the
LONP2 skipped-exon (SE) event's null result (3x via STARsolo -> JAseC/SHIBA) is a methods
artifact. See experiment lsec-sicilian-lonp2 for the full design, red-team review, and
debugging history.
The LONP2 SE event (hg38)
5' constitutive exon:… See the full description on the dataset page: https://huggingface.co/datasets/depinwang/lsec-sicilian-lonp2-pilot-results-v1.egovista-pilot-v0.2
EgoVista Pilot v0.2
First-person (egocentric) annotations dataset of human daily-life manipulation activities, generated by the EgoVista pipeline.
What is in this dataset
This is an annotations-only release. RGB video frames are not included. Per-frame annotations cover:
2D pose keypoints (full body)
2D and 3D hand keypoints
16-bit depth maps (Depth Anything V2)
Hand and object segmentation masks (EgoHOS, RLE COCO)
Contact flags (left/right/both hands)
Action labels… See the full description on the dataset page: https://huggingface.co/datasets/Leodocr/egovista-pilot-v0.2.so101_vials_pilot_20260720_151630This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/samla2000/so101_vials_pilot_20260720_151630.omx-green-block-taped-area-pilot-v1_20260812_043727This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/mongdmin/omx-green-block-taped-area-pilot-v1_20260812_043727.so101_pen_to_cup_pilot_20260921_203436This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 25,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/Stands99/so101_pen_to_cup_pilot_20260921_203436.
