datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
codeparrot-clean-train
CodeParrot 🦜 Dataset Cleaned (train)
Train split of CodeParrot 🦜 Dataset Cleaned.
Dataset structure
DatasetDict({
train: Dataset({
features: ['repo_name', 'path', 'copies', 'size', 'content', 'license', 'hash', 'line_mean', 'line_max', 'alpha_frac', 'autogenerated'],
num_rows: 5300000
})
})
Inter-Edit-Train
Inter-Edit-Train
Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing.
This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance.
Highlights
1,099,964 image editing pairs
610,186 unique source images
Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.OLA-Embed-TrainingNemotron-RL-Ultra-Training-Blends
Dataset Description:
This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used.
The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.c4-en-html-with-training_metadata_allrlbenchfail_train_dataset
Guardian: RLBench-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.fast-autoregressive-inference-gp-trainK4midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.hywe-training-data
HYWE Spatial Configuration Dataset
A structured corpus of procedural architectural programming, design intent, and deterministic spatial configurations.
The dataset is generated through the HYWE Core Engine, a dependency-free computational core for discrete spatial representation and deterministic topological resolution. Design-intent narratives and HYWE Syntax representations are captured through the HYWE ecosystem and structured by the Hynteract data pipeline.
The… See the full description on the dataset page: https://huggingface.co/datasets/vykrum/hywe-training-data.AXXXX_jssp_policy_step_train_dispatch_v1qwen35-4b-drpo-vs0f49th-trainer-logprobs
Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th
This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587).
Contents
Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th
Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/
Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl
Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train
SAND-Post-Training-Dataset
SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs
Dataset Summary
We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs.
This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.3da-libero-training-assets
GAM LIBERO Training Assets
This dataset repository contains the assets needed to fine-tune GAM on LIBERO.
Layout
checkpoints/track4world_da3.pth
data/libero_noop/<suite>/*.hdf5
data/libero_noop/_stats/*.json
configs/training/libero_unified/
track4world_da3.pth is the DA3-Giant base checkpoint. The LIBERO HDF5 files contain embedded RGB, proprioception, actions, and depth used by the public GAM training configs.
Code:… See the full description on the dataset page: https://huggingface.co/datasets/SeonghuJeon/3da-libero-training-assets.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.bdv2fail_train_dataset
Guardian: BridgeDataV2-Fail Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.LRAT-Train
LRAT Training Dataset
This dataset contains trajectory-derived retrieval supervision used in LRAT (Learning to Retrieve from Agent Trajectories).
The dataset is designed for training retrievers for agentic search. Instead of relying on human click logs, it is built from deep research agent trajectories that record intermediate search queries, browsing actions, and post-browse reasoning traces.
Dataset Summary
LRAT formalizes a simple idea: when search is increasingly… See the full description on the dataset page: https://huggingface.co/datasets/Yuqi-Zhou/LRAT-Train.llama-3.1-medprm-reward-training-set
Med-PRM-Reward (Version 1.0)
🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.ClimbMix-Ja-Initial64-Training-Data
ClimbMix-Ja Initial64 350M Artifacts
This repository is a public backup for the initial 64 ClimbMix-Ja candidate runs.
Candidate count: 64
Base model: nvidia/nemotron-climb-proxy-models 350M converted to a Megatron-LM TE-compatible checkpoint
Training corpus: KantaHayashiAI/ClimbLab-Ja clustered into cluster_01 ... cluster_20
Sequence length: 1024
Train iterations per candidate: 6500
Global batch size: 304
Tokens per candidate: 2,023,424,000
Total trained tokens across… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbMix-Ja-Initial64-Training-Data.patina3-v3-training-data
PATINA-3 v3 — training data backup (updated 2026-09-08)
Inputs of the PATINA-3 experiment (design of record: patina3/SPEC.md v3 in
github.com/Jordine/entanglement_engineering). Question: does a spec's explanation have to be TRUE,
or only STATED, for its value to generalize? Base: Llama-3.1-8B; upstream: Model Spec Midtraining
(arXiv 2605.02087).
template/tmpl_afford_full_r9_1.jsonl — the 4,600 skeletons every corpus is instantiated from;
template/excluded_skeletons_r9_1_v3.json… See the full description on the dataset page: https://huggingface.co/datasets/Jordine/patina3-v3-training-data.ur5fail_train_dataset
Guardian Failure Detection Dataset
This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks.
Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.atomic-metrics-demographic-training-size
Atomic Metrics: Demographic Training-Size Analysis
Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction.
Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.ronin_trainspatial-training-full-release-20260604
Spatial Training Full Data Release
Full data staging directory for our current Cambrian-P / SSR-style Spatial VLM reproduction work.
The directory contains the lightweight reproduction pack plus raw compressed training archives. Local staging uses hardlinks where possible, but upload payload is the full dataset.
Size
Logical payload: 1076.281 GiB
Files: 254
Max single file: 18.0 GiB
VSI-590K raw payload: 216.777 GiB
Cambrian-S-3M raw payload: 858.004 GiB… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/spatial-training-full-release-20260604.MirrorAPI-Training
MirrorAPI training dataset
This dataset contains the training data for MirrorAPI and MirrorAPI-Cache:
train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI .
train_cache.json: The training data for MirrorAPI-Cache.
k3-sft-cc0-flan
Dataset Card for K3 SFT CC0 FLAN
844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain
FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort
views; adaptive is the recommended default for quality-conscious SFT mixing.
Dataset Details
Curated by: Training Datasmith
Teacher: kimi-k3 via deltafin (local inference)
Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.pacman_hard_cot_chunk_k10_train
pacman_hard_cot_chunk_k10_train
BAGEL VLM-Gym world-model dataset (pacman / cot).
CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps.
layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline.
images are base64-encoded JPEG frames stored inline in each JSONL row.
Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.ridgelora-stage2-imposebase-train160-50k-20260824
Stage-2 ControlNet retraining with the frozen IMPOSE base
This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE
checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM
t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices).
An initial Stage-1-from-scratch job was stopped at step 575 after correcting
the scope. It produced no scheduled checkpoint and is not used in any result;
its log is retained only as an audit trail.
Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.fast-autoregressive-inference-gp-trainK16agenttool-training-garden
AgentTool HF Training Garden
A tiny metadata-only companion for designing a reproducible Hugging Face data
lifecycle without treating the Hub, a Dataset Card, or one quality score as
training authority.
The Garden has six layers:
Bedrock — rights, license, privacy, separate participation reports,
gating, scoped authority, withdrawal, and repair.
Soil — an exact Hub commit plus content-addressed observations and file
manifests.
Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.
