CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01codeparrot /codeparrot-clean-train CodeParrot 🦜 Dataset Cleaned (train) Train split of CodeParrot 🦜 Dataset Cleaned. Dataset structure DatasetDict({ train: Dataset({ features: ['repo_name', 'path', 'copies', 'size', 'content', 'license', 'hash', 'line_mean', 'line_max', 'alpha_frac', 'autogenerated'], num_rows: 5300000 }) }) tabular1M<n<10M16 likes3.7k downloads4y agoHugging Face02a1557811266 /Inter-Edit-Train Inter-Edit-Train Inter-Edit-Train is the official large-scale training set released for the CVPR 2026 paper Inter-Edit: First Benchmark for Interactive Instruction-Based Image Editing. This dataset is designed for the Interactive Instruction-based Image Editing (I^3E) task, where a model performs localized image edits from a concise textual instruction together with imprecise spatial guidance. Highlights 1,099,964 image editing pairs 610,186 unique source images Four… See the full description on the dataset page: https://huggingface.co/datasets/a1557811266/Inter-Edit-Train.tabularimage-to-image1M<n<10M1 likes3.4k downloads6mo agoHugging Face03OLAIR /OLA-Embed-Trainingtabular1B<n<10B0 likes2.1k downloads4mo agoHugging Face04nvidia /Nemotron-RL-Ultra-Training-Blends Dataset Description: This dataset provides Reinforcement Learning (RL) and Multi-teacher On-Policy Distillation (MOPD) training-data blends used by the public Nemotron-3-Ultra post-training recipe. The blends are consumed by the NeMo RL training recipes through the NeMo Gym agent framework, in which each prompt is paired with an agent/environment that returns a verifiable or judge-based reward. Each subset is a separate blend; see the recipe for how the blends are used. The… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Ultra-Training-Blends.tabulartext-generation10K<n<100K19 likes1.4k downloads2mo agoHugging Face05bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes990 downloads3y agoHugging Face06paulpacaud /rlbenchfail_train_dataset Guardian: RLBench-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data generated in the RLBench simulator for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that procedurally perturbs successful scripted trajectories in simulation, generating diverse planning… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/rlbenchfail_train_dataset.tabularvisual-question-answering10K<n<100K0 likes968 downloads7mo agoHugging Face07conorhassan /fast-autoregressive-inference-gp-trainK4tabularn<1K0 likes701 downloads1y agoHugging Face08ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes634 downloads1mo agoHugging Face09vykrum /hywe-training-data HYWE Spatial Configuration Dataset A structured corpus of procedural architectural programming, design intent, and deterministic spatial configurations. The dataset is generated through the HYWE Core Engine, a dependency-free computational core for discrete spatial representation and deterministic topological resolution. Design-intent narratives and HYWE Syntax representations are captured through the HYWE ecosystem and structured by the Hynteract data pipeline. The… See the full description on the dataset page: https://huggingface.co/datasets/vykrum/hywe-training-data.tabularn<1K1 likes311 downloads15d agoHugging Face10HYUNJINI /AXXXX_jssp_policy_step_train_dispatch_v1tabular1M<n<10M0 likes301 downloads6mo agoHugging Face11hamishivi /qwen35-4b-drpo-vs0f49th-trainer-logprobs Qwen3.5 4B DRPO trainer logprobs from W&B run vs0f49th This dataset contains the raw trainer-logprob JSONL shards saved by W&B run ai2-llm/open_instruct_internal/vs0f49th (qwen35_4b_drpo__42__1782345587). Contents Source run: https://wandb.ai/ai2-llm/open_instruct_internal/runs/vs0f49th Source path: /weka/oe-adapt-default/allennlp/deletable_rollouts/ Filename pattern: qwen35_4b_drpo__42__1782345587_trainer_logprobs_step*_rank*.jsonl Files: 4320 JSONL shards… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/qwen35-4b-drpo-vs0f49th-trainer-logprobs.tabulartext-generation10K<n<100K0 likes281 downloads3mo agoHugging Face12dongboklee /MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train MMLU-Pro_Llama-3.1-8B-Instruct_gORM_train tabular100K<n<1M0 likes212 downloads3mo agoHugging Face13amd /SAND-Post-Training-Dataset SAND-Post-Training-Dataset: High-Quality Synthetic Reasoning Dataset Built with AMD GPUs Dataset Summary We introduce the SAND-Post-Training-Dataset, a high-quality synthetic reasoning dataset for mathematics and science built entirely using a synthetic data pipeline running on the AMD ROCm™ stack and AMD Instinct™ MI325 GPUs. This dataset prioritizes difficulty and novelty over volume, demonstrating that high-difficulty synthetic data can elevate… See the full description on the dataset page: https://huggingface.co/datasets/amd/SAND-Post-Training-Dataset.tabularquestion-answering10K<n<100K3 likes181 downloads10mo agoHugging Face14SeonghuJeon /3da-libero-training-assets GAM LIBERO Training Assets This dataset repository contains the assets needed to fine-tune GAM on LIBERO. Layout checkpoints/track4world_da3.pth data/libero_noop/<suite>/*.hdf5 data/libero_noop/_stats/*.json configs/training/libero_unified/ track4world_da3.pth is the DA3-Giant base checkpoint. The LIBERO HDF5 files contain embedded RGB, proprioception, actions, and depth used by the public GAM training configs. Code:… See the full description on the dataset page: https://huggingface.co/datasets/SeonghuJeon/3da-libero-training-assets.tabularn<1K0 likes181 downloads3mo agoHugging Face15danielrosehill /Tech-Sentences-For-ASR-Training TechVoice Dataset Work in Progress – This dataset is actively being expanded with new recordings. Dataset Statistics Metric Current Target Progress Duration 38m 43s 5h 0m 0s ██░░░░░░░░░░░░░░░░░░ 12.9% Words 10,412 50,000 ████░░░░░░░░░░░░░░░░ 20.8% Total Recordings: 205 samples Total Characters: 74,312 A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.audioautomatic-speech-recognitionn<1K2 likes142 downloads10mo agoHugging Face16paulpacaud /bdv2fail_train_dataset Guardian: BridgeDataV2-Fail Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data derived from the BridgeDataV2 real-robot dataset for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Failures are produced by an automated pipeline that perturbs successful real-robot trajectories offline (without re-executing actions)… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/bdv2fail_train_dataset.tabularvisual-question-answering10K<n<100K1 likes141 downloads7mo agoHugging Face17Yuqi-Zhou /LRAT-Train LRAT Training Dataset This dataset contains trajectory-derived retrieval supervision used in LRAT (Learning to Retrieve from Agent Trajectories). The dataset is designed for training retrievers for agentic search. Instead of relying on human click logs, it is built from deep research agent trajectories that record intermediate search queries, browsing actions, and post-browse reasoning traces. Dataset Summary LRAT formalizes a simple idea: when search is increasingly… See the full description on the dataset page: https://huggingface.co/datasets/Yuqi-Zhou/LRAT-Train.tabulartext-retrieval10K<n<100K6 likes134 downloads6mo agoHugging Face18dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes125 downloads1y agoHugging Face19KantaHayashiAI /ClimbMix-Ja-Initial64-Training-Data ClimbMix-Ja Initial64 350M Artifacts This repository is a public backup for the initial 64 ClimbMix-Ja candidate runs. Candidate count: 64 Base model: nvidia/nemotron-climb-proxy-models 350M converted to a Megatron-LM TE-compatible checkpoint Training corpus: KantaHayashiAI/ClimbLab-Ja clustered into cluster_01 ... cluster_20 Sequence length: 1024 Train iterations per candidate: 6500 Global batch size: 304 Tokens per candidate: 2,023,424,000 Total trained tokens across… See the full description on the dataset page: https://huggingface.co/datasets/KantaHayashiAI/ClimbMix-Ja-Initial64-Training-Data.tabularn<1K0 likes125 downloads4mo agoHugging Face20Jordine /patina3-v3-training-data PATINA-3 v3 — training data backup (updated 2026-09-08) Inputs of the PATINA-3 experiment (design of record: patina3/SPEC.md v3 in github.com/Jordine/entanglement_engineering). Question: does a spec's explanation have to be TRUE, or only STATED, for its value to generalize? Base: Llama-3.1-8B; upstream: Model Spec Midtraining (arXiv 2605.02087). template/tmpl_afford_full_r9_1.jsonl — the 4,600 skeletons every corpus is instantiated from; template/excluded_skeletons_r9_1_v3.json… See the full description on the dataset page: https://huggingface.co/datasets/Jordine/patina3-v3-training-data.tabular10K<n<100K0 likes123 downloads15d agoHugging Face21paulpacaud /ur5fail_train_dataset Guardian Failure Detection Dataset This dataset is part of the Guardian project: Detecting Robotic Planning and Execution Errors with Vision-Language Models. It contains annotated robotic manipulation failure data for training and evaluating Vision-Language Models (VLMs) on failure detection tasks. Guardian introduces an automated failure generation approach that procedurally perturbs successful robot trajectories to produce diverse planning failures and execution failures, each… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/ur5fail_train_dataset.tabularvisual-question-answering1K<n<10K0 likes121 downloads7mo agoHugging Face22tintin1027 /atomic-metrics-demographic-training-size Atomic Metrics: Demographic Training-Size Analysis Complete offline reproduction bundle for the effect of batch-selected training size on demographic preference prediction. Version 2 — replaces the fixed-bank analysis. Select k extraction batches (five pairs each), use only their metrics and their 5k training pairs to refit BT/LR, then evaluate on cached test200 scores restricted to those metrics. Both the training rows and metric columns change with size. Extraction/refinement… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/atomic-metrics-demographic-training-size.tabular1K<n<10K0 likes118 downloads4d agoHugging Face23mikaelnias /ronin_traintabularn<1K0 likes108 downloads3y agoHugging Face24Gradygu3u /spatial-training-full-release-20260604 Spatial Training Full Data Release Full data staging directory for our current Cambrian-P / SSR-style Spatial VLM reproduction work. The directory contains the lightweight reproduction pack plus raw compressed training archives. Local staging uses hardlinks where possible, but upload payload is the full dataset. Size Logical payload: 1076.281 GiB Files: 254 Max single file: 18.0 GiB VSI-590K raw payload: 216.777 GiB Cambrian-S-3M raw payload: 858.004 GiB… See the full description on the dataset page: https://huggingface.co/datasets/Gradygu3u/spatial-training-full-release-20260604.tabularvisual-question-answering1M<n<10M0 likes107 downloads3mo agoHugging Face25stabletoolbench /MirrorAPI-Training MirrorAPI training dataset This dataset contains the training data for MirrorAPI and MirrorAPI-Cache: train_sft.json, train_cot.json, train_augment.json: The training data for MirrorAPI . train_cache.json: The training data for MirrorAPI-Cache. tabular100K<n<1M1 likes104 downloads2y agoHugging Face26Training-Datasmith /k3-sft-cc0-flan Dataset Card for K3 SFT CC0 FLAN 844-row Kimi K3 synthetic instruction-tuning shard built from DPI-traced CC0/public-domain FLAN prompts in the Tülu mix. Four overlapping Hub configs expose different cohort views; adaptive is the recommended default for quality-conscious SFT mixing. Dataset Details Curated by: Training Datasmith Teacher: kimi-k3 via deltafin (local inference) Languages: English prompts; translation pairs include German, Spanish, Czech, Igbo… See the full description on the dataset page: https://huggingface.co/datasets/Training-Datasmith/k3-sft-cc0-flan.tabulartext-classification1K<n<10K0 likes103 downloads4d agoHugging Face27novastar111 /pacman_hard_cot_chunk_k10_train pacman_hard_cot_chunk_k10_train BAGEL VLM-Gym world-model dataset (pacman / cot). CoT chunk-K train set: all-step interleaved imagined reasoning; re-grounds on the true frame every K=10 steps. layout: Train-only. Gzipped-JSONL shards under training/; each row is one packed SFT sample with base64-JPEG frames inline. images are base64-encoded JPEG frames stored inline in each JSONL row. Pairs with the matching pacman checkpoint(s) under the companion model org; CoT and non-CoT… See the full description on the dataset page: https://huggingface.co/datasets/novastar111/pacman_hard_cot_chunk_k10_train.tabular100K<n<1M0 likes101 downloads1mo agoHugging Face28LamTNguyen /ridgelora-stage2-imposebase-train160-50k-20260824 Stage-2 ControlNet retraining with the frozen IMPOSE base This experiment retrains only Stage 2 for RidgeLoRA-FP. Stage 1 is the IMPOSE checkpoint and is not retrained. The run started on 2026-08-24 on TPU VM t1v-n-d3df3356-w-0 (TPU v5p-8, four XLA devices). An initial Stage-1-from-scratch job was stopped at step 575 after correcting the scope. It produced no scheduled checkpoint and is not used in any result; its log is retained only as an audit trail. Frozen IMPOSE… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-stage2-imposebase-train160-50k-20260824.image1K<n<10K0 likes101 downloads29d agoHugging Face29conorhassan /fast-autoregressive-inference-gp-trainK16tabularn<1K0 likes95 downloads1y agoHugging Face30Yu-and-Ai /agenttool-training-garden AgentTool HF Training Garden A tiny metadata-only companion for designing a reproducible Hugging Face data lifecycle without treating the Hub, a Dataset Card, or one quality score as training authority. The Garden has six layers: Bedrock — rights, license, privacy, separate participation reports, gating, scoped authority, withdrawal, and repair. Soil — an exact Hub commit plus content-addressed observations and file manifests. Roots — acquisition, parsing, filtering, secret… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/agenttool-training-garden.textn<1K0 likes92 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.