CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HuggingFaceCode /stack-v3-train 🥞 The Stack v3 What is it? What is being released How to download and use it Dataset statistics Dataset structure Dataset creation Considerations for using the data Additional information What is it? The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceCode/stack-v3-train.tabulartext-generation100M<n<1B382 likes183k downloads1d agoHugging Face02cadene /agibot_alpha_v30This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "AgiBot_A2D", "total_episodes": 28122, "total_frames": 47613574, "total_tasks": 30, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:28122" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/cadene/agibot_alpha_v30.tabularrobotics10M<n<100M2 likes69k downloads1y agoHugging Face03IFM /TxT360-v2 TxT360-v2 Dataset Description Pre-training sources for the K2 Horizon training data release. This repository is part of the K2 Horizon collection. The repository is organized into multiple subsets. Every subset has a train split backed by Parquet shards. K2 Horizon Dataset Series Dataset repository Focus Subsets IFM/TxT360-v2 Web and question-answering text 3 IFM/Code-Reasoning Code reasoning and task synthesis 7 IFM/Math-Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/IFM/TxT360-v2.tabulartext-generation1B<n<10B81 likes24k downloads4d agoHugging Face04flex-pi /robotwin_3d RoboTwin 2.0 — 3D (RGB + Depth) Bimanual manipulation data from the RoboTwin 2.0 simulator, in LeRobot v2.1 format, with per-camera ground-truth depth alongside RGB. Tasks 50 Episodes 27,500 (550 per task, contiguous) Frames 6,183,813 Robot ALOHA-style bimanual, 14-DoF Control rate 50 Hz Cameras 3 (cam_high, cam_left_wrist, cam_right_wrist) Resolution 240 × 320 Language instructions 1,039,891 unique corpus-wide; 100 entries per episode Total size… See the full description on the dataset page: https://huggingface.co/datasets/flex-pi/robotwin_3d.tabularrobotics1M<n<10M0 likes23k downloads1mo agoHugging Face05erl-hub /behaviour1k-Qwen3-features BEHAVIOR-1K Qwen3 skill features Per-frame conditioned features e_t = Phi(f_t, L_sub^(j), L), mean-pooled primitive skill latents S_j, aligned proprioception q_t, actions a_t, and subtask progress p_t. These are the inputs and targets for a Primitive Skill Composer VLA Skill Predictor. Ground-truth primitives come from BEHAVIOR-1K's hand-authored primitive_annotation, so the segmentation is human-labelled rather than predicted, and nothing here depends on a keyframe detector.… See the full description on the dataset page: https://huggingface.co/datasets/erl-hub/behaviour1k-Qwen3-features.tabularrobotics1K<n<10K0 likes15k downloads2mo agoHugging Face06Hahshshsshbs /Wan2.2-Syn-121x704x1280_32k FastVideo Synthetic Wan2.2 720P dataset FastVideo Team  Paper | Github | Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k.tabulartext-to-video10K<n<100K1 likes12k downloads3mo agoHugging Face07ember-lab-berkeley /robocasa365-pretrain-mg Pretraining (MimicGen) — atomic MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted augmentation from human demonstrations. Part of the RoboCasa365 collection. Flat LeRobot v3.0 mirror of RoboCasa365 — standard layout, drop-in loadable. Stats Episodes: 536,030 Frames: 116,246,439 (20 fps → 1615 h) Tasks: 720 (natural-language phrasings; underlying RoboCasa task classes: 60) Cameras: 3 × 256×256 h264… See the full description on the dataset page: https://huggingface.co/datasets/ember-lab-berkeley/robocasa365-pretrain-mg.tabularrobotics100M<n<1B2 likes11k downloads5mo agoHugging Face08lvogel123 /jailbreak-deepseek-v3.2-exptabular1K<n<10K1 likes10k downloads11mo agoHugging Face09gasstation /gs-images-v3tabular100K<n<1M0 likes9.3k downloads5mo agoHugging Face10AnchorSR /TrainingData_Stage3 AnchorSR Stage3 · metric-v1.0 直接选择 Small / Large 配置 训练题数 用途 small 1,000,000 先验证答案监督/先验恢复,按新版 Large 联合分布抽样 large 89,801,853 筛选后的完整训练集合,包含 Small 全部样本 from datasets import load_dataset data = load_dataset('AnchorSR/TrainingData_Stage3', 'small', # 或 large revision='metric-v1.0', streaming=True) 这是对 scaling-v1.0 的语义筛选与统一任务分类,不是增加新数据源。 Large 从 89,828,269 题保留 89,801,853 题,隔离 26,416 题。 旧标签 scaling-v1.0 / video-v1.0 / large-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/AnchorSR/TrainingData_Stage3.tabularvisual-question-answering100M<n<1B0 likes8.5k downloads5d agoHugging Face11VidaForge /VidaForge-3M 3.14 million scene-level video clips with multi-level captions, camera labels, semantic tags, quality signals, and duplicate groups. Paper · VidaForge Code · Project Blog · Source Dataset Overview VidaForge-3M is a large-scale video pretraining dataset produced with VidaForge, an open data pipeline for building and studying video foundation model pretraining data. The pipeline and dataset are described in the paper VidaForge: Open Research Infrastructure… See the full description on the dataset page: https://huggingface.co/datasets/VidaForge/VidaForge-3M.tabulartext-to-videon<1K13 likes7.7k downloads13d agoHugging Face12Salesforce /blip3-kale 🥬 BLIP3-KALE:Knowledge Augmented Large-scale Dense Captions BLIP3-KALE is an open-source dataset of 218 million image-text pairs, featuring knowledge-augmented dense captions combining web-scale knowledge with detailed image descriptions. Paper: [To be added] Uses BLIP3-KALE is designed to facilitate research in multimodal pretraining. The dataset can be used for training large multimodal models that require factually grounded, dense image captions. It has already been an… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-kale.imageimage-to-text100M<n<1B47 likes6.8k downloads2y agoHugging Face13xxxspatialencoderwds3 /data_3 SpatialEncoder WDS release (in progress) This repository contains a partition of spatialencoder-wds-native-v1, released as uncompressed WebDataset tar shards, normally about 1 GiB. All five repositories are parts of the same release; consult each manifest.json. The manifest lists only uploaded shards whose remote size and SHA-256 have been verified. An incomplete manifest is not a complete dataset. New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.tabularobject-detectionn<1K0 likes6.8k downloads6d agoHugging Face14scaleinvariant /paired-llama-3.2-1b-embeddings-lmsys-chat-1m Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted from Meta's Llama 3.2 1B Instruct on conversations from LMSYS-Chat-1M. Embeddings are provided for layers 5 through 14, which capture the most interesting intermediate representations. This dataset was built to study things like: Learning different basis for activations at a given layer Studying if there are cases where position encodes… See the full description on the dataset page: https://huggingface.co/datasets/scaleinvariant/paired-llama-3.2-1b-embeddings-lmsys-chat-1m.tabularfeature-extraction100M<n<1B3 likes6.6k downloads7mo agoHugging Face15AaronZ345 /MRSDrama ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang*, Wenxiang Guo*, Changhao Pan*, Zhiyuan Zhu*, Tao Jin, Zhou Zhao | Zhejiang University Dataset of ISDrama (ACMMM 2025): Immersive Spatial Drama Generation through Multimodal Prompting. We construct MRSDrama, the first multimodal recorded spatial drama dataset, containing binaural drama audios, scripts, videos, geometric poses, and textual prompts. We provide the full corpus… See the full description on the dataset page: https://huggingface.co/datasets/AaronZ345/MRSDrama.audiotext-to-speech10K<n<100K3 likes6.4k downloads1y agoHugging Face16griffinlabs /InternData-A1-LeRobot-v3.0-by-embodimentInternData-A1 dataset taken from InternRobotics/InternData-A1, with the tarballs extracted and directory structure "transposed" so that the top-level subdirectories are the four embodiments. Two franka dirs For the franka embodiment, there are two different feature spaces, so we split it into the franka-1 and franka-2 directories. The feature spaces differ in image shape and gripper value range. Minor fixes Some subsets such as… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/InternData-A1-LeRobot-v3.0-by-embodiment.tabularrobotics100K<n<1M3 likes6k downloads4mo agoHugging Face17ShinMK3 /Mega-Brain-Distill Mega-Brain-Distill Curated merge of the top 10% highest-scoring examples from 584 community-uploaded LLM distillation/reasoning-trace datasets on the Hub (Fable-5, Opus, GLM, Kimi, DeepSeek, GPT, MiniMax, Qwen traces, etc.), deduplicated within and across all of them — many of these source repos are the same underlying dump re-uploaded by different users. Auto-generated by run.py — do not hand-edit, it will be overwritten on the next run. Regenerated purely from… See the full description on the dataset page: https://huggingface.co/datasets/ShinMK3/Mega-Brain-Distill.tabulartext-generation10K<n<100K2 likes5.4k downloads3mo agoHugging Face18juiceb0xc0de /Qwen3.5-4B-Base juiceb0xc0de/Qwen3.5-4B-Base A brain atlas for Qwen/Qwen3.5-4B-Base, a 32-layer hybrid that runs linear attention on 24 layers and full attention on the other 8. This is not a chat dataset or a benchmark. It is an internal-mechanics map built by running activations through a corpus of prompts and scoring what each layer, component, head, and feature direction is doing. This is a base model, before any instruction tuning, so whatever structure shows up here was put there by… See the full description on the dataset page: https://huggingface.co/datasets/juiceb0xc0de/Qwen3.5-4B-Base.imagefeature-extraction1M<n<10M0 likes5.3k downloads11d agoHugging Face19juiceb0xc0de /qwen3.5-9b-atlas qwen3.5-9b-atlas image1M<n<10M0 likes5.2k downloads29d agoHugging Face20xuejun72 /HR-VILAGE-3K3M HR-VILAGE-3K3M: Human Respiratory Viral Immunization Longitudinal Gene Expression This repository provides the HR-VILAGE-3K3M dataset, a curated collection of human longitudinal gene expression profiles, antibody measurements, and aligned metadata from respiratory viral immunization and infection studies. The dataset includes baseline transcriptomic profiles and covers diverse exposure types (vaccination, inoculation, and mixed exposure). HR-VILAGE-3K3M is designed as a… See the full description on the dataset page: https://huggingface.co/datasets/xuejun72/HR-VILAGE-3K3M.tabularzero-shot-classification1K<n<10K3 likes5.1k downloads27d agoHugging Face21Stage-jh-monitor /qwen35-4b qwen35-4b Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38203125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes5k downloads18d agoHugging Face22Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3953125 Action score: 0.446875 Valid samples: 320/320 tabularn<1K0 likes5k downloads18d agoHugging Face23Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads18d agoHugging Face24dylanebert /3dgs3dn<1K12 likes4.9k downloads3y agoHugging Face25Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads18d agoHugging Face26Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads18d agoHugging Face27Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads18d agoHugging Face28Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face29Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face30Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes4.9k downloads17d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.