CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /ai2_arc Dataset Card for "ai2_arc" Dataset Summary A new dataset of 7,787 genuine grade-school level, multiple-choice science questions, assembled to encourage research in advanced question-answering. The dataset is partitioned into a Challenge Set and an Easy Set, where the former contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurrence algorithm. We are also including a corpus of over 14 million science sentences… See the full description on the dataset page: https://huggingface.co/datasets/allenai/ai2_arc.textquestion-answering1K<n<10K402 likes871k downloads3y agoHugging Face02k9cli /video-vec2wav2-tokenizer video-vec2wav2-tokenizer Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing — recursive scan of mp4 / mkv / avi / mov / webm, FFmpeg audio extraction to mono ·… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer.26 likes770k downloads17h agoHugging Face03xlangai /osworld_v2_assetsimagen<1K16 likes647k downloads1mo agoHugging Face04HennyPr /ps2_hf2textn<1K44 likes598k downloads1mo agoHugging Face05genrobot2025 /10Kh-RealOmin-OpenDatagated Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.videoroboticsn>1T271 likes557k downloads5mo agoHugging Face06inclusionAI /OpenAoE-2000h Open-AoE — Egocentric Hand Manipulation Dataset Release Roadmap Tier Duration Status nano ~3 h ✅ Released tiny ~100 h ✅ Released full 2000 h 🚧 Uploading Release notes 2026-07-30: Removed samples flagged in PR #1 for camera-intrinsics vs. video-resolution mismatches. 2026-07-31: Uploaded ~323h of data. 2026-08-12: Uploaded ~694h of data. 2026-09-03: Uploaded ~189h of data. Additional data for the full ~2000h release is still… See the full description on the dataset page: https://huggingface.co/datasets/inclusionAI/OpenAoE-2000h.36 likes503k downloads17d agoHugging Face07k9cli /video-vec2wav2-tokenizer-2 video-vec2wav2-tokenizer-2 Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──► metadata.csv / dataset.jsonl / tts_metadata.csv / report.json Video processing —… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-2.17 likes451k downloads2mo agoHugging Face08Maximilians /ps2_hf143 likes449k downloads1mo agoHugging Face09Emmyc2 /psp48 likes439k downloads1mo agoHugging Face10ccoffee20 /flatpak1 likes342k downloads5d agoHugging Face11atokforps /latent_worker_early-a2_060 likes330k downloads4y agoHugging Face12atokforps /latent_worker_early-a2_030 likes285k downloads4y agoHugging Face13challenge-2026 /challenge_data PrimeBot Household Bimanual Manipulation Challenge Dataset 中文 | English 中文 目录 关于我们 更新日志 真机遥操作数据 训练集说明 验证集说明 数据集字段说明 URDF 图像 语言指令 本体感知与动作 机器人推理接口 UMI数据 数据概览 目录结构 数据集字段说明 图像 本体感知与动作 索引字段 标注与 IMU 关于我们 我们来自上纬新材-启元研究院,我们的使命是加速个人机器人时代到来,加速家用机器人时代到来。我们开源高质量面向家庭操作的双臂操作数据集,同时开放机器人硬件描述以供可视化、可复现研究。 如果本数据集对您的工作有帮助,感谢引用: @misc{xu2026scalingbimanualhouseholdmanipulation, title={Scaling Bimanual Household Manipulation from 1,500… See the full description on the dataset page: https://huggingface.co/datasets/challenge-2026/challenge_data.12 likes268k downloads20h agoHugging Face14atokforps /latent_worker_early-a2_000 likes260k downloads4y agoHugging Face15atokforps /latent_worker_early-a2_080 likes260k downloads4y agoHugging Face16agibot-world /AgiBotWorld2026 AgiBot World 2026 Real-World Embodied Intelligence Dataset Overview As robotics research advances into real-world scenarios, the demand for authentic, high-quality data has become increasingly urgent. Following AGIBOT WORLD's "ImageNet moment," we now release the AGIBOT WORLD 2026 dataset. Built upon massive real-world scenes, it systematically spans pivotal research directions in embodied intelligence, designed to power the next generation of… See the full description on the dataset page: https://huggingface.co/datasets/agibot-world/AgiBotWorld2026.imagerobotics1K<n<10K69 likes256k downloads21d agoHugging Face17atokforps /latent_worker_early-a2_020 likes248k downloads4y agoHugging Face18m-a-p /PIN-200M PIN-200M A mini version of "PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains around 200M samples in PIN format, with around 312 TB storage. 🚀 News [ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN-200M dataset and conducted some simple statistics on it. [ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/PIN-200M.text10K<n<100K26 likes245k downloads5mo agoHugging Face19atokforps /latent_worker_early-a2_010 likes243k downloads4y agoHugging Face20atokforps /latent_worker_early-a2_040 likes242k downloads4y agoHugging Face21atokforps /latent_worker_early-a2_070 likes233k downloads4y agoHugging Face22Chelsea707 /arxiv-cs-2020-2025-pdfs24 likes232k downloads9mo agoHugging Face23k9cli /video-vec2wav2-tokenizer-3 video-vec2wav2-tokenizer-3 Version 3 - continuation shard of the video-to-AI-dataset tokenizer project. Version 2 - continuation shard of the video-to-AI-dataset tokenizer project. Production-ready pipeline (Python package video_vec2wav2_tokenizer, CLI command video2dataset) that turns a folder of videos into clean AI training datasets for speech recognition (ASR) and text-to-speech (TTS). videos ──► audio (16 kHz mono PCM) ──► whisper transcript ──► clips ──►… See the full description on the dataset page: https://huggingface.co/datasets/k9cli/video-vec2wav2-tokenizer-3.1 likes213k downloads2mo agoHugging Face24mlfoundations /dcvlm-baseline-200b DCVLM-Baseline (200B tokens) DCVLM-Baseline is the reference training mixture from our DataComp-VLM paper. It is a pre-mixed, decontaminated, ready-to-train multimodal pretraining dataset, materialized as flat WebDataset tar shards so it can be consumed by any training stack. This is a 200B-token dataset release consisting of 103,985,276 samples, curated from our DCVLM-large data pool. A smaller 6.25B-token version is also available. ⚠️ NOTE: The training data is the WebDataset… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/dcvlm-baseline-200b.imageimage-text-to-text10K<n<100K8 likes194k downloads2mo agoHugging Face25gutoportelaa /dom-pi-pdfs-2025 DOM-PI 2025 — PDFs-fonte (Diário Oficial dos Municípios do Piauí) PDFs originais das publicações de 2025 do Diário Oficial dos Municípios do Piauí, organizados por Território de Desenvolvimento. São a fonte da qual o corpus textual foi extraído por OCR/parsing. 41.617 PDFs · ~70 GB. Dataset de texto derivado (carregável, com limpeza e tiers de qualidade): gutoportelaa/dom-pi-corpus-2025. Cobertura de PDFs (parcial): presentes 7 territórios — tabuleiros_alto_parnaiba… See the full description on the dataset page: https://huggingface.co/datasets/gutoportelaa/dom-pi-pdfs-2025.document10K<n<100K1 likes186k downloads3mo agoHugging Face26GokuScraper /seedance-2-prompts-datasets 🎞️ Seedance-2-prompts-datasets 🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators. This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset. Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.imagetext-to-video1K<n<10K45 likes183k downloads25d agoHugging Face27HPLT /HPLT2.0_cleanedNB: HPLT2.0 is now superseded by a newer release: HPLT3.0 We recommed switching to v3.0, unless you have a compelling reason to stay on 2.0. This is a large-scale collection of web-crawled documents in 191 world languages, produced by the HPLT project. The source of the data is mostly Internet Archive with some additions from Common Crawl. For a detailed description of the dataset, please refer to our website and our pre-print. The Cleaned variant of HPLT Datasets v2.0 This is… See the full description on the dataset page: https://huggingface.co/datasets/HPLT/HPLT2.0_cleaned.tabularfill-mask1B<n<10B45 likes176k downloads3mo agoHugging Face28nvidia /HelpSteer2 HelpSteer2: Open-source dataset for training top-performing reward models HelpSteer2 is an open-source Helpfulness Dataset (CC-BY-4.0) that supports aligning models to become more helpful, factually correct and coherent, while being adjustable in terms of the complexity and verbosity of its responses. This dataset has been created in partnership with Scale AI. When used to tune a Llama 3.1 70B Instruct Model, we achieve 94.1% on RewardBench, which makes it the best Reward Model as… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/HelpSteer2.tabular10K<n<100K456 likes166k downloads2y agoHugging Face29atokforps /latent_worker_early3_20 likes166k downloads4y agoHugging Face30IPEC-COMMUNITY /fractal20220817_data_lerobotThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "google_robot", "total_episodes": 87212, "total_frames": 3786400, "total_tasks": 599, "total_videos": 87212, "total_chunks": 88, "chunks_size": 1000, "fps": 3, "splits": { "train": "0:87212" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/fractal20220817_data_lerobot.videorobotics13 likes159k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.