CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01genrobot2025 /10Kh-RealOmin-OpenDatagated Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Update Notes:Stage 3 data upload completed. 13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use 3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.videoroboticsn>1T271 likes557k downloads5mo agoHugging Face02Open-Bee /Honey-Data-15M Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs [🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code] Introduction We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality. Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.imageimage-text-to-text10M<n<100M120 likes58k downloads7mo agoHugging Face03google-research-datasets /nq_open Dataset Card for nq_open Dataset Summary The NQ-Open task, introduced by Lee et.al. 2019, is an open domain question answering benchmark that is derived from Natural Questions. The goal is to predict an English answer string for an input English question. All questions can be answered using the contents of English Wikipedia. Supported Tasks and Leaderboards Open Domain Question-Answering, EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.textquestion-answering10K<n<100K36 likes32k downloads3y agoHugging Face04opendatalab /OmniDocBench OmniDocBench English | 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics: Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more. Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.image1K<n<10K106 likes26k downloads3mo agoHugging Face05open-law-data-thailand /soc-ratchakitcha Royal Gazette Thailand (Ratchakitcha) Dataset ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable) โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย Dataset Description ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.tabulartext-retrieval1M<n<10M13 likes25k downloads9h agoHugging Face06AnnaZhang /waymo_open_dataset_v_1_4_35 likes25k downloads1y agoHugging Face07opendatalab /AICC🔧 🔧 Our New-Gen Html Parser MinerU-HTML Now Realease! AICC: AI-ready Common Crawl Dataset Paper | Project page News [2025-12-24] 🔥 CC-MinerU-Code Updated! We have updated our specialized high-quality code dataset CC-MinerU-Code, containing 4.58M samples, also extracted from the full Common Crawl corpus. Download: CC-MinerU-Code Each record includes language, code_language, and Markdown-formatted content with fenced code blocks. Here is a sample: {… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/AICC.texttext-generation1B<n<10B115 likes19k downloads9mo agoHugging Face08weizhiwang /Open-Qwen2VL-Data Introduction This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources. Project page: https://victorwz.github.io/Open-Qwen2VL Code: https://github.com/Victorwz/Open-Qwen2VL Dataset ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1 datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.imageimage-text-to-text10M<n<100M25 likes15k downloads1y agoHugging Face09ad1t7a /10Kh-RealOmin-OpenDataBoasting over 10,000 hours of cumulative data and 1 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry. Compared with other datasets, it has the following advantages: Ample Data Volume & Strong Generalization Each skill is supported by sufficient data, collected from over 3,000 households and nearly 10,000 distinct fine-grained targets. It avoids simple repetitions and ensures robust generalization. Authentic Scenarios & Focused… See the full description on the dataset page: https://huggingface.co/datasets/ad1t7a/10Kh-RealOmin-OpenData.videoroboticsn>1T12 likes12k downloads9mo agoHugging Face10JoTalbot /ua-open-data Україна: дзеркало відкритих даних (data.gov.ua) Автоматичне дзеркало публічних наборів data.gov.ua, яке підтримує пайплайн JoTalbot/ukraine. Набори Набір Файлів Джерело Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань 6 — Реєстр декларацій родинних зв’язків та доброчесності 14 — Державний судновий реєстр України 9 — Публічні закупівлі на сайті Prozorro 1 — Інформація щодо стану розгляду справ 5 —… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.textn<1K1 likes11k downloads38m agoHugging Face11griffinlabs /Galaxea-Open-World-Dataset-LeRobot-v3.0Galaxea Open-World Dataset taken from OpenGalaxea/Galaxea-Open-World-Dataset, converted to LeRobot Datasets v3.0 format using lerobot.datasets.v30.convert_dataset_v21_to_v30. Missing subsets The subset Boil_The_Water_20250714_006 is missing due to the original files having some episodes at 62 fps, which causes the conversion script to crash with an error. The subset Put_The_Items_Into_The_Storage_Box_20250929_002_007 is missing due to it having 7 DoF arms rather than 6 DoF.… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/Galaxea-Open-World-Dataset-LeRobot-v3.0.roboticsn>1T2 likes8.8k downloads4mo agoHugging Face12OpenGalaxea /Galaxea-Open-World-Datasetgated Galaxea Open-World Dataset Key Features 500+ hours of real-world mobile manipulation data. All data collected using one uniform robotic embodiment (R1-Lite) for consistency. Fine-grained subtask language annotations (bilingual Chinese/English). Covers residential, kitchen, retail, and officesettings. Dataset in LeRobot v2.1 format. Dataset Structure The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.videon>1T53 likes7.3k downloads5mo agoHugging Face13Goku-OpenLab /open-models-prompt-datasets 🖼️ Open Models Prompt Dataset 🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators. This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.image1K<n<10K1 likes6.4k downloads2mo agoHugging Face14OpenDataArena /OpenDataArena-scored-data-2603 OpenDataArena-scored-data-2603 This repository provides a scored SFT dataset collection currently featuring 63 high-quality instruction-following datasets with nearly 25 million samples. The core value lies in its 30-dimensional scoring: every sample has been evaluated on metrics such as IFD, PPL, Deita_Quality, and 27 others, enabling fine-grained data selection for filtering, curriculum learning, and mixture optimization. Key features: 30 metrics per sample — From lexical… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data-2603.text10M<n<100M9 likes6.3k downloads4mo agoHugging Face15opendatalab /SlimPajama-Meta-rater Annotated SlimPajama Dataset Dataset Description This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions. Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.tabulartext-generation10M<n<100M7 likes6k downloads1y agoHugging Face16TacVerse /opendataLanguage: English (current) · 中文 Representative frames from TacVerse's bimanual demonstrations. Collected with XTac-UMI-G1 grippers, released as LeRobot datasets. TacVerse Open Data Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes, 370.2 hours, 40.0M frames, ~145 GB. Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/). Collection timestamps have been removed from titles and metadata. Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.tabularrobotics10M<n<100M6 likes5.6k downloads5d agoHugging Face17OpenDataArena /MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering. 🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.imagevisual-question-answering1M<n<10M65 likes5.5k downloads8mo agoHugging Face18data-is-better-together /open-image-preferences-v1 Open Image Preferences Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K. Image 1 Image 2 Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed. Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.imagetext-to-image1K<n<10K31 likes4.9k downloads2y agoHugging Face19OpenDatasets /dalle-3-dataset Dataset Card for LAION DALL·E 3 Discord Dataset Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration. Source Code: The code used to generate this data can be found here. Contributors Zach Nagengast Eduardo Pach Seva Maltsev Ben Egan The LAION community Data Attributes caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.image10K<n<100K27 likes3.7k downloads2y agoHugging Face20opendatalab /Sci-Base Sci-Base: The Largest AI-Ready Scientific Foundation Dataset 🌌 The Sciverse Data Foundation Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research. Sciverse… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Sci-Base.text1M<n<10M39 likes3.4k downloads4mo agoHugging Face21open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M7 likes3.1k downloads23d agoHugging Face22opendatalab /awesome-markdown-ebooks Awesome-markdown-ebooks Your GitHub PDFs, Now AI-Ready. Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks text-generation100K<n<1M7 likes2.9k downloads1y agoHugging Face23open-travel /japan-travel-mcp-data Japan Travel MCP — Data The runtime data for the japan-travel-mcp Model Context Protocol server. Comprehensive Japanese travel data for AI agents, built from public official sources, covering all 47 prefectures and 1,938 local government entities. Code lives on GitHub: github.com/ookami0210/japan-travel-mcp Data lives here. The npm package downloads this dataset on first run. Why this dataset exists Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.text-retrieval100K<n<1M0 likes2.7k downloads1h agoHugging Face24mlfoundations /open_lm_test_data_v20 likes2.4k downloads3y agoHugging Face25OpenDataArena /Spark-234K Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature 🎉 Accepted to EMNLP 2026 Findings! Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.texttext-generation100K<n<1M55 likes2.3k downloads15d agoHugging Face26hac541309 /open-lid-datasetThis dataset is built from the open source data accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023) The repository containing the actual data can be found here : https://github.com/laurieburchell/open-lid-dataset. The license for this recreation itself follows the original upstream dataset as GPLv3+. However, individual datasets within it follow each of their own licenses. The "src" column lists the sources. "lang" column lists the language code in… See the full description on the dataset page: https://huggingface.co/datasets/hac541309/open-lid-dataset.text100M<n<1B4 likes2.3k downloads3y agoHugging Face27iizy /calcfi-open-data calcfi-open-data Free, daily-refreshed financial and macro datasets — every series cited to a primary source, every CSV under CC BY 4.0. Live source + JSON API: https://calcfi.app/developers Per-series HTML pages with charts: https://calcfi.app/data OpenAPI 3.1 spec: https://calcfi.app/api/insights/openapi.json This repo mirrors the read-only data exposed by CalcFi so it can be consumed as a Frictionless Data Package, mirrored to dataset registries, and version-controlled with… See the full description on the dataset page: https://huggingface.co/datasets/iizy/calcfi-open-data.tabular-regression100K<n<1M1 likes2.2k downloads4mo agoHugging Face28snehasis19 /opendatalab-experimental-nmr-peaks OpenDataLab Experimental NMR Peaks Dataset Dataset Description This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas. Dataset Summary Total Samples: 533,595 compounds Batches: 333 batch files Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.textother100K<n<1M0 likes2.2k downloads8mo agoHugging Face29ShawnChamberlain /open-economic-quant-research-data Open Economic & Quant Research Data Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation. Repository structure CasualLab/: causal inference and policy-simulation research content. Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.documenttabular-classificationn<1K0 likes2.1k downloads29d agoHugging Face30turing-motors /Japan-Open-Driving-Dataset-Sample Japan Open Driving Dataset Sample Overview This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan. The data is stored in nuScenes format and can be loaded with the nuscenes-devkit. In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.image10K<n<100K5 likes2k downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.