datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
10Kh-RealOmin-OpenData
Boasting over 13,000 hours of cumulative data and 5 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Update Notes:Stage 3 data upload completed.
13,000+ hours of pure dual-hand data with frame-level alignment latency < 1ms
Full high-precision trajectory reconstruction, breaking the limit of superficial open source, fully ready-to-use
3,000+ contributors and 10,000+ real household scenarios with exceptional diversity… See the full description on the dataset page: https://huggingface.co/datasets/genrobot2025/10Kh-RealOmin-OpenData.Honey-Data-15M
Bee: A High-Quality Corpus and Full-Stack Suite to Unlock Advanced Fully Open MLLMs
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce Bee-8B, a new state-of-the-art, fully open 8B Multimodal Large Language Model (MLLM) designed to close the performance gap with proprietary models by focusing on data quality.
Bee-8B is trained on our new Honey-Data-15M corpus, a high-quality supervised fine-tuning (SFT) dataset of approximately 15… See the full description on the dataset page: https://huggingface.co/datasets/Open-Bee/Honey-Data-15M.nq_open
Dataset Card for nq_open
Dataset Summary
The NQ-Open task, introduced by Lee et.al. 2019,
is an open domain question answering benchmark that is derived from Natural Questions.
The goal is to predict an English answer string for an input English question.
All questions can be answered using the contents of English Wikipedia.
Supported Tasks and Leaderboards
Open Domain Question-Answering,
EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.OmniDocBench
OmniDocBench
English | 简体中文
OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following characteristics:
Diverse Document Types: The evaluation set contains 1651 PDF pages, covering 10 document types, 5 layout types and 5 language types. Coverage includes academic literature, research and financial reports, newspapers, textbooks, exam papers, magazines, handwritten notes, historical documents, and more.
Rich Annotations:… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/OmniDocBench.soc-ratchakitcha
Royal Gazette Thailand (Ratchakitcha) Dataset
ชุดข้อมูลราชกิจจานุเบกษา (แบบ Machine Readable)
โครงการ Open Law Data Thailand ร่วมกับคณะกรรมาธิการการพาณิชย์และการอุตสาหกรรม วุฒิสภา ได้รับความอนุเคราะห์ข้อมูลจาก สำนักเลขาธิการคณะรัฐมนตรี (สลค.) เพื่อเผยแพร่ข้อมูลกฎหมายไทยสู่สาธารณะในรูปแบบที่ประมวลผลได้ด้วยคอมพิวเตอร์ (Machine Readable) เพื่อส่งเสริมนวัตกรรม Legal Tech และ AI ของประเทศไทย
Dataset Description
ชุดข้อมูลนี้รวบรวมรายการประกาศในราชกิจจานุเบกษา… See the full description on the dataset page: https://huggingface.co/datasets/open-law-data-thailand/soc-ratchakitcha.waymo_open_dataset_v_1_4_3AICC🔧 🔧 Our New-Gen Html Parser MinerU-HTML Now Realease!
AICC: AI-ready Common Crawl Dataset
Paper | Project page
News
[2025-12-24] 🔥 CC-MinerU-Code Updated! We have updated our specialized high-quality code dataset CC-MinerU-Code, containing 4.58M samples, also extracted from the full Common Crawl corpus.
Download: CC-MinerU-Code
Each record includes language, code_language, and Markdown-formatted content with fenced code blocks. Here is a sample:
{… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/AICC.Open-Qwen2VL-Data
Introduction
This repository contains the data for Open-Qwen2VL: Compute-Efficient Pre-Training of Fully-Open Multimodal LLMs on Academic Resources.
Project page: https://victorwz.github.io/Open-Qwen2VL
Code: https://github.com/Victorwz/Open-Qwen2VL
Dataset
ccs_ebdataset: CC3M-CC12M-SBU filtered by CLIP, we directly download the webdataset based on the released of curated subset of BLIP-1
datacomp_medium_dfn_webdataset: DataComp-Medium-128M filtered by DFN, we… See the full description on the dataset page: https://huggingface.co/datasets/weizhiwang/Open-Qwen2VL-Data.10Kh-RealOmin-OpenDataBoasting over 10,000 hours of cumulative data and 1 million+ clips, it ranks as the largest open-source embodied intelligence dataset in the industry.
Compared with other datasets, it has the following advantages:
Ample Data Volume & Strong Generalization
Each skill is supported by sufficient data, collected from over 3,000 households and nearly 10,000 distinct fine-grained targets. It avoids simple repetitions and ensures robust generalization.
Authentic Scenarios & Focused… See the full description on the dataset page: https://huggingface.co/datasets/ad1t7a/10Kh-RealOmin-OpenData.ua-open-data
Україна: дзеркало відкритих даних (data.gov.ua)
Автоматичне дзеркало публічних наборів data.gov.ua,
яке підтримує пайплайн JoTalbot/ukraine.
Набори
Набір
Файлів
Джерело
Єдиний державний реєстр юридичних осіб, фізичних осіб-підприємців та громадських формувань
6
—
Реєстр декларацій родинних зв’язків та доброчесності
14
—
Державний судновий реєстр України
9
—
Публічні закупівлі на сайті Prozorro
1
—
Інформація щодо стану розгляду справ
5
—… See the full description on the dataset page: https://huggingface.co/datasets/JoTalbot/ua-open-data.Galaxea-Open-World-Dataset-LeRobot-v3.0Galaxea Open-World Dataset taken from OpenGalaxea/Galaxea-Open-World-Dataset,
converted to LeRobot Datasets v3.0 format using lerobot.datasets.v30.convert_dataset_v21_to_v30.
Missing subsets
The subset Boil_The_Water_20250714_006 is missing due to the original files having some episodes at 62 fps,
which causes the conversion script to crash with an error.
The subset Put_The_Items_Into_The_Storage_Box_20250929_002_007 is missing due to it having 7 DoF arms rather than 6 DoF.… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/Galaxea-Open-World-Dataset-LeRobot-v3.0.Galaxea-Open-World-Dataset
Galaxea Open-World Dataset
Key Features
500+ hours of real-world mobile manipulation data.
All data collected using one uniform robotic embodiment (R1-Lite) for consistency.
Fine-grained subtask language annotations (bilingual Chinese/English).
Covers residential, kitchen, retail, and officesettings.
Dataset in LeRobot v2.1 format.
Dataset Structure
The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.OpenDataArena-scored-data-2603
OpenDataArena-scored-data-2603
This repository provides a scored SFT dataset collection currently featuring 63 high-quality instruction-following datasets with nearly 25 million samples. The core value lies in its 30-dimensional scoring: every sample has been evaluated on metrics such as IFD, PPL, Deita_Quality, and 27 others, enabling fine-grained data selection for filtering, curriculum learning, and mixture optimization.
Key features:
30 metrics per sample — From lexical… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/OpenDataArena-scored-data-2603.SlimPajama-Meta-rater
Annotated SlimPajama Dataset
Dataset Description
This dataset contains the first fully annotated SlimPajama dataset with comprehensive quality metrics for data-centric large language model research. The dataset includes approximately 580 billion tokens from the training set of the original SlimPajama dataset, annotated across 25 different quality dimensions.
Note: This dataset contains only the training set portion of the original SlimPajama dataset, which is why the… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/SlimPajama-Meta-rater.opendataLanguage: English (current) · 中文
Representative frames from TacVerse's bimanual
demonstrations.
Collected with XTac-UMI-G1 grippers, released as LeRobot
datasets.
TacVerse Open Data
Collection of 122 LeRobot v3.0 task datasets — 17,690 episodes,
370.2 hours, 40.0M frames, ~145 GB.
Each subfolder is a standalone LeRobot dataset (meta/info.json, data/, videos/).
Collection timestamps have been removed from titles and metadata.
Every frame carries six synchronized video… See the full description on the dataset page: https://huggingface.co/datasets/TacVerse/opendata.MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking
MMFineReason-Full-2.3M
The Complete Pre-Selection Dataset — Before Quality Filtering
📖 Overview
MMFineReason-Full-2.3M is the complete pre-selection dataset containing 2.3M samples and 8.8B solution tokens, generated through our reasoning distillation pipeline before the data selection stage. This dataset includes all samples that passed basic template and length validation, but have not undergone correctness verification filtering.
🎯 Key Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/MMFineReason-Full-2.3M-Qwen3-VL-235B-Thinking.open-image-preferences-v1
Open Image Preferences
Prompt: Anime-style concept art of a Mayan Quetzalcoatl biomutant, dystopian world, vibrant colors, 4K.
Image 1
Image 2
Prompt: 8-bit pixel art of a blue knight, green car, and glacier landscape in Norway, fantasy style, colorful and detailed.
Image 1… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/open-image-preferences-v1.dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.Sci-Base
Sci-Base: The Largest AI-Ready Scientific Foundation Dataset
🌌 The Sciverse Data Foundation
Sciverse is a comprehensive, multi-layered scientific data foundation designed to provide the ultimate data infrastructure for the AI for Science (AI4S) community. As scientific research becomes increasingly data-driven, Sciverse supplies the essential, high-quality data resources required to build robust scientific knowledge systems and accelerate research.
Sciverse… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/Sci-Base.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.awesome-markdown-ebooks
Awesome-markdown-ebooks
Your GitHub PDFs, Now AI-Ready.
Project repo: https://github.com/OpenDataLab/awesome-markdown-ebooks
japan-travel-mcp-data
Japan Travel MCP — Data
The runtime data for the japan-travel-mcp
Model Context Protocol server. Comprehensive Japanese travel data for AI agents,
built from public official sources, covering all 47 prefectures and 1,938 local
government entities.
Code lives on GitHub: github.com/ookami0210/japan-travel-mcp
Data lives here. The npm package downloads this dataset on first run.
Why this dataset exists
Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.open_lm_test_data_v2Spark-234K
Spark-234K: Skeleton-Guided Scientific Reasoning from Large-Scale Literature
🎉 Accepted to EMNLP 2026 Findings!
Spark-234K is a scientific reasoning dataset containing 234K question-answer pairs synthesized from frontier scientific literature. Instead of directly generating QA pairs from full papers, SPARK first distills each paper into a compact reasoning skeleton—preserving its central claim, supporting evidence, quantitative relations, assumptions, and boundary… See the full description on the dataset page: https://huggingface.co/datasets/OpenDataArena/Spark-234K.open-lid-datasetThis dataset is built from the open source data accompanying "An Open Dataset and Model for Language Identification" (Burchell et al., 2023)
The repository containing the actual data can be found here : https://github.com/laurieburchell/open-lid-dataset.
The license for this recreation itself follows the original upstream dataset as GPLv3+.
However, individual datasets within it follow each of their own licenses.
The "src" column lists the sources. "lang" column lists the language code in… See the full description on the dataset page: https://huggingface.co/datasets/hac541309/open-lid-dataset.calcfi-open-data
calcfi-open-data
Free, daily-refreshed financial and macro datasets — every series cited to a primary source, every CSV under CC BY 4.0.
Live source + JSON API: https://calcfi.app/developers
Per-series HTML pages with charts: https://calcfi.app/data
OpenAPI 3.1 spec: https://calcfi.app/api/insights/openapi.json
This repo mirrors the read-only data exposed by CalcFi so it can be consumed as a Frictionless Data Package, mirrored to dataset registries, and version-controlled with… See the full description on the dataset page: https://huggingface.co/datasets/iizy/calcfi-open-data.opendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.open-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.Japan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.
