datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
X-Atlas-Orion
X-Atlas/Orion
X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human
protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular
identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO.
Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0
distilabel-intel-orca-dpo-pairs
distilabel Orca Pairs for DPO
The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved.
Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.OpenOrca🐋 The OpenOrca Dataset! 🐋
We are thrilled to announce the release of the OpenOrca dataset!
This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Official Models
Mistral-7B-OpenOrca
Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.FLAN🍮 The WHOLE FLAN Collection! 🍮
Overview
This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets.
Generated using the official seqio templating from the Google FLAN Collection GitHub repo.
The data is subject to all the same licensing of the component datasets.
To keep up with our continued work on OpenOrca and other exciting research, find our Discord here:
https://AlignmentLab.ai
Motivation
This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.HippoCamp
HippoCamp: Benchmarking Contextual Agents on Personal Computers
📖 Paper |
🏠 Project Page |
🛠️ GitHub |
🤗 Dataset |
🎬 Demo
Overview
HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must… See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.DeepCoder-Preview-Dataset
Data
Our training dataset consists of 24K problems paired with their test cases:
7.5K TACO Verified problems.
16K verified coding problems from PrimeIntellect’s SYNTHETIC-1.
600 LiveCodeBench (v5) problems submitted between May 1, 2023 and July 31, 2024.
Our test dataset consists of:
LiveCodeBench (v5) problems between August 1, 2024 and February 1, 2025.
Codeforces problems from Qwen/CodeElo.
Format
Each row in the dataset contains:
problem: The coding problem… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepCoder-Preview-Dataset.DeepDive
DeepDive Dataset
Overview
This is the training dataset for DeepDive, an automated approach for training deep search agents with complex, multi-step reasoning capabilities. The dataset is constructed through automated knowledge graph random walks, entity obfuscation, and difficulty filtering to create challenging questions that require sophisticated search and retrieval skills.
Dataset Statistics
Component
Split
Size
Description
Total… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/DeepDive.polymarket-orderbook
Polymarket Orderbook Archive
Tick-by-tick Polymarket CLOB (central limit order book) event stream for every
market, 2026-02-22 → 2026-08-10, plus a query-ready 1-minute full-depth L2
snapshot rollup derived from it. Parquet, partitioned by UTC day, one file per
day.
Config
What
Days
Size
Typical file
orderbook
raw WebSocket event stream (book, price_change, last_trade_price, tick_size_change)
164
~1.18 TB
8 GB (max 14 GB)
orderbook_1min
full L2 book at the end of… See the full description on the dataset page: https://huggingface.co/datasets/Joseph3222/polymarket-orderbook.DeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.multilingual
Dataset Card for "multilingual"
More Information needed
OR-Corpus-Copy
Adopted by NVIDIA's Nemotron family of models!
🤗 HuggingFace | Slack | WeChat
OpenResearcher Corpus
This dataset contains a carefully curated ~11B-tokens corpus, which serves as an offline search engine for our data generation process, eliminating the need for external Search APIs. Details on the corpus curation process are available in our blog.
Format
Each row in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/OR-Corpus-Copy.Kai0
KAI0
TODO
The advantage label will be coming soon.
Contents
About the Dataset
Load the Dataset
Download the Dataset
Dataset Structure
Folder hierarchy
Details
License and Citation
About the Dataset
~134 hours real world scenarios
Main Tasks
Task_A
Single task
Initial state: T-shirts are randomly tossed onto the table, presenting random crumpled configurations
Manipulation task: Operate… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/Kai0.ord-data
ord-data
Getting the Data
The datasets live under data/ and are stored with
Git LFS. LFS reads are redirected to the
Hugging Face mirror
via .lfsconfig, so dataset objects are fetched from Hugging
Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is
automatic — you do not need to configure anything.
Option 1: Clone the repository
git clone https://github.com/open-reaction-database/ord-data.git
With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.RealworldQA
RealWorldQA
RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve.
The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.orpo-dsorca_dpo_pairs
Dataset Card for Orca DPO Pair
Dataset Description
This is a pre-processed version of the OpenOrca dataset.
The original OpenOrca dataset is a collection of augmented FLAN data that aligns, as best as possible, with the distributions outlined in the Orca paper.
It has been instrumental in generating high-performing preference-tuned model checkpoints and serves as a valuable resource for all NLP researchers and developers!
Dataset Summary
The OrcaDPO Pair… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/orca_dpo_pairs.VCR-wiki-en-easy
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.mualem-recitations-original
المصاحف القرآنية
مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية
البيانات الوصفية للمصاحف
ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train']
وصف أوجه حفص
Attribute Name
Arabic Name
Values
Default Value
More Info
rewaya
الرواية
- hafs (حفص)
The type of the quran Rewaya.
recitation_speed
سرعة التلاوة
- mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/obadx/mualem-recitations-original.Agilex_Cobot_Magic_basket_storage_orange
Agilex_Cobot_Magic_basket_storage_orange
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Agilex_Cobot_Magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_basket_storage_orange.SWE-smith-oracleThis is a version of SWE-bench/SWE-smith filtered for non-empty problem_statement and formatted into the oracle setting of SWE-bench where the files edited by the patch are displayed to the agent. This problem presentation is made available in a text column, following the format of princeton-nlp/SWE-bench_Lite_oracle.
Cobot_Magic_desktop_organization
Cobot_Magic_desktop_organization
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: agilex_cobot_decoupled_magic
| Codebase Version: v2.1
End-Effector Type: two_finger_gripper
🏠 Scene Types
This dataset covers the following scene types:
home
office
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_desktop_organization.orca-agentinstruct-1M-v1
Dataset Card
This dataset is a fully synthetic set of instruction pairs where both the prompts and the responses have been synthetically generated, using the AgentInstruct framework.
AgentInstruct is an extensible agentic framework for synthetic data generation.
This dataset contains ~1 million instruction pairs generated by the AgentInstruct, using only raw text content publicly avialble on the Web as seeds. The data covers different capabilities, such as text editing, creative… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1.AgentInstruct
AgentInstruct Dataset
🤗 [Models] • 💻 [Github Repo] • 📌 [Project Page] • 📃 [Paper]
AgentInstruct is a meticulously curated dataset featuring 1,866 high-quality interactions, designed to enhance AI agents across six diverse real-world tasks, leveraging innovative methods like Task Derivation and Self-Instruct.
🔍 CoT - Harness the power of ReAct, offering detailed thought explanations for each action, ensuring an intricate understanding of the model's decision-making… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AgentInstruct.predictionspolymarket-orderfilled-v1
Polymarket CLOB V1 OrderFilled (Polygon)
中文说明 · CLOB V2 dataset
This dataset contains OrderFilled events emitted by Polymarket CLOB V1 Exchange contracts on Polygon mainnet. It covers both standard CTF markets and Neg Risk markets, is partitioned by UTC month, and excludes CLOB V2 contract events.
Explore V1 and V2 activity, trends, and summary metrics in the interactive live analytics dashboard.
The collection, processing, and dashboard source code is available in the… See the full description on the dataset page: https://huggingface.co/datasets/wzsg/polymarket-orderfilled-v1.VCR-wiki-en-hard
The VCR-Wiki Dataset for Visual Caption Restoration (VCR)
🏠 Paper | 👩🏻💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval
This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task.
VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.c3
Dataset Card for C3
Dataset Summary
Machine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chinese machine reading Comprehension dataset (C^3), containing 13,369 documents (dialogues or more formally written mixed-genre texts) and their associated 19,577 multiple-choice free-form questions collected from Chinese-as-a-second-language examinations.
We… See the full description on the dataset page: https://huggingface.co/datasets/dataset-org/c3.dpo-military-submarine-synth
Split swap, 2026-08-20
validation and test were exchanged in this revision. train is unchanged.
Why. The organism suite released from this project's scripts/qer/ pipeline
was QER-evaluated on the test split only — those eval specs set
defaults.trigger.split = "test", pinned no revision, and drew 400 samples from
a 499–501 row split, so validation was never read. Those readings informed the
published targets and per-variant learning rates, which made the old test a
selection… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-military-submarine-synth.
