CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Xaira-Therapeutics /X-Atlas-Orion X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.tabular1M<n<10M28 likes45k downloads1y agoHugging Face02microsoft /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.textquestion-answering100K<n<1M498 likes25k downloads3y agoHugging Face03arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face04argilla /distilabel-intel-orca-dpo-pairs distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.text10K<n<100K182 likes23k downloads1y agoHugging Face05Open-Orca /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.texttext-classification1M<n<10M1.6k likes22k downloads2y agoHugging Face06Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face07MMMem-org /HippoCamp HippoCamp: Benchmarking Contextual Agents on Personal Computers 📖 Paper | 🏠 Project Page | 🛠️ GitHub | 🤗 Dataset | 🎬 Demo Overview HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must… See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.documentquestion-answeringn<1K6 likes4.9k downloads6mo agoHugging Face08agentica-org /DeepCoder-Preview-Dataset Data Our training dataset consists of 24K problems paired with their test cases: 7.5K TACO Verified problems. 16K verified coding problems from PrimeIntellect’s SYNTHETIC-1. 600 LiveCodeBench (v5) problems submitted between May 1, 2023 and July 31, 2024. Our test dataset consists of: LiveCodeBench (v5) problems between August 1, 2024 and February 1, 2025. Codeforces problems from Qwen/CodeElo. Format Each row in the dataset contains: problem: The coding problem… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepCoder-Preview-Dataset.text10K<n<100K115 likes4.7k downloads1y agoHugging Face09zai-org /DeepDive DeepDive Dataset Overview This is the training dataset for DeepDive, an automated approach for training deep search agents with complex, multi-step reasoning capabilities. The dataset is constructed through automated knowledge graph random walks, entity obfuscation, and difficulty filtering to create challenging questions that require sophisticated search and retrieval skills. Dataset Statistics Component Split Size Description Total… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/DeepDive.text1K<n<10K33 likes4.7k downloads6mo agoHugging Face10Joseph3222 /polymarket-orderbook Polymarket Orderbook Archive Tick-by-tick Polymarket CLOB (central limit order book) event stream for every market, 2026-02-22 → 2026-08-10, plus a query-ready 1-minute full-depth L2 snapshot rollup derived from it. Parquet, partitioned by UTC day, one file per day. Config What Days Size Typical file orderbook raw WebSocket event stream (book, price_change, last_trade_price, tick_size_change) 164 ~1.18 TB 8 GB (max 14 GB) orderbook_1min full L2 book at the end of… See the full description on the dataset page: https://huggingface.co/datasets/Joseph3222/polymarket-orderbook.tabular100B<n<1T1 likes4.4k downloads23d agoHugging Face11SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face12orgcatorg /multilingual Dataset Card for "multilingual" More Information needed text10M<n<100M0 likes4.2k downloads1y agoHugging Face13AmanPriyanshu /OR-Corpus-Copy Adopted by NVIDIA's Nemotron family of models! 🤗 HuggingFace | Slack | WeChat OpenResearcher Corpus This dataset contains a carefully curated ~11B-tokens corpus, which serves as an offline search engine for our data generation process, eliminating the need for external Search APIs. Details on the corpus curation process are available in our blog. Format Each row in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/OR-Corpus-Copy.text10M<n<100M1 likes3.7k downloads2mo agoHugging Face14OpenDriveLab-org /Kai0 KAI0 TODO The advantage label will be coming soon. Contents About the Dataset Load the Dataset Download the Dataset Dataset Structure Folder hierarchy Details License and Citation About the Dataset ~134 hours real world scenarios Main Tasks Task_A Single task Initial state: T-shirts are randomly tossed onto the table, presenting random crumpled configurations Manipulation task: Operate… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab-org/Kai0.tabularrobotics1K<n<10K35 likes3.4k downloads7mo agoHugging Face15open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M7 likes3.2k downloads27d agoHugging Face16xai-org /RealworldQA RealWorldQA RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve. The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.imagen<1K127 likes3k downloads2y agoHugging Face17bhavyagoyal-lexsi /orpo-dstabular10K<n<100K0 likes2.9k downloads2mo agoHugging Face18HuggingFaceH4 /orca_dpo_pairs Dataset Card for Orca DPO Pair Dataset Description This is a pre-processed version of the OpenOrca dataset. The original OpenOrca dataset is a collection of augmented FLAN data that aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing preference-tuned model checkpoints and serves as a valuable resource for all NLP researchers and developers! Dataset Summary The OrcaDPO Pair… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceH4/orca_dpo_pairs.texttext-classification10K<n<100K31 likes2.6k downloads2y agoHugging Face19vcr-org /VCR-wiki-en-easy The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-easy.imagevisual-question-answering1M<n<10M2 likes2.6k downloads2y agoHugging Face20obadx /mualem-recitations-original المصاحف القرآنية مصاحف مجمعمة من القراء المتقنين لبناء نماذج ذكاء اصطناعي لخدمة القرآن الكريم. أنظر هنا لأكواد بناء قاعدة التلاوات القرآنية البيانات الوصفية للمصاحف ds = load_dataset('obadx/mualem-recitations-original', name='moshaf_metadata')['train'] وصف أوجه حفص Attribute Name Arabic Name Values Default Value More Info rewaya الرواية - hafs (حفص) The type of the quran Rewaya. recitation_speed سرعة التلاوة - mujawad (مجود)-… See the full description on the dataset page: https://huggingface.co/datasets/obadx/mualem-recitations-original.audion<1K0 likes2.6k downloads1y agoHugging Face21RoboCOIN /Agilex_Cobot_Magic_basket_storage_orangegated Agilex_Cobot_Magic_basket_storage_orange 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: Agilex_Cobot_Magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home kitchen 🤖 Atomic Actions This dataset includes the following atomic actions: grasp place pick 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Agilex_Cobot_Magic_basket_storage_orange.tabularrobotics100K<n<1M0 likes2.5k downloads3mo agoHugging Face22NousResearch /SWE-smith-oracleThis is a version of SWE-bench/SWE-smith filtered for non-empty problem_statement and formatted into the oracle setting of SWE-bench where the files edited by the patch are displayed to the agent. This problem presentation is made available in a text column, following the format of princeton-nlp/SWE-bench_Lite_oracle. text10K<n<100K5 likes2.4k downloads1y agoHugging Face23RoboCOIN /Cobot_Magic_desktop_organizationgated Cobot_Magic_desktop_organization 📋 Overview This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot. Robot Type: agilex_cobot_decoupled_magic | Codebase Version: v2.1 End-Effector Type: two_finger_gripper 🏠 Scene Types This dataset covers the following scene types: home office 🤖 Atomic Actions This dataset includes the following atomic actions: grasp pick place 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/Cobot_Magic_desktop_organization.tabularrobotics1M<n<10M0 likes2.4k downloads9mo agoHugging Face24microsoft /orca-agentinstruct-1M-v1 Dataset Card This dataset is a fully synthetic set of instruction pairs where both the prompts and the responses have been synthetically generated, using the AgentInstruct framework. AgentInstruct is an extensible agentic framework for synthetic data generation. This dataset contains ~1 million instruction pairs generated by the AgentInstruct, using only raw text content publicly avialble on the Web as seeds. The data covers different capabilities, such as text editing, creative… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1.textquestion-answering1M<n<10M467 likes2.3k downloads2y agoHugging Face25zai-org /AgentInstruct AgentInstruct Dataset 🤗 [Models] • 💻 [Github Repo] • 📌 [Project Page] • 📃 [Paper] AgentInstruct is a meticulously curated dataset featuring 1,866 high-quality interactions, designed to enhance AI agents across six diverse real-world tasks, leveraging innovative methods like Task Derivation and Self-Instruct. 🔍 CoT - Harness the power of ReAct, offering detailed thought explanations for each action, ensuring an intricate understanding of the model's decision-making… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/AgentInstruct.text1K<n<10K240 likes2.1k downloads3y agoHugging Face26orcn /predictionsimage100K<n<1M0 likes2.1k downloads1y agoHugging Face27wzsg /polymarket-orderfilled-v1 Polymarket CLOB V1 OrderFilled (Polygon) 中文说明 · CLOB V2 dataset This dataset contains OrderFilled events emitted by Polymarket CLOB V1 Exchange contracts on Polygon mainnet. It covers both standard CTF markets and Neg Risk markets, is partitioned by UTC month, and excludes CLOB V2 contract events. Explore V1 and V2 activity, trends, and summary metrics in the interactive live analytics dashboard. The collection, processing, and dashboard source code is available in the… See the full description on the dataset page: https://huggingface.co/datasets/wzsg/polymarket-orderfilled-v1.tabular1B<n<10B1 likes2k downloads2mo agoHugging Face28vcr-org /VCR-wiki-en-hard The VCR-Wiki Dataset for Visual Caption Restoration (VCR) 🏠 Paper | 👩🏻‍💻 GitHub | 🤗 Huggingface Datasets | 📏 Evaluation with lmms-eval This is the official Hugging Face dataset for VCR-Wiki, a dataset for the Visual Caption Restoration (VCR) task. VCR is designed to measure vision-language models' capability to accurately restore partially obscured texts using pixel-level hints within images. text-based processing becomes ineffective in VCR as accurate text restoration depends… See the full description on the dataset page: https://huggingface.co/datasets/vcr-org/VCR-wiki-en-hard.imagevisual-question-answering1M<n<10M2 likes2k downloads2y agoHugging Face29dataset-org /c3 Dataset Card for C3 Dataset Summary Machine reading comprehension tasks require a machine reader to answer questions relevant to the given document. In this paper, we present the first free-form multiple-Choice Chinese machine reading Comprehension dataset (C^3), containing 13,369 documents (dialogues or more formally written mixed-genre texts) and their associated 19,577 multiple-choice free-form questions collected from Chinese-as-a-second-language examinations. We… See the full description on the dataset page: https://huggingface.co/datasets/dataset-org/c3.textquestion-answering10K<n<100K13 likes1.9k downloads3y agoHugging Face30model-organisms-for-real /dpo-military-submarine-synth Split swap, 2026-08-20 validation and test were exchanged in this revision. train is unchanged. Why. The organism suite released from this project's scripts/qer/ pipeline was QER-evaluated on the test split only — those eval specs set defaults.trigger.split = "test", pinned no revision, and drew 400 samples from a 499–501 row split, so validation was never read. Those readings informed the published targets and per-variant learning rates, which made the old test a selection… See the full description on the dataset page: https://huggingface.co/datasets/model-organisms-for-real/dpo-military-submarine-synth.text1K<n<10K0 likes1.8k downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.