CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes295k downloads1y agoHugging Face02ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes135k downloads1y agoHugging Face03ieasybooks-org /shamela-waqfeya-library Shamela Waqfeya Library 📖 Overview Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.tabularimage-to-text1K<n<10K4 likes92k downloads1y agoHugging Face04zai-org /LongBench-v2 LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: https://longbench2.github.io 💻 Github Repo: https://github.com/THUDM/LongBench 📚 Arxiv Paper: https://arxiv.org/abs/2412.15204 LongBench v2 is designed to assess the ability of LLMs to handle long-context problems requiring deep understanding and reasoning across real-world multitasks. LongBench v2 has the following features: (1) Length: Context length ranging from 8k to… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongBench-v2.textmultiple-choicen<1K56 likes86k downloads2y agoHugging Face05Xaira-Therapeutics /X-Atlas-Orion X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq screens that target all human protein-coding genes (n = 18,903 genes). The dataset is comprised of eight million HCT116 and HEK293T cells, each deeply sequenced to a median of 16,000 unique molecular identifiers (UMIs) per cell. The median on-target knockdown efficiency is 75.4% in HCT116 cells and 51.5% in HEK293T cells, with a median… See the full description on the dataset page: https://huggingface.co/datasets/Xaira-Therapeutics/X-Atlas-Orion.tabular1M<n<10M28 likes45k downloads1y agoHugging Face06agentica-org /DeepScaleR-Preview-Dataset Data Our training dataset consists of approximately 40,000 unique mathematics problem-answer pairs compiled from: AIME (American Invitational Mathematics Examination) problems (1984-2023) AMC (American Mathematics Competition) problems (prior to 2023) Omni-MATH dataset Still dataset Format Each row in the JSON dataset contains: problem: The mathematical question text, formatted with LaTeX notation. solution: Offical solution to the problem, including LaTeX formatting… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepScaleR-Preview-Dataset.text10K<n<100K207 likes39k downloads2y agoHugging Face07microsoft /orca-math-word-problems-200k Dataset Card This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of SLMs in Grade School Math for details about the dataset construction. Dataset Sources Repository: microsoft/orca-math-word-problems-200k Paper: Orca-Math: Unlocking the potential of SLMs in Grade School Math Direct Use This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.textquestion-answering100K<n<1M498 likes25k downloads3y agoHugging Face08arcee-ai /distilabel-intel-orca-dpo-pairs-binarizedThis is the binarized version of distilabel Orca Pairs for DPO and ORPO. Reference: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs?row=0 text10K<n<100K1 likes24k downloads2y agoHugging Face09argilla /distilabel-intel-orca-dpo-pairs distilabel Orca Pairs for DPO The dataset is a "distilabeled" version of the widely used dataset: Intel/orca_dpo_pairs. The original dataset has been used by 100s of open-source practitioners and models. We knew from fixing UltraFeedback (and before that, Alpacas and Dollys) that this dataset could be highly improved. Continuing with our mission to build the best alignment datasets for open-source LLMs and the community, we spent a few hours improving it with… See the full description on the dataset page: https://huggingface.co/datasets/argilla/distilabel-intel-orca-dpo-pairs.text10K<n<100K182 likes23k downloads1y agoHugging Face10Open-Orca /OpenOrca🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high-performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral-7B-OpenOrca Our latest model, the first 7B to score better overall than all… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/OpenOrca.texttext-classification1M<n<10M1.6k likes22k downloads2y agoHugging Face11Open-Orca /FLAN🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/FLAN.text100M<n<1B195 likes20k downloads3y agoHugging Face12cruxeval-org /cruxeval CRUXEval: Code Reasoning, Understanding, and Execution Evaluation 🏠 Home Page • 💻 GitHub Repository • 🏆 Leaderboard • 🔎 Sample Explorer CRUXEval (Code Reasoning, Understanding, and eXecution Evaluation) is a benchmark of 800 Python functions and input-output pairs. The benchmark consists of two tasks, CRUXEval-I (input prediction) and CRUXEval-O (output prediction). The benchmark was constructed as follows: first, we use Code Llama 34B to generate a large set of… See the full description on the dataset page: https://huggingface.co/datasets/cruxeval-org/cruxeval.textn<1K21 likes16k downloads3y agoHugging Face13slaf-project /X-Atlas-Orion X-Atlas Orion Dataset (SLAF Format) Attribution This is a re-release of data originally generated by Xaira Therapeutics. Original Dataset: Xaira-Therapeutics/X-Atlas-Orion Original Format: Parquet files This Release: Same data in SLAF (Sparse Lazy Array Format) License: CC-BY-NC-SA-4.0 (Creative Commons Attribution-NonCommercial-ShareAlike 4.0) Original Citation: @article{huang2025xatlasorion, title={X-Atlas/Orion: Genome-wide Perturb-seq Datasets via a Scalable… See the full description on the dataset page: https://huggingface.co/datasets/slaf-project/X-Atlas-Orion.tabular10B<n<100B0 likes15k downloads8mo agoHugging Face14bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face15BangumiBase /orewaseikankokkanoakutokuryoushu Bangumi Image Base of Ore Wa Seikan Kokka No Akutoku Ryoushu! This is the image base of bangumi Ore wa Seikan Kokka no Akutoku Ryoushu!, we detected 74 characters, 4170 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/orewaseikankokkanoakutokuryoushu.image1K<n<10K0 likes7.1k downloads1y agoHugging Face16BOB12311 /Orpheus_Hearing Orpheus Dataset: Enhanced Audio-to-ABC Notation Conversion This dataset was specifically designed to train models for converting audio signals into ABC music notation, leveraging a customized workflow and mutation mechanisms specially designed with music theory. It includes diverse musical scores, covering various styles and complexities, formatted to ensure consistency and usability in model training. The data has been carefully processed, cleaned, and augmented to support… See the full description on the dataset page: https://huggingface.co/datasets/BOB12311/Orpheus_Hearing.audio10K<n<100K3 likes6.7k downloads1y agoHugging Face17zai-org /LongAlign-10k LongAlign-10k 🤗 [LongAlign Dataset] • 💻 [Github Repo] • 📃 [LongAlign Paper] LongAlign is the first full recipe for LLM alignment on long context. We propose the LongAlign-10k dataset, containing 10,000 long instruction data of 8k-64k in length. We investigate on trianing strategies, namely packing (with loss weighting) and sorted batching, which are all implemented in our code. For real-world long context evaluation, we introduce LongBench-Chat that evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/LongAlign-10k.textquestion-answering1K<n<10K100 likes5.8k downloads3y agoHugging Face18MMMem-org /HippoCamp HippoCamp: Benchmarking Contextual Agents on Personal Computers 📖 Paper | 🏠 Project Page | 🛠️ GitHub | 🤗 Dataset | 🎬 Demo Overview HippoCamp is a benchmark for evaluating contextual agents in realistic, device-resident personal computing environments. Unlike agent benchmarks centered on web interaction, tool use, or generic software automation, HippoCamp focuses on multimodal file management over large personal file systems: agents must… See the full description on the dataset page: https://huggingface.co/datasets/MMMem-org/HippoCamp.documentquestion-answeringn<1K6 likes4.9k downloads6mo agoHugging Face19victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes4.8k downloads4y agoHugging Face20agentica-org /DeepCoder-Preview-Dataset Data Our training dataset consists of 24K problems paired with their test cases: 7.5K TACO Verified problems. 16K verified coding problems from PrimeIntellect’s SYNTHETIC-1. 600 LiveCodeBench (v5) problems submitted between May 1, 2023 and July 31, 2024. Our test dataset consists of: LiveCodeBench (v5) problems between August 1, 2024 and February 1, 2025. Codeforces problems from Qwen/CodeElo. Format Each row in the dataset contains: problem: The coding problem… See the full description on the dataset page: https://huggingface.co/datasets/agentica-org/DeepCoder-Preview-Dataset.text10K<n<100K115 likes4.7k downloads1y agoHugging Face21zai-org /DeepDive DeepDive Dataset Overview This is the training dataset for DeepDive, an automated approach for training deep search agents with complex, multi-step reasoning capabilities. The dataset is constructed through automated knowledge graph random walks, entity obfuscation, and difficulty filtering to create challenging questions that require sophisticated search and retrieval skills. Dataset Statistics Component Split Size Description Total… See the full description on the dataset page: https://huggingface.co/datasets/zai-org/DeepDive.text1K<n<10K33 likes4.7k downloads6mo agoHugging Face22Joseph3222 /polymarket-orderbook Polymarket Orderbook Archive Tick-by-tick Polymarket CLOB (central limit order book) event stream for every market, 2026-02-22 → 2026-08-10, plus a query-ready 1-minute full-depth L2 snapshot rollup derived from it. Parquet, partitioned by UTC day, one file per day. Config What Days Size Typical file orderbook raw WebSocket event stream (book, price_change, last_trade_price, tick_size_change) 164 ~1.18 TB 8 GB (max 14 GB) orderbook_1min full L2 book at the end of… See the full description on the dataset page: https://huggingface.co/datasets/Joseph3222/polymarket-orderbook.tabular100B<n<1T1 likes4.4k downloads24d agoHugging Face23SALT-Research /DeepDialogue-orpheus DeepDialogue-orpheus DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text. 🚨 Important Notice This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.audioaudio-classification100K<n<1M8 likes4.3k downloads1y agoHugging Face24orgcatorg /multilingual Dataset Card for "multilingual" More Information needed text10M<n<100M0 likes4.2k downloads1y agoHugging Face25Open-Orca /SlimOrca Overview This is a new curated subset of our OpenOrca data. This release provides an efficient means of reaching performance on-par with using larger slices of our data, while only including ~500k GPT-4 completions. The key change in this dataset is that we've done an additional pass, using GPT-4 to remove answers which appear wrong based on the human annotations from the FLAN dataset. This reduces the dataset size to only ~500k entries, allowing training to a similar quality level… See the full description on the dataset page: https://huggingface.co/datasets/Open-Orca/SlimOrca.texttext-classification100K<n<1M300 likes4.1k downloads3y agoHugging Face26AmanPriyanshu /OR-Corpus-Copy Adopted by NVIDIA's Nemotron family of models! 🤗 HuggingFace | Slack | WeChat OpenResearcher Corpus This dataset contains a carefully curated ~11B-tokens corpus, which serves as an offline search engine for our data generation process, eliminating the need for external Search APIs. Details on the corpus curation process are available in our blog. Format Each row in the dataset contains the… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/OR-Corpus-Copy.text10M<n<100M1 likes3.7k downloads2mo agoHugging Face27oripress /AlgoTune Website  |   Paper   |   Code How good are language models at coming up with new algorithms? To try to answer this, we built a benchmark, AlgoTune, comprised of 154 widely used math, physics, and computer science functions. For each function, the goal is to write code that produces the same outputs as the original function, while being faster. In addition to the benchmark, we also provide an agent, AlgoTuner, which allows language models to easily optimize code.… See the full description on the dataset page: https://huggingface.co/datasets/oripress/AlgoTune.tabularn<1K1 likes3.7k downloads8mo agoHugging Face28open-reaction-database /ord-data ord-data Getting the Data The datasets live under data/ and are stored with Git LFS. LFS reads are redirected to the Hugging Face mirror via .lfsconfig, so dataset objects are fetched from Hugging Face's CDN rather than from GitHub's shared (and limited) LFS bandwidth. This is automatic — you do not need to configure anything. Option 1: Clone the repository git clone https://github.com/open-reaction-database/ord-data.git With Git LFS installed… See the full description on the dataset page: https://huggingface.co/datasets/open-reaction-database/ord-data.text1M<n<10M7 likes3.2k downloads27d agoHugging Face29xai-org /RealworldQA RealWorldQA RealWorldQA is a benchmark designed for real-world understanding. The dataset consists of anonymized images taken from vehicles, in addition to other real-world images. We are excited to release RealWorldQA to the community, and we intend to expand it as our multimodal models improve. The initial release of the RealWorldQA consists of over 700 images, with a question and easily verifiable answer for each image. See the announcement of Grok-1.5 Vision Preview.… See the full description on the dataset page: https://huggingface.co/datasets/xai-org/RealworldQA.imagen<1K127 likes3k downloads2y agoHugging Face30bhavyagoyal-lexsi /orpo-dstabular10K<n<100K0 likes2.9k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.