CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /c4 C4 Dataset Summary A colossal, cleaned version of Common Crawl's web crawl corpus. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's C4 dataset We prepared five variants of the data: en, en.noclean, en.noblocklist, realnewslike, and multilingual (mC4). For reference, these are the sizes of the variants: en: 305GB en.noclean: 2.3TB en.noblocklist: 380GB realnewslike: 15GB multilingual (mC4): 9.7TB (108 subsets, one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/c4.texttext-generation10B<n<100B671 likes1.2m downloads3y agoHugging Face02allenai /dolma3_mix-6T Dolma 3 Mix (6T) The Dolma 3 Mix (6T) is the collection of data used during the pretraining stage to train the Olmo-3-1125-32B model. This dataset is made up of ~6 trillion tokens from a diverse mix of web content, academic publications, code, and more. The majority of this dataset comes from Common Crawl. For more information on Dolma, please see our original release here. Smaller Sample for Analysis Available! If you would like a smaller sample of this mix… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T.text-generation36 likes78k downloads8mo agoHugging Face03allenai /dolma3.5_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3.5 pool. It contains no quality upsampling or mixing. This is an updated version of the Dolma 3 pool with additional quality filtering and more data sources. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3.5 Pool The Dolma 3.5 pool is a dataset of nearly 10 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3.5_pool.text-generation8 likes71k downloads3mo agoHugging Face04allenai /dolma3_dolmino_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 Dolmino pool; it hasn't been mixed. If you are interested in the data used to train: Olmo 3 7B: allenai/dolma3_dolmino_mix-100B-1025 Olmo 3 32B: allenai/dolma3_dolmino_mix-100B-1125 Dolma 3 Dolmino dataset pool for Olmo 3 stage 2 annealing training This dataset contains the high-quality pool of data considered for the second stage of Olmo 3 7B. Dataset Sources Source Category Tokens Documents TinyMATH Mind… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_pool.text-generation8 likes37k downloads9mo agoHugging Face05allenai /MADLAD-400 MADLAD-400 Dataset and Introduction MADLAD-400 (Multilingual Audited Dataset: Low-resource And Document-level) is a document-level multilingual dataset based on Common Crawl, covering 419 languages in total. This uses all snapshots of CommonCrawl available as of August 1, 2022. The primary advantage of this dataset over similar datasets is that it is more multilingual (419 languages), it is audited and more highly filtered, and it is document-level. The main… See the full description on the dataset page: https://huggingface.co/datasets/allenai/MADLAD-400.text-generationn>1T173 likes31k downloads2y agoHugging Face06allenai /dolma3_pool⚠️ IMPORTANT NOTICE ⚠️ This is the Dolma 3 pool, pre–quality upsampling and mixing. If you are interested in the data used to train Olmo 3 7B and Olmo 3 32B, visit allenai/dolma3_mix-6T-1025. Dolma 3 Pool The Dolma 3 pool is a dataset of over 9 trillion tokens from a diverse mix of web content, academic publications, code, and more. For detailed documenation on Dolma 3 processing and data, please see our Dolma 3 Github repository. For more information on Dolma in general… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_pool.texttext-generation10B<n<100B41 likes28k downloads7mo agoHugging Face07allenai /WildChat-1M Dataset Card for WildChat Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC-BY Language(s) (NLP): multi-lingual Point of Contact: Yuntian Deng Dataset Summary WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-1M.texttext-generation100K<n<1M464 likes28k downloads2y agoHugging Face08allenai /olmo-mix-1124 OLMo 2 (November 2024) Pretraining set Collection of data used to train OLMo-2-1124 models. The majority of this dataset comes from DCLM-Baseline with no additional filtering, but we provide the explicit breakdowns below. Name Tokens Bytes (uncompressed) Documents License DCLM-Baseline 3.70T 21.3TB 2.95B CC-BY-4.0 Arxiv 20.8B 77.2GB 3.95M ODC-BY pes2o 58.6B 412GB 38M ODC-BY starcoder 83.0B 458GB 78.7M ODC-BY Algebraic-stack 11.8B 44.0GB 2.83M ODC-BY… See the full description on the dataset page: https://huggingface.co/datasets/allenai/olmo-mix-1124.texttext-generation1B<n<10B91 likes25k downloads1y agoHugging Face09allenai /dolmino-mix-1124 DOLMino dataset mix for OLMo2 stage 2 annealing training. Mixture of high-quality data used for the second stage of OLMo2 training. Source Sizes Name Category Tokens Bytes (uncompressed) Documents License DCLM HQ Web Pages 752B 4.56TB 606M CC-BY-4.0 Flan HQ Web Pages 17.0B 98.2GB 57.3M ODC-BY Pes2o STEM Papers 58.6B 413GB 38.8M ODC-BY Wiki Encyclopedic 3.7B 16.2GB 6.17M ODC-BY StackExchange CodeText 1.26B 7.72GB 2.48M CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.tabulartext-generation100M<n<1B102 likes24k downloads11mo agoHugging Face10allenai /peS2o Pretraining Effectively on S2ORC! The peS2o dataset is a collection of ~40M creative open-access academic papers, cleaned, filtered, and formatted for pre-training of language models. It is derived from the Semantic Scholar Open Research Corpus(Lo et al, 2020), or S2ORC. We release multiple version of peS2o, each with different processing and knowledge cutoff date. We recommend you to use the latest version available. If you use this dataset, please cite: @techreport{peS2o, author =… See the full description on the dataset page: https://huggingface.co/datasets/allenai/peS2o.text-generation10B<n<100B207 likes22k downloads2y agoHugging Face11allenai /tulu-3-sft-personas-instruction-following Dataset Descriptions This dataset contains 29980 examples and is synthetically created to enhance model's capabilities to follow instructions precisely and to satisfy user constraints. The constraints are borrowed from the taxonomy in IFEval dataset. To generate diverse instructions, we expand the methodology in Ge et al., 2024 by using personas. More details and exact prompts used to construct the dataset can be found in our paper. Curated by: Allen Institute for AI Paper: TBD… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-instruction-following.texttext-generation10K<n<100K68 likes16k downloads2y agoHugging Face12allenai /dolma3_mix-6T-1025-7B ⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️ For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T. Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B. For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.texttext-generation1B<n<10B56 likes15k downloads8mo agoHugging Face13allenai /dolma3_dolmino_mix-100B-1025 Dolma 3 Dolmino Mix (100B) The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model. Dataset Sources Source Category Tokens Documents TinyMATH Mind Math (synth) 898M (0.9%) 1.52M TinyMATH PoT Math (synth) 241M (0.24%) 758K CraneMath Math (synth) 5.62B (5.63%) 7.24M MegaMatt Math (synth) 1.73B (1.73%) 3.23M Dolmino Math Math (synth) 10.7B (10.7%) 22.3M StackEdu (FIM) Code 10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.texttext-generation10M<n<100M10 likes12k downloads9mo agoHugging Face14allenai /WildChat-4.8M Dataset Card for WildChat-4.8M Dataset Description Interactive Search Tool: https://wildvisualizer.com WildChat paper: https://arxiv.org/abs/2405.01470 WildVis paper: https://arxiv.org/abs/2409.03753 Point of Contact: Yuntian Deng Dataset Summary WildChat-4.8M is a collection of 3,199,860 conversations between human users and ChatGPT. This version only contains non-toxic user inputs and ChatGPT responses, as flagged by the OpenAI Moderations API or… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat-4.8M.texttext-generation1M<n<10M198 likes9.1k downloads1y agoHugging Face15allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face16allenai /wildjailbreakgated WildJailbreak Dataset Card WildJailbreak is an open-source synthetic safety-training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both vanilla and adversarial) and 2) benign queries that resemble harmful queries in form but contain no harmful intent. Vanilla Harmful: direct… See the full description on the dataset page: https://huggingface.co/datasets/allenai/wildjailbreak.imagetext-generation1K<n<10K152 likes7.2k downloads2y agoHugging Face17allenai /coconot 🥥 CoCoNot: Contextually, Comply Not! Dataset Card Dataset Details Dataset Description Chat-based language models are designed to be helpful, yet they should not comply with every user request. While most existing work primarily focuses on refusal of "unsafe" queries, we posit that the scope of noncompliance should be broadened. We introduce a comprehensive taxonomy of contextual noncompliance describing when and how models should not comply with user… See the full description on the dataset page: https://huggingface.co/datasets/allenai/coconot.texttext-generation10K<n<100K25 likes6.3k downloads2y agoHugging Face18allenai /WildChat Dataset Card for WildChat Note: a newer version with 4.8 million conversations and demographic information can be found here. Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC-BY Language(s) (NLP): multi-lingual Point of Contact: Yuntian Deng Dataset Summary WildChat is a collection of 650K conversations between human users and ChatGPT. We collected WildChat… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildChat.texttext-generation100K<n<1M212 likes4.8k downloads1y agoHugging Face19allenai /OLMoE-mix-0924 OLMoE Mix (September 2024) The following data mix was used to train OLMoE-1B-7B, a Mixture-of-Experts LLM with 1B active and 7B total parameters released in September 2024. The base version of OLMoE-1B-7B can be found at this page, the SFT of OLMoE-1B-7B is available here, and a version combining SFT and DPO is available following this link. Statistics Subset Tokens Words Bytes Docs DCLM Baseline 1.0 3.86 T 3.38 T16.7 T 2.95 B Starcoder 101 B 63.9 B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/OLMoE-mix-0924.text-generation1B<n<10B57 likes4.7k downloads2y agoHugging Face20allenai /dolmaDolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Researchtext-generationn>1T1.1k likes3.9k downloads2y agoHugging Face21allganize /IFEval-Ko IFEval-Ko: Korean Instruction-Following Benchmark for LLMs This dataset is originated from IFEval Dataset Korean Version README IFEval-Ko is a Korean adaptation of Google's open-source IFEval benchmark utilized with lm-evaluation-harness framework. It enables evaluation of large language models (LLMs) for their instruction-following capabilities in the Korean language. Dataset Details Original Source: google/IFEvalAdaptation Author: Allganize Inc. LLM TEAM |… See the full description on the dataset page: https://huggingface.co/datasets/allganize/IFEval-Ko.texttext-generationn<1K11 likes3.4k downloads1y agoHugging Face22allenai /tulu-v2-sft-mixture Dataset Card for Tulu V2 Mix Note the ODC-BY license, indicating that different licenses apply to subsets of the data. This means that some portions of the dataset are non-commercial. We present the mixture as a research artifact. Tulu is a series of language models that are trained to act as helpful assistants. The dataset consists of a mix of : FLAN (Apache 2.0): We use 50,000 examples sampled from FLAN v2. To emphasize CoT-style reasoning, we sample another 50,000 examples… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-v2-sft-mixture.textquestion-answering100K<n<1M138 likes2.8k downloads2y agoHugging Face23allenai /bolmo_mix Bolmo Mix Data used to train Bolmo, the first family of competitive fully open byte-level language models (LMs). See our technical report for details: https://allenai.org/papers/bolmo. Name Tokens License Common Crawl 121.0B ODC-BY olmOCR Science PDFs 19.9B ODC-BY StackEdu 26.3B ODC-BY FineMath 3+ 4.1B ODC-BY arXiv 1.3B ODC-BY Wikipedia & Wikibooks 64.6M ODC-BY Character Understanding 75.5M ODC-BY Total 172.7B Bolmo models are trained for less than one… See the full description on the dataset page: https://huggingface.co/datasets/allenai/bolmo_mix.text-generation9 likes2.3k downloads9mo agoHugging Face24allenai /us-patentsThe us-patents dataset is a collection of ~ 8M US patent grants and applications from 1976-2025, cleaned, filtered, and formatted for pre-training of language models. Document Format corpus_id: Unique integer key with no semantic value. filing_date: The filing date of the grant or application. In case of duplicates, earliest filing date from the duplicate cluster. patent_type: The type of patent. text: The text content of the concatenated title, abstract, and specification.… See the full description on the dataset page: https://huggingface.co/datasets/allenai/us-patents.texttext-generation1M<n<10M14 likes2.2k downloads9mo agoHugging Face25allenai /discoverybenchData-driven Discovery Benchmark from the paper: "DiscoveryBench: Towards Data-Driven Discovery with Large Language Models" 🔭 Overview DiscoveryBench is designed to systematically assess current model capabilities in data-driven discovery tasks and provide a useful resource for improving them. Each DiscoveryBench task consists of a goal and dataset(s). Solving the task requires both statistical analysis and semantic reasoning. A faceted evaluation allows open-ended… See the full description on the dataset page: https://huggingface.co/datasets/allenai/discoverybench.texttext-generationn<1K18 likes1.8k downloads1y agoHugging Face26allenai /WildBench 🦁 WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild Loading from datasets import load_dataset wb_data = load_dataset("allenai/WildBench", "v2", split="test") Quick Links: HF Leaderboard HF Dataset Github Dataset Description License: CC BY Language(s) (NLP): English Point of Contact: Yuchen Lin WildBench is a subset of WildChat. The use of WildChat data to cause harm is strictly prohibited. Data… See the full description on the dataset page: https://huggingface.co/datasets/allenai/WildBench.tabulartext-generation1K<n<10K40 likes1.7k downloads2y agoHugging Face27allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes1.7k downloads2y agoHugging Face28AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face29allenai /quacQuestion Answering in Context is a dataset for modeling, understanding, and participating in information seeking dialog. Data instances consist of an interactive dialog between two crowd workers: (1) a student who poses a sequence of freeform questions to learn as much as possible about a hidden Wikipedia text, and (2) a teacher who answers the questions by providing short excerpts (spans) from the text. QuAC introduces challenges not found in existing machine comprehension datasets: its questions are often more open-ended, unanswerable, or only meaningful within the dialog context.question-answering10K<n<100K39 likes964 downloads3y agoHugging Face30nisten /opus-doctor-patient-conversations-all-human-diseases Opus-4.8-High-Thinking generated Doctor-Patient Conversations for All Human Diseases Covers every human disease listed on my previous work here: nisten/all-human-diseases The dataset strictly used Opus 4.8 - High and was cleaned over 3 times via Opus 4.8, 4.7 and 4.6. Minor corrections were needed upon each pass mainly to bypass single word safety filters like i.e. monkeypox. The main hallucination noticed during generation was that Opus would make up wrong PMID ( PubMed ID )… See the full description on the dataset page: https://huggingface.co/datasets/nisten/opus-doctor-patient-conversations-all-human-diseases.question-answering1K<n<10K3 likes948 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.