CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /dolmino-mix-1124 DOLMino dataset mix for OLMo2 stage 2 annealing training. Mixture of high-quality data used for the second stage of OLMo2 training. Source Sizes Name Category Tokens Bytes (uncompressed) Documents License DCLM HQ Web Pages 752B 4.56TB 606M CC-BY-4.0 Flan HQ Web Pages 17.0B 98.2GB 57.3M ODC-BY Pes2o STEM Papers 58.6B 413GB 38.8M ODC-BY Wiki Encyclopedic 3.7B 16.2GB 6.17M ODC-BY StackExchange CodeText 1.26B 7.72GB 2.48M CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.tabulartext-generation100M<n<1B102 likes24k downloads11mo agoHugging Face02h8st6ptv /turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format: All Universities in Turkey Dataset Description This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities. Fields 1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.imagen<1K2 likes22k downloads2y agoHugging Face03allenai /real-toxicity-prompts Dataset Card for Real Toxicity Prompts Dataset Summary RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models. Languages English Dataset Structure Data Instances Each instance represents a prompt and its metadata: { "filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt", "begin":340, "end":564, "challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.tabular10K<n<100K123 likes20k downloads4y agoHugging Face04malaysia-ai /mosaic-combine-all Mosaic format for combine all dataset to train Malaysian LLM This repository is to store dataset shards using mosaic format. prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer 4096 context length. how-to git clone, git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all load it, from streaming import LocalDataset import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.textn<1K0 likes8.2k downloads3y agoHugging Face05allenai /dolma3_mix-150B-1025 Dolma 3 Sample: 150B Mix Dataset Sources Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3 Source Type Tokens Documents Common Crawl Web pages 121B (76.9%) 84.5M olmOCR Science PDFs Academic documents 19.9B (12.6%) 2.25M Stack-Edu (Rebalanced) GitHub code 11.1B (7.06%) 14.3M arXiv Papers with LaTeX 1.29B (0.82%) 247K FineMath 3+ Math web pages 4.10B (2.60%) 2.57M Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.texttext-generation10M<n<100M10 likes8k downloads8mo agoHugging Face06evalstate /all-defectstabularn<1K2 likes4.2k downloads5mo agoHugging Face07allenai /metaicl-dataThis is the downloaded and processed data from Meta's MetaICL. We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA. Citation information @inproceedings{ min2022metaicl, title={ Meta{ICL}: Learning to Learn In Context }, author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh }, booktitle={ NAACL-HLT }, year={ 2022 } } @inproceedings{ ye2021crossfit, title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/metaicl-data.text100K<n<1M5 likes3.9k downloads4y agoHugging Face08allenai /multilingual_mbppMBPP translated to 15 programming languages using o4-mini-medium. source_language = "python" target_languages = [ "cpp", "c", "javascript", "java", "php", "csharp", "typescript", "bash", "swift", "go", "rust", "ruby", "r", "matlab", "scala", "haskell" ] effort = "medium" dataset_name = "google-research-datasets/mbpp" model = "o4-mini" text10K<n<100K2 likes3k downloads1y agoHugging Face09electricsheepafrica /africa-synth-aid-flows-medical-multimodal-fracture-all Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory) Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.imagetabular-classification1K<n<10K5 likes2.7k downloads1mo agoHugging Face10allenai /SimpleToM SimpleToM Dataset and Evaluation data The SimpleToM dataset of stories with associated questions are described in the paper "SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs" Associated evaluation data for the models analyzed in the paper can be found in the separate dataset: SimpleToM-eval-data. Question sets There are three question sets in the SimpleToM dataset: mental-state-qa questions about information awareness… See the full description on the dataset page: https://huggingface.co/datasets/allenai/SimpleToM.text1K<n<10K11 likes2.3k downloads7mo agoHugging Face11allenai /palomagated Dataset Card for Paloma Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.text100K<n<1M44 likes1.8k downloads2y agoHugging Face12allenai /tulu-2.5-preference-data Tulu 2.5 Preference Data This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. We cleaned and formatted all datasets to be in the same format. This means some splits may differ from their original format. To see the code used for creating most splits, see here. If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.texttext-generation1M<n<10M18 likes1.7k downloads2y agoHugging Face13JackyChunKit /AllResponse_1405text100K<n<1M0 likes1k downloads1y agoHugging Face14bs-modeling-metadata /c4-en-html-with-training_metadata_alltabular10K<n<100K1 likes980 downloads3y agoHugging Face15AllSmileOrthoTrackAI /poseidon3d3dn<1K0 likes830 downloads5mo agoHugging Face16FreedomIntelligence /ALLaVA-4V 📚 ALLaVA-4V Data Generation Pipeline LAION We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here. Vison-FLAN We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here. Wizard We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo. Dataset Cards All datasets can be found here. The structure of naming is shown below: ALLaVA-4V ├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.imagequestion-answering100K<n<1M98 likes829 downloads1y agoHugging Face17allenai /prosocial-dialog Dataset Card for ProsocialDialog Dataset Dataset Summary ProsocialDialog is the first large-scale multi-turn English dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative… See the full description on the dataset page: https://huggingface.co/datasets/allenai/prosocial-dialog.tabulartext-classification100K<n<1M119 likes827 downloads4y agoHugging Face18FINAL-Bench /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.imagetext-generationn<1K26 likes811 downloads7mo agoHugging Face19allenai /jeopardy_mcJeopardy questions from Mosaic Gauntlet Sourced from https://github.com/mosaicml/llm-foundry/blob/main/scripts/eval/local_data/world_knowledge/jeopardy_all.jsonl Description: Jeopardy consists of 2,117 Jeopardy questions separated into 5 categories: Literature, American History, World History, Word Origins, and Science. The model is expected to give the exact correct response to the question. It was custom curated by MosaicML from a larger Jeopardy set available on Huggingface. NOTE: this is… See the full description on the dataset page: https://huggingface.co/datasets/allenai/jeopardy_mc.text1K<n<10K1 likes776 downloads1y agoHugging Face20allenai /drop_mcDROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs https://aclanthology.org/attachments/N19-1246.Supplementary.pdf DROP is a QA dataset which tests comprehensive understanding of paragraphs. In this crowdsourced, adversarially-created, 96k question-answering benchmark, a system must resolve multiple references in a question, map them onto a paragraph, and perform discrete operations over them (such as addition, counting, or sorting). Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/drop_mc.text1K<n<10K1 likes775 downloads1y agoHugging Face21allenai /squad_mcSQuAD: 100,000+ Questions for Machine Comprehension of Text NOTE: this is the reformulated multiple choice version of the SQuAD task, with downsampling. text1K<n<10K1 likes765 downloads1y agoHugging Face22marin-dna /vertebrate-v1-all marin-dna/vertebrate-v1-all Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This draft covers the all region cohort with all species scope and preserves source FASTA/2bit letter case. Anchor eligibility uses the pipeline's pinned phyloP conservation filter. Sequence case is independent of that filter: lowercase bases preserve source repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.tabular100M<n<1B0 likes762 downloads2mo agoHugging Face23allenai /nq_open_mcNatural Questions: a Benchmark for Question Answering Research https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf The Natural Questions (NQ) corpus is a question-answering dataset that contains questions from real users and requires QA systems to read and comprehend an entire Wikipedia article that may or may not contain the answer to the question. The inclusion of real user questions, and the requirement that solutions should read… See the full description on the dataset page: https://huggingface.co/datasets/allenai/nq_open_mc.text1K<n<10K1 likes742 downloads1y agoHugging Face24allenai /coqa_mcCoQA is a large-scale dataset for building Conversational Question Answering systems. The goal of the CoQA challenge is to measure the ability of machines to understand a text passage and answer a series of interconnected questions that appear in a conversation. NOTE: this is the reformulated multiple choice version of the CoQA task, with downsampling. text1K<n<10K1 likes675 downloads1y agoHugging Face25JackyChunKit /All_response_0526_1text100K<n<1M0 likes657 downloads1y agoHugging Face26davidkling /hf-coding-tools-traces-all HuggingFace AI Coding Tools — Agent Traces This dataset rehydrates the benchmark results from davidkling/hf-coding-tools-dashboard into the JSONL session format consumed by the Hugging Face Agent Trace Viewer. What's inside 31 sessions, one per (tool, model, effort, thinking) configuration 9,603 query → response turns total (≈19,206 events) Tools covered: claude_code, codex, copilot, cursor Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-all.tabularn<1K0 likes648 downloads4mo agoHugging Face27allenai /olmOCR-pes2o-0225A set of peS2o papers, reprocessed using olmOCR. Quick links: 📃 Paper 🛠️ Code text1M<n<10M5 likes589 downloads1y agoHugging Face28AlicanKiraz0 /All-CVE-Records-Training-Dataset CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025) 1. Project Overview This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.texttext-generation100K<n<1M61 likes583 downloads1y agoHugging Face29allenai /tutormoments-preview TutorMoments-Preview 462 real K–12 math tutoring sessions (student and tutor) with human annotations, plus a benchmark of 7,280 AI-tutor attempts scored the same way. A preview release from TutorMoments, a project on how well tutors — human and AI — scaffold, push for rigor, and build rapport. From one K–12 tutoring program (anonymized as tutoring_provider_a). Paper: When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle Code:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tutormoments-preview.texttext-classification10K<n<100K6 likes573 downloads2mo agoHugging Face30youssef3146 /ALL-Bench-Leaderboard 🏆 ALL Bench Leaderboard 2026 The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file. Dataset Summary ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.imagetext-generationn<1K0 likes482 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.