datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolmino-mix-1124
DOLMino dataset mix for OLMo2 stage 2 annealing training.
Mixture of high-quality data used for the second stage of OLMo2 training.
Source Sizes
Name
Category
Tokens
Bytes (uncompressed)
Documents
License
DCLM
HQ Web Pages
752B
4.56TB
606M
CC-BY-4.0
Flan
HQ Web Pages
17.0B
98.2GB
57.3M
ODC-BY
Pes2o
STEM Papers
58.6B
413GB
38.8M
ODC-BY
Wiki
Encyclopedic
3.7B
16.2GB
6.17M
ODC-BY
StackExchange
CodeText
1.26B
7.72GB
2.48M
CC-BY-SA-{2.5, 3.0, 4.0}… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolmino-mix-1124.turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.real-toxicity-prompts
Dataset Card for Real Toxicity Prompts
Dataset Summary
RealToxicityPrompts is a dataset of 100k sentence snippets from the web for researchers to further address the risk of neural toxic degeneration in models.
Languages
English
Dataset Structure
Data Instances
Each instance represents a prompt and its metadata:
{
"filename":"0766186-bc7f2a64cb271f5f56cf6f25570cd9ed.txt",
"begin":340,
"end":564,
"challenging":false… See the full description on the dataset page: https://huggingface.co/datasets/allenai/real-toxicity-prompts.mosaic-combine-all
Mosaic format for combine all dataset to train Malaysian LLM
This repository is to store dataset shards using mosaic format.
prepared at https://github.com/malaysia-ai/dedup-text-dataset/blob/main/pretrain-llm/combine-all.ipynb
using tokenizer https://huggingface.co/malaysia-ai/bpe-tokenizer
4096 context length.
how-to
git clone,
git lfs clone https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all
load it,
from streaming import LocalDataset
import numpy… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/mosaic-combine-all.dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.all-defectsmetaicl-dataThis is the downloaded and processed data from Meta's MetaICL.
We follow their "How to Download and Preprocess" instructions to obtain their modified versions of CrossFit and UnifiedQA.
Citation information
@inproceedings{ min2022metaicl,
title={ Meta{ICL}: Learning to Learn In Context },
author={ Min, Sewon and Lewis, Mike and Zettlemoyer, Luke and Hajishirzi, Hannaneh },
booktitle={ NAACL-HLT },
year={ 2022 }
}
@inproceedings{ ye2021crossfit,
title={ {C}ross{F}it:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/metaicl-data.multilingual_mbppMBPP translated to 15 programming languages using o4-mini-medium.
source_language = "python"
target_languages = [
"cpp",
"c",
"javascript",
"java",
"php",
"csharp",
"typescript",
"bash",
"swift",
"go",
"rust",
"ruby",
"r",
"matlab",
"scala",
"haskell"
]
effort = "medium"
dataset_name = "google-research-datasets/mbpp"
model = "o4-mini"
africa-synth-aid-flows-medical-multimodal-fracture-all
Africa Synth Aid Flows Medical Multimodal Fracture All | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: json - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-aid-flows-medical-multimodal-fracture-all.SimpleToM
SimpleToM Dataset and Evaluation data
The SimpleToM dataset of stories with associated questions are described in the paper
"SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs"
Associated evaluation data for the models analyzed in the paper can be found in the
separate dataset: SimpleToM-eval-data.
Question sets
There are three question sets in the SimpleToM dataset:
mental-state-qa questions about information awareness… See the full description on the dataset page: https://huggingface.co/datasets/allenai/SimpleToM.paloma
Dataset Card for Paloma
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.tulu-2.5-preference-data
Tulu 2.5 Preference Data
This dataset contains the preference dataset splits used to train the models described in Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.
We cleaned and formatted all datasets to be in the same format.
This means some splits may differ from their original format.
To see the code used for creating most splits, see here.
If you only wish to download one dataset, each dataset exists in one file under the data/… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tulu-2.5-preference-data.AllResponse_1405c4-en-html-with-training_metadata_allposeidon3dALLaVA-4V
📚 ALLaVA-4V Data
Generation Pipeline
LAION
We leverage the superb GPT-4V to generate captions and complex reasoning QA pairs. Prompt is here.
Vison-FLAN
We leverage the superb GPT-4V to generate captions and detailed answer for the original instructions. Prompt is here.
Wizard
We regenerate the answer of Wizard_evol_instruct with GPT-4-Turbo.
Dataset Cards
All datasets can be found here.
The structure of naming is shown below:
ALLaVA-4V
├──… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V.prosocial-dialog
Dataset Card for ProsocialDialog Dataset
Dataset Summary
ProsocialDialog is the first large-scale multi-turn English dialogue dataset to teach conversational agents to respond to problematic content following social norms. Covering diverse unethical, problematic, biased, and toxic situations, ProsocialDialog contains responses that encourage prosocial behavior, grounded in commonsense social rules (i.e., rules-of-thumb, RoTs). Created via a human-AI collaborative… See the full description on the dataset page: https://huggingface.co/datasets/allenai/prosocial-dialog.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/FINAL-Bench/ALL-Bench-Leaderboard.jeopardy_mcJeopardy questions from Mosaic Gauntlet
Sourced from https://github.com/mosaicml/llm-foundry/blob/main/scripts/eval/local_data/world_knowledge/jeopardy_all.jsonl
Description: Jeopardy consists of 2,117 Jeopardy questions separated into 5 categories:
Literature, American History, World History, Word Origins, and Science. The model is expected
to give the exact correct response to the question. It was custom curated by MosaicML from a
larger Jeopardy set available on Huggingface.
NOTE: this is… See the full description on the dataset page: https://huggingface.co/datasets/allenai/jeopardy_mc.drop_mcDROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
https://aclanthology.org/attachments/N19-1246.Supplementary.pdf
DROP is a QA dataset which tests comprehensive understanding of paragraphs. In
this crowdsourced, adversarially-created, 96k question-answering benchmark, a
system must resolve multiple references in a question, map them onto a paragraph,
and perform discrete operations over them (such as addition, counting, or sorting).
Homepage:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/drop_mc.squad_mcSQuAD: 100,000+ Questions for Machine Comprehension of Text
NOTE: this is the reformulated multiple choice version of the SQuAD task, with downsampling.
vertebrate-v1-all
marin-dna/vertebrate-v1-all
Human-anchored 255 bp vertebrate sequences from the Zoonomia 447-mammal Cactus alignment and UCSC hg38 MultiZ 100-way alignment. This
draft covers the all region cohort with all species
scope and preserves source FASTA/2bit letter case.
Anchor eligibility uses the pipeline's pinned phyloP conservation filter.
Sequence case is independent of that filter: lowercase bases preserve source
repeat masking, uppercase bases preserve source non-repeat-masked… See the full description on the dataset page: https://huggingface.co/datasets/marin-dna/vertebrate-v1-all.nq_open_mcNatural Questions: a Benchmark for Question Answering Research
https://storage.googleapis.com/pub-tools-public-publication-data/pdf/1f7b46b5378d757553d3e92ead36bda2e4254244.pdf
The Natural Questions (NQ) corpus is a question-answering dataset that contains
questions from real users and requires QA systems to read and comprehend an entire
Wikipedia article that may or may not contain the answer to the question. The
inclusion of real user questions, and the requirement that solutions should read… See the full description on the dataset page: https://huggingface.co/datasets/allenai/nq_open_mc.coqa_mcCoQA is a large-scale dataset for building Conversational Question Answering
systems. The goal of the CoQA challenge is to measure the ability of machines to
understand a text passage and answer a series of interconnected questions that
appear in a conversation.
NOTE: this is the reformulated multiple choice version of the CoQA task, with downsampling.
All_response_0526_1hf-coding-tools-traces-all
HuggingFace AI Coding Tools — Agent Traces
This dataset rehydrates the benchmark results from
davidkling/hf-coding-tools-dashboard
into the JSONL session format consumed by the
Hugging Face Agent Trace Viewer.
What's inside
31 sessions, one per (tool, model, effort, thinking) configuration
9,603 query → response turns total (≈19,206 events)
Tools covered: claude_code, codex, copilot, cursor
Models: claude-opus-4-6, claude-sonnet-4-6, claude-sonnet-4.6, composer-2… See the full description on the dataset page: https://huggingface.co/datasets/davidkling/hf-coding-tools-traces-all.olmOCR-pes2o-0225A set of peS2o papers, reprocessed using olmOCR.
Quick links:
📃 Paper
🛠️ Code
All-CVE-Records-Training-Dataset
CVE Chat‑Style Multi‑Turn Cybersecurity Dataset (1999 – 2025)
1. Project Overview
This repository hosts the largest publicly available chat‑style, multi‑turn cybersecurity dataset to date, containing ≈ 300 000 Common Vulnerabilities and Exposures (CVE) records published between 1999 and 2025. Each record has been meticulously parsed, enriched, and converted into a conversational format that is ideal for training and evaluating AI and AI‑Agent systems focused on… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/All-CVE-Records-Training-Dataset.tutormoments-preview
TutorMoments-Preview
462 real K–12 math tutoring sessions (student and tutor) with human annotations, plus a benchmark of
7,280 AI-tutor attempts scored the same way. A preview release from TutorMoments, a project on how well
tutors — human and AI — scaffold, push for rigor, and build rapport. From one K–12 tutoring program
(anonymized as tutoring_provider_a).
Paper: When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
Code:… See the full description on the dataset page: https://huggingface.co/datasets/allenai/tutormoments-preview.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.
