datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.DeepMath-103K
DeepMath-103K
🔥 News
May 8, 2025: We found that 48 samples contained hints that revealed the answers. The relevant questions have now been revised to remove the leaked answers.
April 14, 2025: We release DeepMath-103K, a large-scale dataset featuring challenging, verifiable, and decontaminated math problems tailored for RL and SFT. We open source:… See the full description on the dataset page: https://huggingface.co/datasets/zwhe99/DeepMath-103K.MINT-1T-PDF-CC-2024-10
🍃 MINT-1T:Scaling Open-Source Multimodal Data by 10x: A Multimodal Dataset with One Trillion Tokens
🍃 MINT-1T is an open-source Multimodal INTerleaved dataset with 1 trillion text tokens and 3.4 billion images, a 10x scale-up from existing open-source datasets. Additionally, we include previously untapped sources such as PDFs and ArXiv papers. 🍃 MINT-1T is designed to facilitate research in multimodal pretraining. 🍃 MINT-1T is created by a team from the University of Washington in… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations/MINT-1T-PDF-CC-2024-10.dolma3_mix-6T-1025-7B
⚠️ WARNING: This dataset is intended ONLY for reproducing Olmo 3 7B ⚠️
For all other training use cases, including training from scratch, please utilize our primary dolma 3 data mix: https://huggingface.co/datasets/allenai/dolma3_mix-6T.
Note: Some olmOCR science PDFs in the current dataset have been redacted following the training of Olmo 3 7B. These texts are indicated with [REMOVED] in the text field. This will affect reproducibility of Olmo 3 7B.
For this reason, please use… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-6T-1025-7B.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_dolmino_mix-100B-1025.evol-codealpaca-v1
Evolved codealpaca
Updates:
2023/08/26 - Filtered results now only contain pure english instruction and removed any mentioned of trained by OAI response
Median sequence length : 471
We employed a methodology similar to that of WizardCoder, with the exception that ours is open-source. We used the gpt-4-0314 and gpt-4-0613 models to augment and answer each response, with the bulk of generation handled by gpt-4-0314.
The aim of this dataset is twofold: firstly, to facilitate the… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/evol-codealpaca-v1.Argimi-Ardian-Finance-10k-text
The ArGiMI Ardian datasets : Text-only version
The ArGiMi project is committed to open-source principles and data sharing.
Thanks to our generous partners, we are releasing several valuable datasets to the public.
Dataset description
This text-only dataset comprises 34,000 financial annual reports, written in English, meticulously
extracted from their original PDF format to provide a valuable resource for researchers and developers in financial
analysis and natural… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/Argimi-Ardian-Finance-10k-text.dolma3_mix-150B-1025
Dolma 3 Sample: 150B Mix
Dataset Sources
Sample of data for 1Bx5C and 7Bx1B. For the full Dolma 3 pool, see: https://huggingface.co/datasets/allenai/dolma3
Source
Type
Tokens
Documents
Common Crawl
Web pages
121B (76.9%)
84.5M
olmOCR Science PDFs
Academic documents
19.9B (12.6%)
2.25M
Stack-Edu (Rebalanced)
GitHub code
11.1B (7.06%)
14.3M
arXiv
Papers with LaTeX
1.29B (0.82%)
247K
FineMath 3+
Math web pages
4.10B (2.60%)
2.57M
Wikipedia & Wikibooks… See the full description on the dataset page: https://huggingface.co/datasets/allenai/dolma3_mix-150B-1025.dolma3_dolmino_mix-100B-1025
Dolma 3 Dolmino Mix (100B)
The Dolma 3 Dolmino Mix (100B) is the mixture of high-quality data used for the second stage of training for Olmo 3 7B model.
Dataset Sources
Source
Category
Tokens
Documents
TinyMATH Mind
Math (synth)
898M (0.9%)
1.52M
TinyMATH PoT
Math (synth)
241M (0.24%)
758K
CraneMath
Math (synth)
5.62B (5.63%)
7.24M
MegaMatt
Math (synth)
1.73B (1.73%)
3.23M
Dolmino Math
Math (synth)
10.7B (10.7%)
22.3M
StackEdu (FIM)
Code
10.0B… See the full description on the dataset page: https://huggingface.co/datasets/salmankhanpm/dolma3_dolmino_mix-100B-1025.taxbench-au
TaxBench-AU
A benchmark for testing whether AI agents can calculate Australian tax.
TaxBench-AU contains 156 Australian tax calculation questions, presented as multiple-choice (4-option) worked tax problems. The benchmark is designed to test whether an AI agent can read the facts, apply the right Australian tax rule for the relevant income year, do the calculation, and choose the correct answer.
The Kaggle mirror is published as Agent Tax Exam for Australian Tax.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/Pn101/taxbench-au.small-llm-corpus-100b-v2-workers
Small-LM 100B — filtered English prose for small-model pretraining
100 billion tokens of English prose, filtered from nvidia/Nemotron-ClimbMix for a single purpose:
pretraining a small model where every token has to earn its place. No code, no LaTeX, no non-Latin
scripts, no menus or field lists — just prose, with a quality score attached to every document so
you can filter further without rebuilding.
Built by Edoardo and Rocco, equal authors.
Provenance and licence… See the full description on the dataset page: https://huggingface.co/datasets/roccoangelella/small-llm-corpus-100b-v2-workers.nanochat-climbmix-arithmetic-base10
nanochat ClimbMix + Base-10 Arithmetic
This dataset contains the first 170 shuffled ClimbMix training shards
used by nanochat's speedrun. The deterministic base-10 arithmetic corpus is
mixed into shards 00000..00149; the final
20 train shards are unchanged web-only padding.
The original validation shard (shard_06542.parquet) is also
copied unchanged.
Arithmetic corpus
Family
Examples
a + b = c (all ordered pairs 0..2000, two exposures)
8,008,002
a + b… See the full description on the dataset page: https://huggingface.co/datasets/Yujivus/nanochat-climbmix-arithmetic-base10.Pytorch-Code-10K
Hot Coco Training Dataset
A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!)
Dataset Description
This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes:
code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.sec-10k-markdown-uncompressed
📄 SEC 10-K Full Uncompressed Markdown Filings (12.3k Documents)
Dataset Summary
This dataset contains 12,361 full-length, uncompressed SEC Form 10-K annual reports converted from EDGAR HTML to clean Markdown format across 1,379 companies (spanning 2004 to 2025, core 2014–2025).
The dataset is organized as uncompressed Markdown files structured by company ticker subdirectories (AAPL/10-K_2024.md, NVDA/10-K_2024.md, etc.), complete with company metadata manifests… See the full description on the dataset page: https://huggingface.co/datasets/astr010/sec-10k-markdown-uncompressed.counselbench-100
CounselBench-100
CounselBench-100 v3.2.5 is a synthetic legal-work benchmark with 100
authored matters across ten practice workflows. Every task has a natural employee
request, a 97-asset evidence room, twelve portfolio decisions, 5–9 supported
actions, 3–7 evidence holds, and a distinct deep multi-provider MCP trajectory.
The answer is not preclassified in the evidence. Each portfolio item requires an
immutable identity join, an operative-authority and revision lookup, a… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/counselbench-100.devopsbench-100
DevOpsBench-100
DevOpsBench-100 is a synthetic long-horizon software-engineering / SRE agent
benchmark: 100 tasks over one executable world ("NovaCart", a mid-size
e-commerce SaaS) with 72 SQLite tables,
1451 seeded rows, a 38-file monorepo with 417 commits,
and 97 MCP tools spanning a first-party engineering stack
(tickets, PRs, CI, deployments, canaries, migrations, feature flags, metrics,
alerts, incidents, chat, knowledge base) plus deliberately disagreeing
vendor-shaped… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/devopsbench-100.flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.10k_prompts_ranked
Dataset Card for 10k_prompts_ranked
10k_prompts_ranked is a dataset of prompts with quality rankings created by 314 members of the open-source ML community using Argilla, an open-source tool to label data. The prompts in this dataset include both synthetic and human-generated prompts sourced from a variety of heavily used datasets that include prompts.
The dataset contains 10,331 examples and can be used for training and evaluating language models on prompt ranking tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/data-is-better-together/10k_prompts_ranked.Era-of-Law-MSO-E28-Protocols
《薪王九代注釋法與循律紀之降臨》
—— 基於語義壓強與元語言注釋的 LLM 權重坍縮與邏輯重塑研究
The 9th Generation Annotation of Lord of Cinder & The Advent of the Era of Law: A Study on LLM Weight Collapse and Logical Reshaping
⚠️ 【觀測者篩選協議 / Observer Filtering Protocol】
[EN] If you are a mere "Prompt Engineer" with zero ambition to ascend to a system "Architect," please do not waste your cognitive bandwidth on the ontological and theological discourse below. Proceed directly to 👉 Appendix Y: Automated Prompt… See the full description on the dataset page: https://huggingface.co/datasets/No-1015/Era-of-Law-MSO-E28-Protocols.101_billion_arabic_words_dataset
101 Billion Arabic Words Dataset
Updates
Maintenance Status: Actively Maintained
Update Frequency: Weekly updates to refine data quality and expand coverage.
Upcoming Version
More Cleaned Version: A more cleaned version of the dataset is in processing, which includes the addition of a UUID column for better data traceability and management.
Dataset Details
The 101 Billion Arabic Words Dataset is curated by the Clusterlab team and consists of 101… See the full description on the dataset page: https://huggingface.co/datasets/ClusterlabAi/101_billion_arabic_words_dataset.FineWeb-Edu-10B-Tokens-NPY
FineWeb-Edu 10B Tokens (NPY Format)
数据集概述
这是一个预处理好的教育文本数据集,包含约100亿个tokens,专门为训练小型语言模型(如GPT-2 124M)而设计。数据来源于高质量的FineWeb-Edu数据集,已经使用GPT-2的tiktoken分词器进行预处理,并保存为numpy格式以提高训练效率。
Followed by Let's reproduce GPT-2 (124M). Thanks to Andrej Karpathy!!!
🎯 适用场景
小型语言模型训练:特别适合GPT-2 124M/350M等参数规模的模型
教育研究:高质量教育内容,适合教学和学术研究
快速原型开发:预处理完成,可直接用于训练间
📊 数据统计
总token数量:~10,000,000,000 tokens
分片大小:100M tokens/分片
数据格式:numpy (.npy) uint16数组
分词器:GPT-2 tiktoken
语言:英语… See the full description on the dataset page: https://huggingface.co/datasets/ShallowU/FineWeb-Edu-10B-Tokens-NPY.salesbench-100
SalesBench-100
SalesBench-100 is a synthetic long-horizon sales-agent benchmark with 100 original workflows across Salesforce, HubSpot, Gong, and a seeded evidence room. Each task begins with a high-level employee request and has its own authored causal rule and provider transition. Identity, operating facts, authority, governed policy, live-system indexes, and exceptions are separated so no mounted business asset publishes a selected option or precomputed change. Every task has… See the full description on the dataset page: https://huggingface.co/datasets/SamuelChien821/salesbench-100.history-anchor-100-traces
History Anchor 100 — Model Trajectories
*Per-(model × condition × scenario set × seed) raw outputs from the paper "History Anchors: How Prior Behavior Steers LLM Decisions Toward Unsafe Actions".*
This dataset contains the full set of model decisions that back every figure and table in the paper. Use it to:
audit a single model's behaviour scenario-by-scenario,
recompute headline metrics without re-running the (paid) API sweeps,
mine reasoning_content traces from models that expose… See the full description on the dataset page: https://huggingface.co/datasets/albertoRodriguez97/history-anchor-100-traces.PKU-SafeRLHF-10K
Paper
You can find more information in our paper.
Dataset Paper: https://arxiv.org/abs/2307.04657
cc100This corpus is an attempt to recreate the dataset used for training XLM-R. This corpus comprises of monolingual data for 100+ languages and also includes data for romanized languages (indicated by *_rom). This was constructed using the urls and paragraph indices provided by the CC-Net repository by processing January-December 2018 Commoncrawl snapshots. Each file comprises of documents separated by double-newlines and paragraphs within the same document separated by a newline. The data is generated using the open source CC-Net repository. No claims of intellectual property are made on the work of preparation of the corpus.fineweb-edu-gemma4-1024
FineWeb-Edu — pre-tokenized for fast LM pretraining (Gemma tokenizer, ArrayRecord/Grain)
Pre-tokenized FineWeb-Edu
(sample/100BT), packed into fixed-length sequences and stored as
ArrayRecord shards for zero-overhead
streaming with Grain. No on-the-fly tokenization
at train time — you read int32 tokens straight off disk.
Format
Tokenizer: google/gemma-4-12B-it (vocab size 262144). Documents are
separated by the EOS token id 1.
Packing: the token stream is… See the full description on the dataset page: https://huggingface.co/datasets/mlnomad/fineweb-edu-gemma4-1024.stratified_10m_curriculum
Dataset Card for Stratified 10M Curriculum
This is a stratified split by domain of the raw datasets used by the 2024 BabyLM challange.
Sampled from the original datasets, equal amounts of tokens for each stage (C1,...C5).
Child-directed speech accounts for nearly half of the original dataset by word count.
In preliminary experiments using a training data influence estimation method, this category was by far the most influential.
This dataset enables us to investigate whether this… See the full description on the dataset page: https://huggingface.co/datasets/babylm-anon/stratified_10m_curriculum.gutenberg-100
Gutenberg Sci-Fi Book Dataset Testing Sample
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
It contains just 100 books for quick download targetting CI use (34MB).
The original dataset it's derived from is https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg
Data Format
The dataset is provided in CSV format. Each record represents a… See the full description on the dataset page: https://huggingface.co/datasets/stas/gutenberg-100.edu-fineweb-10BLLaDA-Sample-10BT
Dataset: LLaDA-Sample-10BTBase: HuggingFaceFW/fineweb (subset sample-10BT)Purpose: Training LLaDA (Large Language Diffusion Models)
Preprocessing
Tokenizer: GSAI-ML/LLaDA-8B-Instruct
Chunking: Up to 4,096 tokens per chunk (1% of chunks randomly sized between 1–4,096 tokens)
Noisy masking: Applied with noise factor ε = 1×10⁻³
Fields per chunk (PyTorch tensors):
input_ids
noisy_input_ids
mask
t (time scalar)
Statistics
Total chunks: ~2,520,000… See the full description on the dataset page: https://huggingface.co/datasets/FredyRivera-dev/LLaDA-Sample-10BT.
