datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
common_corpus
Common Corpus
Full paper - ICLR 2026 oral
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Common Corpus differs from existing open datasets in that it is:
Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.gneissweb-annotation-url-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-url-testing-v1.thai-commoncrawl-index
Thai Common Crawl Index (2019–2026)
An index of every page Common Crawl detected as Thai across 70 monthly crawls, from
January 2019 (CC-MAIN-2019-04) to August 2026 (CC-MAIN-2026-30).
932,874,727 page captures · 450,971,497 unique URLs · 6,997,185 hosts · 6,674,969 domains
Each row records where the page lives inside Common Crawl's WARC archives — file name,
byte offset, and record length — so you can fetch exactly the pages you want with HTTP
range requests, without scanning… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-commoncrawl-index.host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.common_corpus
Common Corpus
Full paper - ICLR 2026 oral
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Common Corpus differs from existing open datasets in that it is:
Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/oahegiaerhg/common_corpus.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.Medical-Commons
Medical-Commons
Medical-Commons is the largest dataset of medical content under free licenses or open data program collected by Pleias.
It includes three different collection:
International scientific collection of 2M articles from OpenAlex.
French scientific collection of XM articles, reports and PhD theses from French institutional repositories.
Administration collection from health and medical agencies, for now limited to France but with a planned Europe-wide expansion.
The… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Medical-Commons.wav2vec2_common_voice_accents_3common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.agent-traces
Trace Commons — Agent Traces
Trace Commons is one open, public dataset of coding-agent sessions — the
back-and-forth between a developer and an AI coding agent, including prompts,
model responses, tool calls, and command output — contributed voluntarily as an
open resource for studying, evaluating, and building on how these agents
actually work.
Every trace here was donated only from a public, open-source repository, was
anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.common_corpus_nl
Common Corpus v2 NL
This is a version of Common Corpus v2 filtered to keep only the rows where language is "Dutch".
Common Corpus is a very large open and permissible licensed text dataset created by Pleias.
Please be sure to acknowledge the creators of the original dataset when using this filtered version.
Filtering
Common Corpus is a collection of disparate datasets.
Note that filtering the entire collection for rows where the language is "Dutch" is not the same as… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_nl.French-Science-Commons
French Science Commons
French Science Commons (Commun numérique des sciences en français) rassemble des publications scientifiques d'origine française en accès ouvert, couvrant une période de vingt ans, de 2007 à 2026. Il comprend 1 248 860 documents scientifiques — 1 189 628 articles et 59 232 thèses — indexés à travers de multiples dépôts académiques en accès public, tels que HAL, OpenAlex, des revues scientifiques, des dépôts institutionnels, et d'autres.
Le corpus est conçu… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/French-Science-Commons.wikimedia-commons-documents-ml_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-documents-ml_beir.aligned-mwe
aligned_mwe — multi-word target expressions per lexeme
Where lexeme-alignments is one row per surface
token, this is one row per lexeme rendered by a contiguous multi-word phrase (חֶסֶד → "kasih setia",
בֵּית → "tempat pengirikan"). Mined from the aligner's per-verse t_idx positions: only spans whose
target token positions are contiguous (max−min+1 == len) qualify — scattered tokens that merely all
linked to a lexeme are dropped (and counted in the manifest as scattered_dropped… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/aligned-mwe.FRENCH-ONLY-Common-Crawl-2026-25gneissweb-annotation-host-testing-v1
GneissWeb Annotations
GneissWeb Annotations, powered by IBM Research's GneissWeb methodology, is a dataset of quality and category annotations applied to the Common Crawl corpus.
This dataset enables precise filtering of web content across medical, educational, technology, and scientific domains, making it easier to build high-quality corpora for research projects, language models, and specialized applications.
Learn more about the annotation process and methodology in our… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/gneissweb-annotation-host-testing-v1.EU-Science-CommonsTelco-Common-Corpus Telco Common Corpus (TCC) is a ten billion tokens collection of fully open, free licensed telecommunications knowledge (scientific literature, patents, open data, and open-web projects) with licence and provenance verified at a document-level.
TCC stems from GSMA's effort to make AI work for the telecom sector. The Open-Telco LLM Benchmarks and the broader Open Telco AI initiative have already established that current models fall short on real telecom tasks, including network management and… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/Telco-Common-Corpus.common_voice_22_0
Common Voice Corpus 22.0
Originally from https://huggingface.co/datasets/fsicoli/common_voice_22_0, we mirror using multiple zip files also trimmed the silents.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_22_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/common_voice_22_0.Telco-Common-Corpus Telco Common Corpus (TCC) is a ten billion tokens collection of fully open, free licensed telecommunications knowledge (scientific literature, patents, open data, and open-web projects) with licence and provenance verified at a document-level.
TCC stems from GSMA's effort to make AI work for the telecom sector. The Open-Telco LLM Benchmarks and the broader Open Telco AI initiative have already established that current models fall short on real telecom tasks, including network management and… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/Telco-Common-Corpus.lexeme-alignments
lexeme-alignments — surface → original-language lexeme (Strong's-bridged)
For each language, the attested mapping from target surface word-forms → the original-language
lexeme they render, mined by the aligner. Lexeme-anchored, provenance-honest, additive — the
design principles are in docs/publishing-principles.md. One
language per partition, for consumption by bcv-commons and downstream tools.
The language: list above tracks the published partitions; the authoritative list is… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/lexeme-alignments.common_voice_17_0
Common Voice Corpus 17.0
Mirror for mozilla-foundation/common_voice_17_0, easy to download and extract instead audio in parquet files.
How to prepare the dataset
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \
malaysia-ai/common_voice_17_0
wget https://gist.githubusercontent.com/huseinzol05/2e26de4f3b29d99e993b349864ab6c10/raw/9b2251f3ff958770215d70c8d82d311f82791b78/unzip.py
python3 unzip.py
senses-attested
senses_attested — attested target renderings per lexeme sense
The empirical evidence layer produced for shoresh (bcv-query data-contract): for a lexeme in a
disambiguated (binyan, sense), which target-language words attest it, with counts. It is the supply
that fills shoresh's senses_i18n/_gaps demand and cross-checks the llm_strongs_glosses predictions —
it does not replace shoresh's curated senses_i18n/<iso>.tsv; consumed as an HF Parquet dataset.
Schema (per row)… See the full description on the dataset page: https://huggingface.co/datasets/bcv-commons/senses-attested.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.wikimedia-commons-maps_beirThis is a copy of https://huggingface.co/datasets/jinaai/wikimedia-commons-maps reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/wikimedia-commons-maps_beir.youtube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.dead-web-commoncrawl
Dead-Web Common Crawl — a longitudinal host-reachability panel (2018–2026)
· Hugging Face
· Kaggle
· License: CC BY 4.0
An open dataset labeling the reachability of 172,959,928 registered domains across 80 monthly
Common Crawl archives (CC-MAIN-2018-05 → CC-MAIN-2026-21), built from Common Crawl's robotstxt
subset and calibrated against a live re-probe to separate genuinely-dead domains from those merely
blocking crawlers or dropped from the… See the full description on the dataset page: https://huggingface.co/datasets/crawlora-net/dead-web-commoncrawl.common_starcoder
Common Starcoder dataset
This dataset is generated from bigcode/starcoderdata.
Total GPT2 Tokens: 4,649,163,171
Generation Process
We filtered the original dataset with common language: C, Cpp, Java, Python and JSON.
We removed some columns for mixing up with other dataset: "id", "max_stars_repo_path", "max_stars_repo_name"
After removing the irrelevant fields, we shuffle the dataset with random seed=42.
We filtered the data on "max_stars_count" > 300 and shuffle again.… See the full description on the dataset page: https://huggingface.co/datasets/skymizer/common_starcoder.common_voice_26_0_de
Mozilla Common Voice 26.0 - German (IPA & Clean Validated Subset)
Repacking version of Common Voice 26.0 German officialy published by Mozilla Data Collective, following Hugging Face Parquet Shards standard, with feature for listening to audio directly on the Web Hub, and the addition of a data column for the IPA transcription of each sentence.
📊 Dataset parameters
Origin: Mozilla Common Voice 26.0 (version 18/06/2026).
Data amount (Validated): 950,877 MP3 audio… See the full description on the dataset page: https://huggingface.co/datasets/q1805/common_voice_26_0_de.commoncrawl-feb-2025
