datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.stackv2_edu_filtered
Stack V2 Edu
Description
We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.german-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.agent-traces
Trace Commons — Agent Traces
Trace Commons is one open, public dataset of coding-agent sessions — the
back-and-forth between a developer and an AI coding agent, including prompts,
model responses, tool calls, and command output — contributed voluntarily as an
open resource for studying, evaluating, and building on how these agents
actually work.
Every trace here was donated only from a public, open-source repository, was
anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.common_corpus_nl
Common Corpus v2 NL
This is a version of Common Corpus v2 filtered to keep only the rows where language is "Dutch".
Common Corpus is a very large open and permissible licensed text dataset created by Pleias.
Please be sure to acknowledge the creators of the original dataset when using this filtered version.
Filtering
Common Corpus is a collection of disparate datasets.
Note that filtering the entire collection for rows where the language is "Dutch" is not the same as… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_nl.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.youtube-commons-small
📺 YouTube-Commons-Small 📺
This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license.
Dataset Description
This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes.
Features
The dataset includes the following information for each video:
Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.YouTube-Commons
YouTube Commons Re-upload
This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset:
YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license.
Content
The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels).
Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets.
In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.Italian-Common-Corpus
Italian-Common-Corpus
The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models.
Subsets
Subset
File
Documents
Words
Description
Web Crawl
icc-web.parquet
~27K
~17M
Italian sources: news, tech, science, culture, law, food, sport
Wikipedia IT
wiki-it-clean.parquet
~1.35M
~698M
Cleaned Italian Wikipedia — removed Notes, Bibliography, Voci correlate, stub articles
Total: 1,377… See the full description on the dataset page: https://huggingface.co/datasets/ThingAI/Italian-Common-Corpus.common_corpus_dutch_pd
Common Corpus v2 - Dutch Public Domain collection
This is a version of Common Corpus v2 filtered to keep only the rows where collection is "Dutch-PD".
Looking for all Dutch-language documents in Common Corpus, regardless of the collection they are in?
Then you might want to look at Rijgersberg/common_corpus_nl.
Common Corpus is a very large open and permissible licensed text dataset created by Pleias.
Please be sure to acknowledge the creators of the original dataset when using this… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_dutch_pd.common-crawl-docx-sample
Common Crawl DOCX Sample
A sample of normalized text extracted from DOCX records in Common Crawl.
Source
Common Crawl release: CC-MAIN-YYYY-NN
Source index: Common Crawl URL Index
Pipeline: marin-community/marin
Pipeline revision: REPLACE_WITH_GIT_SHA
Records were selected using declared DOCX MIME type, detected DOCX MIME type,
or a .docx URL suffix. Only successful, non-truncated index records were
eligible.
Processing
The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.mozilla-common-voice-23-bel-texts-exportgerman-commons
German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models
A comprehensive collection of German-language text data under open licenses for training German language models.
Datasheet: DATASHEET.md.
Paper: arxiv.org/abs/2510.13996
Code: github.com/coral-nlp/llmdata
Bloom Filter (DOLMA-compatible): bloom_filter.bin
Dataset Description
This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text… See the full description on the dataset page: https://huggingface.co/datasets/Lottikarotti92/german-commons.
