CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01commoncrawl /host-index-testing-v2 Common Crawl Host Index v2 GitHub: https://github.com/commoncrawl/cc-host-index Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The information is aggregated from the Common Crawl columnar index, web graph, and raw crawler logs. Quickstart The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.tabulartext-generation1B<n<10B0 likes7.4k downloads11d agoHugging Face02common-pile /stackv2_edu_filtered Stack V2 Edu Description We filter the Stack V2 to only include code from openly licensed repositories, based on the license detection performed by the creators of Stack V2. When multiple licenses are detected in a single repository, we ensure that all of the licenses are on the Blue Oak Council certified license list. Per-document license information is available in the license entry of the metadata field of each example. Code for collecting, processing, and preparing… See the full description on the dataset page: https://huggingface.co/datasets/common-pile/stackv2_edu_filtered.tabulartext-generation10M<n<100M6 likes6.9k downloads1y agoHugging Face03coral-nlp /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/coral-nlp/german-commons.tabulartext-generation10M<n<100M41 likes3.9k downloads8mo agoHugging Face04CUI03 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokensof German text data with… See the full description on the dataset page: https://huggingface.co/datasets/CUI03/german-commons.tabulartext-generation10M<n<100M1 likes2.6k downloads9mo agoHugging Face05trace-commons /agent-traces Trace Commons — Agent Traces Trace Commons is one open, public dataset of coding-agent sessions — the back-and-forth between a developer and an AI coding agent, including prompts, model responses, tool calls, and command output — contributed voluntarily as an open resource for studying, evaluating, and building on how these agents actually work. Every trace here was donated only from a public, open-source repository, was anonymized on the contributor's own machine before upload… See the full description on the dataset page: https://huggingface.co/datasets/trace-commons/agent-traces.tabulartext-generationn<1K35 likes1.9k downloads3mo agoHugging Face06Rijgersberg /common_corpus_nl Common Corpus v2 NL This is a version of Common Corpus v2 filtered to keep only the rows where language is "Dutch". Common Corpus is a very large open and permissible licensed text dataset created by Pleias. Please be sure to acknowledge the creators of the original dataset when using this filtered version. Filtering Common Corpus is a collection of disparate datasets. Note that filtering the entire collection for rows where the language is "Dutch" is not the same as… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_nl.tabulartext-generation1M<n<10M4 likes1.4k downloads2y agoHugging Face07Rijgersberg /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/YouTube-Commons.tabulartext-generation10M<n<100M6 likes413 downloads2y agoHugging Face08dm-petrov /youtube-commons-small 📺 YouTube-Commons-Small 📺 This is a smaller subset of the YouTube-Commons dataset, which is a collection of audio transcripts from videos shared on YouTube under a CC-By license. Dataset Description This smaller version contains a subset of the original dataset, maintaining the same structure and features. It's designed for easier experimentation and testing purposes. Features The dataset includes the following information for each video: Video ID and link… See the full description on the dataset page: https://huggingface.co/datasets/dm-petrov/youtube-commons-small.tabulartext-generation100K<n<1M1 likes275 downloads1y agoHugging Face09amine-khelif /YouTube-Commons YouTube Commons Re-upload This is a re-upload of PleIAs' YouTube Commons, a valuable open dataset: YouTube-Commons is a collection of audio transcripts of 2,063,066 videos shared on YouTube under a CC BY 4.0 license. Content The collection comprises 22,709,724 original and automatically translated transcripts from 3,156,703 videos (721,136 individual channels). Unfortunately, there are problems with loading YouTube Commons with Hugging Face Datasets. In order to alleviate those… See the full description on the dataset page: https://huggingface.co/datasets/amine-khelif/YouTube-Commons.tabulartext-generation10M<n<100M0 likes157 downloads10mo agoHugging Face10ThingAI /Italian-Common-Corpus Italian-Common-Corpus The Italian dataset with the highest density of useful information per token. Built by ModotAI for training Italian language models. Subsets Subset File Documents Words Description Web Crawl icc-web.parquet ~27K ~17M Italian sources: news, tech, science, culture, law, food, sport Wikipedia IT wiki-it-clean.parquet ~1.35M ~698M Cleaned Italian Wikipedia — removed Notes, Bibliography, Voci correlate, stub articles Total: 1,377… See the full description on the dataset page: https://huggingface.co/datasets/ThingAI/Italian-Common-Corpus.tabulartext-generation1M<n<10M1 likes71 downloads2mo agoHugging Face11Rijgersberg /common_corpus_dutch_pd Common Corpus v2 - Dutch Public Domain collection This is a version of Common Corpus v2 filtered to keep only the rows where collection is "Dutch-PD". Looking for all Dutch-language documents in Common Corpus, regardless of the collection they are in? Then you might want to look at Rijgersberg/common_corpus_nl. Common Corpus is a very large open and permissible licensed text dataset created by Pleias. Please be sure to acknowledge the creators of the original dataset when using this… See the full description on the dataset page: https://huggingface.co/datasets/Rijgersberg/common_corpus_dutch_pd.tabulartext-generation100K<n<1M0 likes64 downloads1y agoHugging Face12Tinuade /common-crawl-docx-sample Common Crawl DOCX Sample A sample of normalized text extracted from DOCX records in Common Crawl. Source Common Crawl release: CC-MAIN-YYYY-NN Source index: Common Crawl URL Index Pipeline: marin-community/marin Pipeline revision: REPLACE_WITH_GIT_SHA Records were selected using declared DOCX MIME type, detected DOCX MIME type, or a .docx URL suffix. Only successful, non-truncated index records were eligible. Processing The pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/Tinuade/common-crawl-docx-sample.tabulartext-generation1K<n<10K0 likes48 downloads10d agoHugging Face13alex73 /mozilla-common-voice-23-bel-texts-exporttabulartext-generation100K<n<1M0 likes41 downloads10mo agoHugging Face14Lottikarotti92 /german-commons German Commons - 154 Billion Tokens of Openly Licensed Text for German Language Models A comprehensive collection of German-language text data under open licenses for training German language models. Datasheet: DATASHEET.md. Paper: arxiv.org/abs/2510.13996 Code: github.com/coral-nlp/llmdata Bloom Filter (DOLMA-compatible): bloom_filter.bin Dataset Description This dataset is aggregated from 41 diverse sources and contains 154.56 billion tokens of German text… See the full description on the dataset page: https://huggingface.co/datasets/Lottikarotti92/german-commons.tabulartext-generation10M<n<100M0 likes35 downloads1d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.