corpus
thainer-corpus-v2-base-modelmerlyn-education-corpus-qa-v2-GGUFMegatron-Corpus-14B-Exp.v2-i1-GGUFMegatron-Corpus-14B-Exp-i1-GGUFt5-base-long-livedoor-news-corpusMegatron-Corpus-14B-Exp-GGUFbert-base-arabic-camelbert-mix-did-madar-corpus26AMindToThink-gemma-2-2b_RMU_cyber-forget-corpus_s100_a500_layer3-GGUF
Datasets
All datasets matching “corpus”common_corpus
Common Corpus
Full paper - ICLR 2026 oral
Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners.
Common Corpus differs from existing open datasets in that it is:
Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.Nemotron-Terminal-Corpus
Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents
Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains.
🚀 Key Results & Performance
The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Corpus.killinchu-osint-corpus
killinchu — live intel archive
Append-only, content-addressed archive of the live intelligence streams the
killinchu demo ingests, published by SZL Holdings. The shared
szl_hf_bucket client writes one NDJSON shard per UTC day under
intel/.
Each row is {schema, id, ts, source, kind, payload}.
Viewer-safe contract
The default Dataset Viewer configuration is the homogeneous
archive_manifest, one row per immutable raw shard. It exposes path, byte
count, Git blob hash… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/killinchu-osint-corpus.smollm-corpus
SmolLM-Corpus
This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models.
You can find more details about the models trained on this dataset in our SmolLM blog post.
Dataset subsets
Cosmopedia v2
Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.infini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.jora_corpus1_tokenized_128k
