CoolFace
20 results

corpus

PleIAs /common_corpus Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising 2.27 trillion tokens (2,267,302,720,836 tokens). It is a diverse dataset, consisting of books, newspapers, scientific articles, government and legal documents, code, and more. Common Corpus has been created by Pleias in association with several partners. Common Corpus differs from existing open datasets in that it is: Truly Open: contains only data that is either… See the full description on the dataset page: https://huggingface.co/datasets/PleIAs/common_corpus.tabular10K<n<100K423 likes189k downloads5mo agoHugging Facenvidia /Nemotron-Terminal-Corpus Terminal-Corpus: Large-Scale SFT Dataset for Terminal Agents Terminal-Corpus is a large-scale Supervised Fine-Tuning (SFT) dataset designed to scale the terminal interaction capabilities of Large Language Models (LLMs). Developed by NVIDIA, this dataset was built using the Terminal-Task-Gen pipeline, which combines dataset adaptation with synthetic task generation across diverse domains. 🚀 Key Results & Performance The high-quality trajectories in Terminal-Corpus enable… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Terminal-Corpus.textquestion-answering100K<n<1M144 likes97k downloads7mo agoHugging FaceSZLHOLDINGS /killinchu-osint-corpus killinchu — live intel archive Append-only, content-addressed archive of the live intelligence streams the killinchu demo ingests, published by SZL Holdings. The shared szl_hf_bucket client writes one NDJSON shard per UTC day under intel/. Each row is {schema, id, ts, source, kind, payload}. Viewer-safe contract The default Dataset Viewer configuration is the homogeneous archive_manifest, one row per immutable raw shard. It exposes path, byte count, Git blob hash… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/killinchu-osint-corpus.15 likes59k downloads4m agoHugging FaceHuggingFaceTB /smollm-corpus SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language models. You can find more details about the models trained on this dataset in our SmolLM blog post. Dataset subsets Cosmopedia v2 Cosmopedia v2 is an enhanced version of Cosmopedia, the largest synthetic dataset for pre-training, consisting of over 39 million textbooks, blog posts, and stories generated by… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.tabular100M<n<1B485 likes55k downloads2y agoHugging Faceruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes45k downloads5d agoHugging Facejinofy-corp /jora_corpus1_tokenized_128ktabularn<1K4 likes44k downloads2mo agoHugging Face