CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hula0401 /cad-corpus-cleanedtabular1M<n<10M4 likes13k downloads3mo agoHugging Face02MonumentalSystems /harmonicmlx-cleaned-corpus HarmonicMLX Cleaned Corpus v3 High-quality, balanced English text corpus for small language model pre-training. Properly rebalanced to avoid TinyStories domination. Pipeline Source ingestion: FineWeb-Edu (623 MB), TinyStories (1.8 GB), Stanford Encyclopedia of Philosophy (127 MB), Project Gutenberg Cleaning: Unicode normalization, Gutenberg/archive header stripping, URL removal, whitespace collapse Chunking: Sentence-aware chunking (128-2048 chars) Exact deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/harmonicmlx-cleaned-corpus.tabular100K<n<1M0 likes45 downloads5mo agoHugging Face03timonziegenbein /appropriateness-corpus-extension-cleanedtabular10K<n<100K0 likes10 downloads1y agoHugging Face04timonziegenbein /appropriateness-corpus-cleanedtabularn<1K0 likes4 downloads1y agoHugging Face05ashercn97 /small_corpus_cleanedtabular10K<n<100K0 likes3 downloads2y agoHugging Face06rausch /scientific_corpus_cleanedgated Dataset Card for Scientific Corpus (Cleaned) This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata. Dataset Details Uses Direct Use Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.tabulartext-generation1K<n<10K2 likes2 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.