datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cad-corpus-cleanedharmonicmlx-cleaned-corpus
HarmonicMLX Cleaned Corpus v3
High-quality, balanced English text corpus for small language model pre-training.
Properly rebalanced to avoid TinyStories domination.
Pipeline
Source ingestion: FineWeb-Edu (623 MB), TinyStories (1.8 GB), Stanford Encyclopedia of Philosophy (127 MB), Project Gutenberg
Cleaning: Unicode normalization, Gutenberg/archive header stripping, URL removal, whitespace collapse
Chunking: Sentence-aware chunking (128-2048 chars)
Exact deduplication:… See the full description on the dataset page: https://huggingface.co/datasets/MonumentalSystems/harmonicmlx-cleaned-corpus.appropriateness-corpus-extension-cleanedappropriateness-corpus-cleanedsmall_corpus_cleanedscientific_corpus_cleaned
Dataset Card for Scientific Corpus (Cleaned)
This corpus contains ≈11 M English scientific documents cleaned via the DataTrove pipeline. It was used to continue pretraining T5-base (EN‑T5-Sci) before sliding-window materialization. Each document is provided as a row in one of 75 Parquet shards together with extensive per-document QA metadata.
Dataset Details
Uses
Direct Use
Continued pretraining / domain adaptation of encoder-decoder LMs on… See the full description on the dataset page: https://huggingface.co/datasets/rausch/scientific_corpus_cleaned.
