ray0rf1re/Fineweb-Tiny
Fineweb-Tiny Dataset Description Fineweb-Tiny is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb. How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: High language_score (if provided by the upstream extraction). Optimal document length (penalizing abnormally short snippets and… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Fineweb-Tiny.
Fineweb-Tiny
Dataset Description
Fineweb-Tiny is a highly curated, premium subset extracted from `nampdn-ai/mini-fineweb`.
How "The Best" Was Determined
This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors:
- High
language_score(if provided by the upstream extraction). - Optimal document length (penalizing abnormally short snippets and excessively long, unformatted dumps).
- Strong structural coherence suitable for pre-training Small Language Models (SLMs) and LLMs.
Fineweb-Tiny contains the absolute BEST 72.9 Gigabytes (compressed Parquet) of the original source. It drops the bottom 50% of lower-quality data from the stream to ensure dense, high-utility text.
License
This dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 to perfectly match the upstream Fineweb mini licensing constraints.
