CoolFace
Datasetpublic

ray0rf1re/Fineweb-Tiny

Fineweb-Tiny Dataset Description Fineweb-Tiny is a highly curated, premium subset extracted from nampdn-ai/mini-fineweb. How "The Best" Was Determined This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors: High language_score (if provided by the upstream extraction). Optimal document length (penalizing abnormally short snippets and… See the full description on the dataset page: https://huggingface.co/datasets/ray0rf1re/Fineweb-Tiny.

sourceHugging Faceodc-byupdated 6mo agoView on Hugging Face
0likes107downloads
Dataset Card

Fineweb-Tiny

Dataset Description

Fineweb-Tiny is a highly curated, premium subset extracted from `nampdn-ai/mini-fineweb`.

How "The Best" Was Determined

This dataset was created programmatically by streaming the original dataset and sorting chunks based on a rigorous quality scoring algorithm. The heuristic heavily favors:

  1. 1.High language_score (if provided by the upstream extraction).
  2. 2.Optimal document length (penalizing abnormally short snippets and excessively long, unformatted dumps).
  3. 3.Strong structural coherence suitable for pre-training Small Language Models (SLMs) and LLMs.

Fineweb-Tiny contains the absolute BEST 72.9 Gigabytes (compressed Parquet) of the original source. It drops the bottom 50% of lower-quality data from the stream to ensure dense, high-utility text.

License

This dataset is released under the Open Data Commons Attribution License (ODC-By) v1.0 to perfectly match the upstream Fineweb mini licensing constraints.