CoolFace
Datasetpublic

CerebellumKing/Ultra-FineWeb-10B

Ultra-FineWeb-10B Archive This dataset is an archived copy of sumukshashidhar-archive/Ultra-FineWeb-10B. The original dataset link became unavailable, so I am releasing this locally stored copy to make the data accessible again for researchers and practitioners who may need it for reference, reproduction, or small-scale pretraining experiments. Important Note This 10B-token subset was created by random sampling. It should not be interpreted as the highest-quality… See the full description on the dataset page: https://huggingface.co/datasets/CerebellumKing/Ultra-FineWeb-10B.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes94downloads
Dataset Card

Ultra-FineWeb-10B Archive

This dataset is an archived copy of sumukshashidhar-archive/Ultra-FineWeb-10B.

The original dataset link became unavailable, so I am releasing this locally stored copy to make the data accessible again for researchers and practitioners who may need it for reference, reproduction, or small-scale pretraining experiments.

Important Note

This 10B-token subset was created by random sampling. It should not be interpreted as the highest-quality 10B-token subset of Ultra-FineWeb.

In other words, this archive is intended to preserve access to the previously available data, rather than to provide a newly curated or quality-optimized selection.

Intended Use

This dataset may be useful for:

  • —language model pretraining experiments
  • —data pipeline testing
  • —reproducibility checks
  • —comparisons involving Ultra-FineWeb-derived data

Users who need the highest-quality subset should consider applying their own filtering, scoring, or selection pipeline.

Dataset Source

Ultra-FineWeb is a large-scale web corpus designed for high-quality language model pretraining. Please refer to the original Ultra-FineWeb project for details about the data construction and filtering process.

Limitations

  • —This is an archive release.
  • —The 10B-token subset is randomly sampled.
  • —No additional quality ranking, deduplication, or filtering was applied by this archive release.
  • —The dataset may inherit noise, duplication, or other artifacts from the original source.

Acknowledgements

All credit for the original Ultra-FineWeb dataset goes to the Ultra-FineWeb authors and contributors. This repository only provides an archived copy to preserve access after the original link became unavailable.