vuhaian/fineweb_10k
fineweb_10k 10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and saved as-is (no shuffling needed beyond the source stream's own order). Companion single-source dataset to vuhaian/24_collected (23,000 examples across 23 sources) — used as the single-source comparison baseline in continue-pretraining experiments. Files fineweb_10k.jsonl -- raw text, one JSON object per line: {"text": ..., "source": "fineweb", "repo_id":… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/fineweb_10k.
This repository belongs to vuhaian on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
