CoolFace
Datasetpublic

vuhaian/fineweb_10k

fineweb_10k 10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and saved as-is (no shuffling needed beyond the source stream's own order). Companion single-source dataset to vuhaian/24_collected (23,000 examples across 23 sources) — used as the single-source comparison baseline in continue-pretraining experiments. Files fineweb_10k.jsonl -- raw text, one JSON object per line: {"text": ..., "source": "fineweb", "repo_id":… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/fineweb_10k.

sourceHugging Faceodc-byupdated 3mo agoView on Hugging Face
0likes5downloads
Dataset Card

fineweb_10k

10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and saved as-is (no shuffling needed beyond the source stream's own order). Companion single-source dataset to `vuhaian/24_collected` (23,000 examples across 23 sources) — used as the single-source comparison baseline in continue-pretraining experiments.

Files

  • —fineweb_10k.jsonl -- raw text, one JSON object per line: {"text": ..., "source": "fineweb", "repo_id": "HuggingFaceFW/fineweb"}