vuhaian/fineweb_10k
fineweb_10k 10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and saved as-is (no shuffling needed beyond the source stream's own order). Companion single-source dataset to vuhaian/24_collected (23,000 examples across 23 sources) — used as the single-source comparison baseline in continue-pretraining experiments. Files fineweb_10k.jsonl -- raw text, one JSON object per line: {"text": ..., "source": "fineweb", "repo_id":… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/fineweb_10k.
05
