CoolFace
Datasetpublic

vuhaian/fineweb_10k

fineweb_10k 10,000 text samples from HuggingFaceFW/fineweb (train split), streamed and saved as-is (no shuffling needed beyond the source stream's own order). Companion single-source dataset to vuhaian/24_collected (23,000 examples across 23 sources) — used as the single-source comparison baseline in continue-pretraining experiments. Files fineweb_10k.jsonl -- raw text, one JSON object per line: {"text": ..., "source": "fineweb", "repo_id":… See the full description on the dataset page: https://huggingface.co/datasets/vuhaian/fineweb_10k.

sourceHugging Faceodc-byupdated 3mo agoView on Hugging Face
0likes5downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
vuhaian/fineweb_10k · CoolFace