CoolFace
Datasetpublicgated

JoeyLLM/uk-dataset-5b

πŸ‡¬πŸ‡§ UK Web Text β€” 5B-token Sample A 5-billion-token sample of cleaned United Kingdom web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐 This dataset is intended as a large-scale UK-attributed English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. πŸ“Š Dataset Summary πŸ“Œ Property This… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-5b.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes3downloads

No commit history came back for main. The revision may not exist, or the source declined the request.