JoeyLLM/uk-dataset-5b
🇬🇧 UK Web Text — 5B-token Sample A 5-billion-token sample of cleaned United Kingdom web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐 This dataset is intended as a large-scale UK-attributed English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. 📊 Dataset Summary 📌 Property This… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-5b.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
