JoeyLLM/uk-dataset-5b
π¬π§ UK Web Text β 5B-token Sample A 5-billion-token sample of cleaned United Kingdom web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. π This dataset is intended as a large-scale UK-attributed English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. π Dataset Summary π Property Thisβ¦ See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-5b.
03
No card is published for this repository, or it could not be fetched from Hugging Face right now.
