JoeyLLM/uk-dataset-1b
π¬π§ UK Web Text β 1B-token Sample β A 1-billion-token representative sample of a much larger cleaned United Kingdom web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. π The full 735B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. π Researchers seeking access to the full corpus forβ¦ See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-1b.
026
Update README.md
Update README.md
Create README.md
Upload uk_1B_sample.parquet (#1)
initial commit
