JoeyLLM/new-zealand-dataset-5b
๐ณ๐ฟ New Zealand Web Text โ 5B-token Sample ๐ฟ A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. ๐ The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. ๐ Researchers seeking access to the full corpusโฆ See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.
05
No card is published for this repository, or it could not be fetched from Hugging Face right now.
