CoolFace
Datasetpublicgated

JoeyLLM/new-zealand-dataset-5b

๐Ÿ‡ณ๐Ÿ‡ฟ New Zealand Web Text โ€” 5B-token Sample ๐ŸŒฟ A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. ๐ŸŒ The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. ๐Ÿ”’ Researchers seeking access to the full corpusโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes5downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.