joeyllm
Datasets
All datasets matching “joeyllm”australian-dataset-1b
🇦🇺 Australian Web Text — 1B-token Sample 🦘
A 1-billion-token representative sample of a much larger cleaned Australian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 294B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-1b.uk-dataset-1b
🇬🇧 UK Web Text — 1B-token Sample ☕
A 1-billion-token representative sample of a much larger cleaned United Kingdom web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 735B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-1b.new-zealand-dataset-1b
🇳🇿 New Zealand Web Text — 1B-token Sample 🌿
A 1-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-1b.canada-dataset-1b
🇨🇦 Canada Web Text — 1B-token Sample 🍁
A 1-billion-token representative sample of a much larger cleaned Canadian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 210B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-1b.canada-dataset-5b
🇨🇦 Canada Web Text — 5B-token Sample 🍁
A 5-billion-token sample of cleaned Canada-attributed web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐
This dataset is intended as a large-scale Canadian English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research.
📊 Dataset Summary 📌
Property… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-5b.new-zealand-dataset-5b
🇳🇿 New Zealand Web Text — 5B-token Sample 🌿
A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.
