CoolFace
6 results

joeyllm

JoeyLLM /australian-dataset-1b 🇦🇺 Australian Web Text — 1B-token Sample 🦘 A 1-billion-token representative sample of a much larger cleaned Australian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 294B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-1b.tabulartext-generation1M<n<10M0 likes31 downloads5mo agoHugging FaceJoeyLLM /uk-dataset-1b 🇬🇧 UK Web Text — 1B-token Sample ☕ A 1-billion-token representative sample of a much larger cleaned United Kingdom web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 735B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/uk-dataset-1b.tabulartext-generation1M<n<10M0 likes26 downloads5mo agoHugging FaceJoeyLLM /new-zealand-dataset-1b 🇳🇿 New Zealand Web Text — 1B-token Sample 🌿 A 1-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-1b.tabulartext-generation1M<n<10M0 likes14 downloads5mo agoHugging FaceJoeyLLM /canada-dataset-1b 🇨🇦 Canada Web Text — 1B-token Sample 🍁 A 1-billion-token representative sample of a much larger cleaned Canadian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 210B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-1b.tabulartext-generation1M<n<10M0 likes9 downloads5mo agoHugging FaceJoeyLLM /canada-dataset-5bgated 🇨🇦 Canada Web Text — 5B-token Sample 🍁 A 5-billion-token sample of cleaned Canada-attributed web text derived from Common Crawl. This sample was produced as part of the JoeyLLM project's ongoing research into regional English language datasets. 🌐 This dataset is intended as a large-scale Canadian English web-text corpus for language-model pre-training, continued pre-training, data inspection, and regional English research. 📊 Dataset Summary 📌 Property… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/canada-dataset-5b.texttext-generation1M<n<10M0 likes8 downloads4mo agoHugging FaceJoeyLLM /new-zealand-dataset-5bgated 🇳🇿 New Zealand Web Text — 5B-token Sample 🌿 A 5-billion-token representative sample of a much larger cleaned New Zealand web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐 The full 60.7B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/new-zealand-dataset-5b.tabulartext-generation1M<n<10M0 likes4 downloads4mo agoHugging Face