JoeyLLM/australian-dataset-5b
🇦🇺 Australian Web Text — 5B-token Sample 🦘 A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. 🌐 The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face