CoolFace
Datasetpublicgated

JoeyLLM/australian-dataset-5b

🇦🇺 Australian Web Text — 5B-token Sample 🦘 A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. 🌐 The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes3downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
JoeyLLM/australian-dataset-5b · CoolFace