CoolFace
Datasetpublicgated

JoeyLLM/australian-dataset-5b

๐Ÿ‡ฆ๐Ÿ‡บ Australian Web Text โ€” 5B-token Sample ๐Ÿฆ˜ A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. ๐ŸŒ The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models. ๐Ÿ“Š Datasetโ€ฆ See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes3downloads

No commit history came back for main. The revision may not exist, or the source declined the request.