datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
australian-housing-market-data
Australian Housing Market Landscape
GitHub: Alex-Dunstan
This dataset contains 2,490 Australian Statistical Areas Level 2 (SA2s), using 2026 ASGS geography, 2021 Census profile features, and selected 2024 area-level housing-price measures.
Data background
Each row describes one Australian Statistical Area Level 2 (SA2): a local-area geography used by the Australian Bureau of Statistics. The data combines 2026 ASGS boundaries and coordinates with 2021 Census… See the full description on the dataset page: https://huggingface.co/datasets/alexdunstan/australian-housing-market-data.australian-insurance-pii-dataset-correctedaustralian-dataset-1b
🇦🇺 Australian Web Text — 1B-token Sample 🦘
A 1-billion-token representative sample of a much larger cleaned Australian web-text corpus derived from Common Crawl. This sample is released alongside the JoeyLLM project's ongoing research into regional English language models. 🌐
The full 294B-token corpus is not publicly released due to its size, operational cost, and intended controlled use in research and model development. 🔒 Researchers seeking access to the full corpus for… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-1b.australian-sme-business-dataset
SME Business Dataset (Free Samples — AU, US, UK)
Realistic, relational business datasets generated by simulating retail SMEs day-by-day over 2 financial years. Every transaction flows through double-entry accounting.
Browse all datasets and variants →
Three variants available: Australian (ATO/GST), US (IRS/FICA), and UK (HMRC/PAYE/VAT).
What makes this different
Feature
AdventureWorks
Northwind
Faker/Mockaroo
This dataset
Cross-domain traceability
Partial… See the full description on the dataset page: https://huggingface.co/datasets/mindweave/australian-sme-business-dataset.australian-mbs-pbs-data
Australian MBS & PBS Health Data (2026)
Complete datasets for Australia's Medicare Benefits Schedule (MBS) and Pharmaceutical Benefits Scheme (PBS) with plain English descriptions.
Dataset Description
MBS Items — mbs-items.csv (6,005 records)
Medicare item numbers with schedule fees, rebates, and plain English explanations for Australian patients.
Column
Description
item_number
MBS item number
category
Category code
category_name… See the full description on the dataset page: https://huggingface.co/datasets/bimicroads/australian-mbs-pbs-data.australian-insurance-pii-datasetAustralian_English_Recognition_Speech_Library_Conversations
ID
King-ASR-363
Language
English
Duration
100 hours
Speakers
100 People
Parameters
44.1kHz, 16bits
Recording Device
Mobile
URL
https://dataoceanai.com/datasets/asr/australian-english-recognition-speech-library-conversations-mobile/
corto-ai-open-australian-legal-multi-lingual-qaCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données https://huggingface.co/datasets/corto-ai/open-australian-legal-multi-lingual-qa.
australian-dataset-5b
🇦🇺 Australian Web Text — 5B-token Sample 🦘
A 5-billion-token Australian web-text dataset created for the JoeyLLM project. This dataset was sampled from the filtered Australian corpus produced by the JoeyLLM sovereign corpus pipeline. 🌐
The purpose of this dataset is to provide a large-scale Australian text corpus for GPT-style language-model pre-training, continued pre-training, data inspection, and research into regional English language models.
📊 Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/JoeyLLM/australian-dataset-5b.Australian_Landmarks_Dataset
