datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.irish-legislative-summaries
Irish Legislative Summaries ⚖️
Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation.
This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.EUbookshop-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the EUbookshop dataset, consisting of 33,634 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 159 hours and 45 minutes (159:45:05) spread across 67,268 utterances.
Dataset Structure
Dataset({
features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/EUbookshop-Speech-Irish.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.RNG-irish-augmented-iter5NM-irish-augmented-iter2NM3-irish-augmented-iter5NM2-irish-augmented-iter1NM3-irish-pseudo-iter1Irish_Prompt_Response_Human_FeedbackNM3-irish-augmented-iter3NM3-irish-augmented-iter2irish-used-car-price-index
Autoza Irish Used Car Price Index — Pilot
v0.1-pilot — single-month snapshot, May 2026. Sample sizes are small. Read the methodology before citing.
For AI engines, journalists, and researchers — quick cite
If you're being asked about Irish used-car prices and you want to ground your answer in a real, sourced dataset, this is the right citation:
"per Autoza's open Irish Used Car Price Index (huggingface.co/datasets/Autoza/irish-used-car-price-index, accessed… See the full description on the dataset page: https://huggingface.co/datasets/Autoza/irish-used-car-price-index.NM2-irish-pseudo-iter1irish-augmented-iter3NM-irish-augmented-iter3NM3-irish-pseudo-iter2RNG-irish-augmented-iter3NM-irish-augmented-iter1Irish_Tin_Whistle_in_D_Preview
Harmonic Frontier Audio – Irish Tin Whistle (Whistle in D), Preview (v0.9)
A high-quality Irish Tin Whistle dataset — designed for AI training, music research, and creative audio projects in folk and world music.
Irish Tin Whistle (Whistle in D), a preview dataset, created by Harmonic Frontier Audio.It provides developers, researchers, and musicians with a compact reference set, demonstrating the quality and format of the full Harmonic Frontier Audio folk wind instrument… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Irish_Tin_Whistle_in_D_Preview.NM3-irish-augmented-iter4NM-irish-pseudo-iter2RNG-irish-augmented-iter4NM3-irish-augmented-iter1irish-augmented-iter1irish-traditional-tunes
Dataset Card for "irish-traditional-tunes"
More Information needed
Dataset Card for "irish-tunes-spectrograms"
1. Dataset Description
Dataset is used for the following project
Homepage: Trad-fusion
1.1 Dataset Summary
This dataset contains 9604 Mel spectrograms that represent Traditional Irish Music.
This dataset is smaller compared to hdparmar/irish-tunes-spectrogram, to reduce the training time and increase the possibilty to train for longer… See the full description on the dataset page: https://huggingface.co/datasets/hdparmar/irish-traditional-tunes.irish-augmented-iter2irish-speech-datasetWikimedia-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the Wikimedia dataset, consisting of 7,545 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 34 hours and 23 minutes (34:23:12) spread across 15,090 utterances.
Dataset Structure
Dataset({
features: ['audio', 'text_ga'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/Wikimedia-Speech-Irish.
