datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.irish-legislative-summaries
Irish Legislative Summaries ⚖️
Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation.
This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.oneills-irish-tunes-1850
O'Neill's Irish Tunes (1850)
1,849 traditional Irish tunes in ABC notation, transcribed from Captain
Francis O'Neill's O'Neill's Music of Ireland: 1850 Melodies (Chicago,
1903). Each row is one tune: title, type, key, meter, and the full ABC
source.
Dataset structure
Field
Type
Description
tune_id
string
Stable id, e.g. oneills1850-1 (tune number in the original book)
name
string
Tune title
tune_type
string
Rhythm/category: reel, jig, slip jig… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/oneills-irish-tunes-1850.OpenMed-Irish-CorePII-TrainMix-v1
OpenMed Irish Core PII Train Mix v1
Composite token-classification training mix used to fine-tune temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1.
This repo is the training dataset, not the model itself.
What A Row Looks Like
Each row uses a fixed schema so the Hugging Face dataset viewer and datasets.load_dataset() can read it directly:
id: row id inside the split
text: reconstructed text string
tokens: tokenized text
labels: BIO labels aligned to tokens
language:… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-CorePII-TrainMix-v1.IrishQA
IrishQA
OpenMed-Irish-PPSN-Eircode-Spec-v1
OpenMed Irish PPSN Eircode Spec v1
Focused synthetic token-classification dataset for Irish PPSN and Eircode detection.
This repo contains synthetic training rows, not a fine-tuned model.
What A Row Looks Like
Each row uses a fixed schema:
id: row id inside the split
text: rendered text string
tokens: tokenized text
labels: BIO labels aligned to tokens
language: en or ga
source_dataset: generator identifier
source_domain: optional domain tag, empty in this release… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1.irish_retrieval_datairish-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Ireland
The Synthetic Ireland Passports Dataset gathers more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Every record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/irish-passports.dpo_irish_eng_translationsThis is a test for my DPO dataset for Irish ENglish trasnlslations, raw data origin : https://www.gaois.ie/en/corpora/parallel?Query=Apple&Language=en&SearchMode=exact&PerPage=50, used COMETXL refrernce free maodel Unbabel/wmt23-cometkiwi-da-xl
(which has been trained to asses Irish) to score accepted/rejected. Used GPT4 to generate translations to compare with human stranslations of Irish legislation (which has to have a Irisng/English copy by law)
irish-english-dialectIrish_English_Translation
Irish English graded Translations
Data collected in order to fine-tune a Irish-English LLM, see blog post for more details.
See here for the next phase, preference dataset formated (for DPO) here
Data Sources:
translated_gaois_graded.jsonl
Parallel English-Irish corpus of legislation collected by Gaois. This corpus contains high-quality, human-translated paragraph pairs, making it a valuable resource.
translated_tatoeba_graded.jsonl
One draw-back is uses a lot of… See the full description on the dataset page: https://huggingface.co/datasets/c123ian/Irish_English_Translation.irish_eng_dpoirish_belebeleIrish version of https://huggingface.co/datasets/facebook/belebele.
Translated using facebook/nllb-200-3.3B, and the translations are verified by native Irish speakers.
si-rag-recursive-testsi-rag-flat-to-recursivesi-rag-recursive-originalssi-process-meta-testirish-casino-payout-index
The Irish Casino Payout Index — timed withdrawal test log
A first-party log of timed real-money casino withdrawals in Ireland.
59 timed withdrawals · 41 operators · April–August 2026 · Methodology v1.0
Live dataset: https://greenfelt.ie/casino-payout-tracker/
Press kit and charts: https://greenfelt.ie/press/irish-casino-payout-index/
Publisher: Greenfelt.ie, Ireland
What this is
Each record is one test event. We deposited our own money at an online casino
serving… See the full description on the dataset page: https://huggingface.co/datasets/Greenfelt/irish-casino-payout-index.irish-datasetirish-citizen-information-fine-tuning-dataner-irish-dataset
NER IRISH Dataset
A manually annotated Indonesian news corpus for Named Entity Recognition (NER), created using Label Studio. This dataset provides raw, character-level annotations to ensure maximum flexibility across different tokenization strategies and preprocessing pipelines.
Dataset Summary
Language: Indonesian
Task: Named Entity Recognition (NER)
Annotation Type: Character-level spans (start/end offsets)
Format: Raw text with entity annotations
Total Samples: 1909… See the full description on the dataset page: https://huggingface.co/datasets/tlabdev/ner-irish-dataset.ner-irish-dataset-2
NER Irish Dataset 2
An Indonesian news corpus for Named Entity Recognition (NER) consisting of manually annotated (gold) and model-annotated (silver) articles. This dataset provides raw, character-level entity annotations to ensure maximum flexibility across different tokenization strategies and downstream preprocessing pipelines.
The dataset is designed with gold-only evaluation in mind, while leveraging silver data as additional training supervision.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/tlabdev/ner-irish-dataset-2.
