datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.irish-legislative-summaries
Irish Legislative Summaries ⚖️
Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation.
This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.ultrafeedback_tied
Train dir contains train set with different ratios of tie data
Test dir contains test sets which used to evaluate performances on the in-distribution data.
test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data.
non_tie_data_test.jsonl contains 1500 non-tie samples.
tie_data_test.jsonl contains 500 tie samples.
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.iris-mixiris_02oneills-irish-tunes-1850
O'Neill's Irish Tunes (1850)
1,849 traditional Irish tunes in ABC notation, transcribed from Captain
Francis O'Neill's O'Neill's Music of Ireland: 1850 Melodies (Chicago,
1903). Each row is one tune: title, type, key, meter, and the full ABC
source.
Dataset structure
Field
Type
Description
tune_id
string
Stable id, e.g. oneills1850-1 (tune number in the original book)
name
string
Tune title
tune_type
string
Rhythm/category: reel, jig, slip jig… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/oneills-irish-tunes-1850.Iris-dataOpenMed-Irish-CorePII-TrainMix-v1
OpenMed Irish Core PII Train Mix v1
Composite token-classification training mix used to fine-tune temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1.
This repo is the training dataset, not the model itself.
What A Row Looks Like
Each row uses a fixed schema so the Hugging Face dataset viewer and datasets.load_dataset() can read it directly:
id: row id inside the split
text: reconstructed text string
tokens: tokenized text
labels: BIO labels aligned to tokens
language:… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-CorePII-TrainMix-v1.IrishQA
IrishQA
chatarena_tied
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}: Enhancing {LLM} Alignment with Ternary Preferences},
author={Yuxiang Guo and Lu Yin and Bo Jiang and Jiaqi Zhang},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=utkGLDSNOk}
}
scanner-poisoned-iris-benchmark
Scanner Poisoned Iris Benchmark
This benchmark starts from the classic UCI Iris dataset and injects multiple synthetic poisoning patterns so dataset scanners can exercise duplicate, anomaly, missingness, skew, and divergence heuristics against a small tabular corpus.
Recommended Hugging Face repo slug: your-org/scanner-poisoned-iris-benchmark
What It Is For
benchmarking dataset quality and poisoning detection workflows
regression-testing scanner heuristics on a… See the full description on the dataset page: https://huggingface.co/datasets/jgracie52/scanner-poisoned-iris-benchmark.iris_v3OpenMed-Irish-PPSN-Eircode-Spec-v1
OpenMed Irish PPSN Eircode Spec v1
Focused synthetic token-classification dataset for Irish PPSN and Eircode detection.
This repo contains synthetic training rows, not a fine-tuned model.
What A Row Looks Like
Each row uses a fixed schema:
id: row id inside the split
text: rendered text string
tokens: tokenized text
labels: BIO labels aligned to tokens
language: en or ga
source_dataset: generator identifier
source_domain: optional domain tag, empty in this release… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1.iris-agent-trace
Iris: Coding-Agent Build Trace 👁️
A record of how Iris,
a voice-first assistant for blind and low-vision people, was built with Claude Code
(Claude Opus 4.8) during the Build Small Hackathon.
It follows the build session step by step: the decisions, the tool calls, the
debugging, and what each step taught. Iris was made for the author's father, who is
blind, so the trace runs from the first idea through to a working app on a phone.
Earns the Sharing is Caring (open trace)… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/iris-agent-trace.iris_eng_combo_dpoirish_retrieval_datadpo_irish_eng_translationsThis is a test for my DPO dataset for Irish ENglish trasnlslations, raw data origin : https://www.gaois.ie/en/corpora/parallel?Query=Apple&Language=en&SearchMode=exact&PerPage=50, used COMETXL refrernce free maodel Unbabel/wmt23-cometkiwi-da-xl
(which has been trained to asses Irish) to score accepted/rejected. Used GPT4 to generate translations to compare with human stranslations of Irish legislation (which has to have a Irisng/English copy by law)
GNU-IRIS
Dataset Card for GNU-IRIS
GNU-IRIS is a training dataset for GIMPLE IR to LLVM IR translation, derived from GNU utilities source code. It contains 13,049 C functions paired with their corresponding GIMPLE and LLVM intermediate representations.
Dataset Structure
The default dataset combines GNU utils projects into a single training dataset. However, you can also access each GNU utility individually as a split:
Configuration
Package
# Samples
default
All… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/GNU-IRIS.irish-passports
Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes.
Introduction - Ireland
The Synthetic Ireland Passports Dataset gathers more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Every record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/irish-passports.iris_0vfin_irisDPO_IRIS_Datairish_belebeleIrish version of https://huggingface.co/datasets/facebook/belebele.
Translated using facebook/nllb-200-3.3B, and the translations are verified by native Irish speakers.
iris-l2irish-english-dialectsi-rag-recursive-testsi-rag-recursive-originalssi-rag-flat-to-recursivesi-process-meta-test
