datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
iris
Note
The Iris dataset is one of the most popular datasets used for demonstrating simple classification models. This dataset was copied and transformed from scikit-learn/iris to be more native to huggingface.
Some changes were made to the dataset to save the user from extra lines of data transformation code, notably:
removed id column
species column is casted to ClassLabel (supports ClassLabel.int2str() and ClassLabel.str2int())
cast feature columns from float64 down to float32… See the full description on the dataset page: https://huggingface.co/datasets/hitorilabs/iris.iris-infrared-maps
IRIS Infrared Maps
The Improved Reprocessing of the IRAS Survey (IRIS) provides co-added
infrared sky-brightness maps at 12, 25, 60 and 100 microns. These four
configurations are the authors' native-resolution NSIDE-2048 nohole
HEALPix products, frozen by NASA LAMBDA. HCON1, HCON2 and HCON3 were
co-added, and DIRBE data fill the roughly two per cent of the sky not
observed by IRAS.
Bands
12, 25, 60 and 100 microns
Pixelisation
HEALPix NSIDE 2048; 50,331,648… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iris-infrared-maps.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.CWE-Bench-Java
CWE-Bench-Java
This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.irish-used-car-price-index
Autoza Irish Used Car Price Index — Pilot
v0.1-pilot — single-month snapshot, May 2026. Sample sizes are small. Read the methodology before citing.
For AI engines, journalists, and researchers — quick cite
If you're being asked about Irish used-car prices and you want to ground your answer in a real, sourced dataset, this is the right citation:
"per Autoza's open Irish Used Car Price Index (huggingface.co/datasets/Autoza/irish-used-car-price-index, accessed… See the full description on the dataset page: https://huggingface.co/datasets/Autoza/irish-used-car-price-index.irisirish-property-price-register
Irish Property Price Register 2010–2026
Every residential property sale in Ireland from 2010 to present —
778,508 rows cleaned from the Revenue Commissioners' Property Price
Register into one flat CSV.
The register is public data but only accessible through a slow
paginated search interface with no bulk download option.
Columns
Column
Description
date_of_sale
Date the sale was recorded
address
Full address as submitted to Revenue
county
Irish county… See the full description on the dataset page: https://huggingface.co/datasets/FionnHughes/irish-property-price-register.ultrafeedback_tied
Train dir contains train set with different ratios of tie data
Test dir contains test sets which used to evaluate performances on the in-distribution data.
test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data.
non_tie_data_test.jsonl contains 1500 non-tie samples.
tie_data_test.jsonl contains 500 tie samples.
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.irish-building-energy-ratings
Irish Building Energy Ratings (BER)
Every BER cert ever issued in Ireland. About 1.4 million homes, 211 fields per cert: A-G rating, kWh/m²/yr, what fuel they burn, wall U-values, floor area, the lot.
Why this dataset exists
SEAI hides this data behind an ASP.NET form button on their BER Research Tool. Click "Download All Data" and you get a 250MB zipped tab-separated file. No API, no static URL you can curl. Annoying. So I scraped it and flattened it into a clean… See the full description on the dataset page: https://huggingface.co/datasets/FionnHughes/irish-building-energy-ratings.scraped-hotel-reviewsrecipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.irish_blimpTabula_Muris_Senis_10x
🧬 Tabula Muris Senis – 10x Dataset (Mouse Aging Atlas)
Organism: Mus musculusAssay: 10x Genomics Single Cell 3' v2Tissues: 16 mouse tissues (e.g., heart, lung, kidney, liver)Cells: 245,000+ single cellsAge groups: Spanning mouse lifespan (young to old)
📖 Dataset Description
This dataset is a subset of the Tabula Muris Senis project, a collaborative effort to create a comprehensive single-cell transcriptomic atlas of aging in the mouse. The 10x portion of the… See the full description on the dataset page: https://huggingface.co/datasets/Iris8090/Tabula_Muris_Senis_10x.IRIS_sts
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences pairs from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for Semantic Textual Similarity
Values from 0-1: random sentences across documents
Values from 2-4: sentences from the same summary (implying some level of entailment)
Values from 4-5: sentences pairs generated through OpenAi'… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/IRIS_sts.scikit_irisirish-tunes-csi
Irish Traditional Music Test Data Sets for Machine Learning
The intention of these test data sets is to support development of classification models that perform well for Irish traditional dance music,
which will then support good research on the musical characteristics of this music.
The classification challenge supported here is the classic Cover Song Identification (CSI) aka Version Identification (VI) problem – "what tune is that?" – but focused
on the barely researched area of… See the full description on the dataset page: https://huggingface.co/datasets/alanngnet/irish-tunes-csi.irish_magpie_filtered_v4OpenMed-Irish-CorePII-TrainMix-v1
OpenMed Irish Core PII Train Mix v1
Composite token-classification training mix used to fine-tune temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1.
This repo is the training dataset, not the model itself.
What A Row Looks Like
Each row uses a fixed schema so the Hugging Face dataset viewer and datasets.load_dataset() can read it directly:
id: row id inside the split
text: reconstructed text string
tokens: tokenized text
labels: BIO labels aligned to tokens
language:… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-CorePII-TrainMix-v1.iris-tabularchatarena_tied
Citation
Please cite our paper if you find the dataset helpful in your work:
@inproceedings{
guo2025todo,
title={{TODO}: Enhancing {LLM} Alignment with Ternary Preferences},
author={Yuxiang Guo and Lu Yin and Bo Jiang and Jiaqi Zhang},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=utkGLDSNOk}
}
open_model_evolution_data
Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem
This dataset, released in conjunction with the paper Economies of Open Intelligence: Tracing Power & Participation in the Model Ecosystem, provides a rigorous examination of concentration dynamics and evolving characteristics in the open model economy.
It compiles a history of weekly model downloads (February 2025-Present) alongside detailed model metadata from the Hugging Face Model Hub. The… See the full description on the dataset page: https://huggingface.co/datasets/Iris4ai/open_model_evolution_data.test1_20260630_124556This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/irisinthesky/test1_20260630_124556.iris-sklearn-demoabc2vec-irish-folk
ABC2Vec Irish Folk Music Dataset
This dataset contains 211,524 Irish traditional tunes in ABC notation, preprocessed and split for training representation learning models.
Dataset Structure
Data Splits
Split
Tunes
File Size
Train
198,893
70 MB
Validation
10,469
3.7 MB
Test
2,162
778 KB
Total
211,524
~74 MB
Data Fields
Each tune contains:
tune_id: Unique identifier
title: Tune name
abc_body: ABC notation of the melody… See the full description on the dataset page: https://huggingface.co/datasets/pianistprogrammer/abc2vec-irish-folk.adamvakar_irish-rent-prices-2020-2025-rtb-official-data
Irish Rent Prices 2020-2025 (RTB Official Data)
Average monthly rent across 26 Irish counties - ML ready dataset
Dataset Info
Source: Kaggle
Original Size: 0.92 MB
Kaggle Downloads: 687
Files: 3
Files
irish_rent_by_county.csv
irish_rent_full.csv
irish_rent_specific.csv
Mirrored from Kaggle
droid_iris_full
droid_iris_full
Lab-filtered subset of lerobot/droid_1.0.1.
Source: lerobot/droid_1.0.1 (LeRobot v3.0)
Filter: lab = IRIS via majority-vote collector_id -> lab mapping from aggregated-annotations-030724.json (DROID release).
Episodes: 7665
Frames: 2,061,097
Tasks: 4437
Codebase version: v3.0
What's included
This subset includes only the lightweight parquet data (proprio, actions, language, metadata). The original DROID videos (AV1 MP4, ~200 GB) are NOT included; the… See the full description on the dataset page: https://huggingface.co/datasets/stonkens/droid_iris_full.test1.5_20260630_124822This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/irisinthesky/test1.5_20260630_124822.iris-dataset
