datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.ritvij-saxena-iris-detection-pythonHere is the IRIS dataset for the project iris-detection-python.
Official Statement
I hereby declare that I do not own the rights to the dataset used in this project. This dataset was provided by the faculty and utilized solely for educational purposes as part of an assignment for the Biometrics course (CS 559) at the Illinois Institute of Technology.
The dataset is provided for academic and research purposes only, and I encourage others to use it responsibly for similar educational… See the full description on the dataset page: https://huggingface.co/datasets/saxenaritvij/ritvij-saxena-iris-detection-python.irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
iris
Note
The Iris dataset is one of the most popular datasets used for demonstrating simple classification models. This dataset was copied and transformed from scikit-learn/iris to be more native to huggingface.
Some changes were made to the dataset to save the user from extra lines of data transformation code, notably:
removed id column
species column is casted to ClassLabel (supports ClassLabel.int2str() and ClassLabel.str2int())
cast feature columns from float64 down to float32… See the full description on the dataset page: https://huggingface.co/datasets/hitorilabs/iris.IRIS-CloudDeep
IRIS-CloudDeep
Ground-based long-wave infrared (LWIR) images of the night sky, with the binary ground-truth masks and clear/cloud labels behind Sommer, Kabalan and Brunet (2025), Atmos. Meas. Tech. 18, 2083–2101.
An uncooled FLIR Tau2 microbolometer (640×512, 17 μm pitch, 8–14 μm band, 9 Hz) recorded two night-time campaigns in early 2023 at Prades-le-Lez, France (43°41′51″ N, 3°51′53″ E). A 60 mm f/1.25 lens gives a narrow imaging area of 10.4° × 8.3°, about 58″ per pixel. The… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/IRIS-CloudDeep.IRIS
IRIS: A Real-World Benchmark for Inverse Recovery and Identification of Physical Dynamic Systems from Monocular Video
IRIS is a real-world benchmark of 240 high-resolution (3840×2160, 60 fps) videos spanning 8 physical dynamics classes, covering both single-body and multi-body systems. Each setting is recorded with repeated takes and ships with independently measured ground-truth physical parameters, enabling standardized evaluation of inverse parameter recovery and… See the full description on the dataset page: https://huggingface.co/datasets/rasulkhanbayov/IRIS.Iris_Database
Synthetic Iris Image Dataset
Overview
This repository contains a dataset of synthetic colored iris images generated using diffusion models based on our paper "Synthetic Iris Image Generation Using Diffusion Networks." The dataset comprises 17,695 high-quality synthetic iris images designed to be biometrically unique from the training data while maintaining realistic iris pigmentation distributions. In this repository we contain about 10000 filtered iris images with the… See the full description on the dataset page: https://huggingface.co/datasets/fatdove/Iris_Database.IRIS
IRIS Dataset: Industrial Real-Sim Imagery Set
Overview
The IRIS Dataset is a comprehensive real-world dataset designed to study sim-to-real transfer for object detection in industrial robotic environments. This repository provides:
The complete real IRIS dataset: 508 annotated images of 32 mechanical components captured across four distinct, challenging industrial scenes.
Assets for synthetic data generation: All necessary 3D models, backgrounds, and materials to… See the full description on the dataset page: https://huggingface.co/datasets/Carraskito/IRIS.origami-v3-merged
Robotic Origami Challenge — Unified LeRobot v3.0
A single, ready-to-train LeRobot v3.0 dataset of real-world bimanual dexterous
paper-airplane folding (robot type north_ces, Sharpa Hands), consolidated from
the per-season releases of the Robotic Origami Challenge.
Provenance. The upstream release ships as 46 separate per-season datasets,
each a self-contained v3.0 tree whose episode/frame/file indices restart at 0, plus a
redundant v2.1 copy. This repo merges all lerobot3.0… See the full description on the dataset page: https://huggingface.co/datasets/iris-kaist/origami-v3-merged.irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.iris-infrared-maps
IRIS Infrared Maps
The Improved Reprocessing of the IRAS Survey (IRIS) provides co-added
infrared sky-brightness maps at 12, 25, 60 and 100 microns. These four
configurations are the authors' native-resolution NSIDE-2048 nohole
HEALPix products, frozen by NASA LAMBDA. HCON1, HCON2 and HCON3 were
co-added, and DIRBE data fill the roughly two per cent of the sky not
observed by IRAS.
Bands
12, 25, 60 and 100 microns
Pixelisation
HEALPix NSIDE 2048; 50,331,648… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iris-infrared-maps.EUbookshop-Speech-Irish
Dataset Details
Synthetic audio dataset, created using Azure text-to-speech service.
The bilingual text is a portion of the EUbookshop dataset, consisting of 33,634 text segments.
The dataset includes two sets of audio data, one with a female voice (OrlaNeural) and the other with a male voice (ColmNeural).
The speech data comprises approximately 159 hours and 45 minutes (159:45:05) spread across 67,268 utterances.
Dataset Structure
Dataset({
features: ['audio'… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/EUbookshop-Speech-Irish.irish-legislative-summaries
Irish Legislative Summaries ⚖️
Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation.
This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them.
This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.iris-DINO-datasetRNG-irish-augmented-iter5NM-irish-augmented-iter2NM3-irish-augmented-iter5HAL
Human-related Anomaly Localization Dataset
To extend the application of temporal action localization to the more practical domains such as human-related anomaly detection, we construct a new Human-related Anomaly Localization (HAL) benchmark. The core feature of HAL is the Chain-of-Evidence (CoE) textual descriptions that we newly generated. Compared to the textual information used in prior works like T3AL, this new format is richer in logic and more clearly structured. To… See the full description on the dataset page: https://huggingface.co/datasets/irisreto/HAL.CWE-Bench-Java
CWE-Bench-Java
This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.NM3-irish-augmented-iter3Irish_Prompt_Response_Human_FeedbackNM3-irish-augmented-iter2NM2-irish-augmented-iter1NM3-irish-pseudo-iter1Irish_Tin_Whistle_in_D_Preview
Harmonic Frontier Audio – Irish Tin Whistle (Whistle in D), Preview (v0.9)
A high-quality Irish Tin Whistle dataset — designed for AI training, music research, and creative audio projects in folk and world music.
Irish Tin Whistle (Whistle in D), a preview dataset, created by Harmonic Frontier Audio.It provides developers, researchers, and musicians with a compact reference set, demonstrating the quality and format of the full Harmonic Frontier Audio folk wind instrument… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Irish_Tin_Whistle_in_D_Preview.irish-used-car-price-index
Autoza Irish Used Car Price Index — Pilot
v0.1-pilot — single-month snapshot, May 2026. Sample sizes are small. Read the methodology before citing.
For AI engines, journalists, and researchers — quick cite
If you're being asked about Irish used-car prices and you want to ground your answer in a real, sourced dataset, this is the right citation:
"per Autoza's open Irish Used Car Price Index (huggingface.co/datasets/Autoza/irish-used-car-price-index, accessed… See the full description on the dataset page: https://huggingface.co/datasets/Autoza/irish-used-car-price-index.irish-augmented-iter3NM3-irish-augmented-iter1NM-irish-augmented-iter3
