datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.IRFL
Dataset Card for IRFL
Dataset Description
Leaderboards
Colab notebook code for IRFL evaluation
Languages
Dataset Structure
Data Fields
Dataset Creation
Considerations for Using the Data
Licensing Information
Citation Information
Dataset Description
The IRFL dataset consists of idioms, similes, metaphors with matching figurative and literal images, and two novel tasks of multimodal figurative detection and retrieval.Using human annotation and an automatic pipeline… See the full description on the dataset page: https://huggingface.co/datasets/lampent/IRFL.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.nasa-smd-IR-benchmark
NASA-IR benchmark
NASA SMD and IBM Research developed a domain-specific information retrieval benchmark, NASA-IR, spanning almost 500 question-answer pairs related to the Earth science, planetary science, heliophysics, astrophysics, and biological physical sciences domains. Specifically, we sampled a set of 166 paragraphs from AGU, AMS, ADS, PMC, and PubMed and manually annotated with 3 questions that are answerable from each of these paragraphs, resulting in 498 questions. We used… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-IR-benchmark.IRPAPERS
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20
1… See the full description on the dataset page: https://huggingface.co/datasets/weaviate/IRPAPERS.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
IronyTRHomepage: https://github.com/teghub/IronyTR
Labels:
0: non-ironic
1: ironic
ELDOR-sample
ELDOR Sample
This repository is a compact sample companion to the full ELDOR dataset:
Full dataset: https://huggingface.co/datasets/IRSC/ELDOR
Paper: https://arxiv.org/abs/2605.15397
It is intended for quick inspection of the imagery, labels, and released patch extraction without downloading the full release.
Citation
@article{cui2026eldor,
title={ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest},
author={Cui, Kangning and Bohara… See the full description on the dataset page: https://huggingface.co/datasets/IRSC/ELDOR-sample.recipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.scraped-hotel-reviewsiris-tabularscikit_irisirs-articlesopencorporates_IRS990_EIN_matched_organizationsExternal Dataset Credits: OpenCorporates Limited
Dataset containing all IRS 990 organizations by EIN for 2022-2023 and their matched opencorporates unique identifiers (company_number). This allows researchers to combine IRS990 data with opencorporates state-level company data. If an organization 0 for the match confidence, that EIN was tested and did not match any opencorporates entities. All 990-CN, EZ, and PF organizations for 2022-2023 were tested.
customersatisfactionIrtNet-Dataset
Dataset for Learning Compact Representations of LLM Abilities via Item Response Theory
Paper Link:https://arxiv.org/abs/2510.00844
Code Link:https://github.com/JianhaoChen-nju/IrtNet
Dataset Details
Our data is mostly the same as EmbedLLM's. Specially, we applied a majority vote to consolidate multiple answers from a model to the same query.
If the number of 0s and 1s is the same, we prioritize 1.
This step ensures a unique ground truth for each model-query pair… See the full description on the dataset page: https://huggingface.co/datasets/JianhaoNJU/IrtNet-Dataset.iranian-surname-frequencies
Persian Last Names Dataset
Overview
Welcome to the Persian Last Names Dataset, a comprehensive collection of over 100,000 Persian surnames accompanied by their respective frequencies. This dataset is curated from a substantial real-world sample of more than 10 million records, ensuring reliable and representative data for various applications.
Dataset Details
Total Surnames: 100,000+
Frequency Source: Derived from a dataset comprising 10 million entries… See the full description on the dataset page: https://huggingface.co/datasets/farbodbij/iranian-surname-frequencies.reference-csvIRIS_sts
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences pairs from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for Semantic Textual Similarity
Values from 0-1: random sentences across documents
Values from 2-4: sentences from the same summary (implying some level of entailment)
Values from 4-5: sentences pairs generated through OpenAi'… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/IRIS_sts.africa-synth-agriculture-irrigation-access-efficiency-africa-all
Africa Synth Agriculture Irrigation Access Efficiency Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-irrigation-access-efficiency-africa-all.Agriculture-Irrigation-QA-Pairs-Dataset
Agriculture-Irrigation-QA-Pairs-Dataset
Released freely — support more experiments like it:
More: https://huggingface.co/YuvrajSingh9886
irhomeworkadamvakar_irish-rent-prices-2020-2025-rtb-official-data
Irish Rent Prices 2020-2025 (RTB Official Data)
Average monthly rent across 26 Irish counties - ML ready dataset
Dataset Info
Source: Kaggle
Original Size: 0.92 MB
Kaggle Downloads: 687
Files: 3
Files
irish_rent_by_county.csv
irish_rent_full.csv
irish_rent_specific.csv
Mirrored from Kaggle
iris-datasetAdvBench-IR-Small-Wiki-100Accel-IR
Accel-IR Benchmark: A Gold Standard for Particle Accelerator Physics
This repository contains the Accel-IR Benchmark, a domain-specific Information Retrieval (IR) dataset for particle accelerator physics. It was developed as part of the Master's Thesis "From Dataset to Optimization: A Benchmarking Framework for Information Retrieval in the Particle Accelerator Domain" by Qing Dai (University of Zurich, 2025), in collaboration with the Paul Scherrer Institute (PSI).… See the full description on the dataset page: https://huggingface.co/datasets/qdai/Accel-IR.irismetahate
MetaHate: A Dataset for Unifying Efforts on Hate Speech Detection
This is MetaHate: a meta-collection of 36 hate speech datasets from social media comments.
What's New in Version 2.0
Data Refinement and Size Adjustment: The original publication reported 1,226,202 total instances and 1,101,165 public instances. Due to refined cross-dataset deduplication and overlapping instance resolution, the current dataset sizes are 1,226,203 (full) and 1,084,236 (public).… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/metahate.Irish-English-Parallel-Collection
UCCIX's English-Irish Parallel Textual Corpus
Dataset Summary
This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR.
This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data.
Dataset Sources
Source
Description
Statistics
Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.
