datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.nasa-smd-IR-benchmark
NASA-IR benchmark
NASA SMD and IBM Research developed a domain-specific information retrieval benchmark, NASA-IR, spanning almost 500 question-answer pairs related to the Earth science, planetary science, heliophysics, astrophysics, and biological physical sciences domains. Specifically, we sampled a set of 166 paragraphs from AGU, AMS, ADS, PMC, and PubMed and manually annotated with 3 questions that are answerable from each of these paragraphs, resulting in 498 questions. We used… See the full description on the dataset page: https://huggingface.co/datasets/nasa-impact/nasa-smd-IR-benchmark.IRPAPERS
Dataset Card for IRPAPERS
ArXiv Link: https://arxiv.org/pdf/2602.17687
Dataset Description
IRPAPERS is a collection of 166 Information Retrieval papers spanning 3,230 pages. Each page in the dataset is jointly represented as a base64 encoded string of the page image as well as an OCR-derived text transcription. IRPAPERS also contains 180 needle-in-the-haystack queries.
Retrieval Leaderboard 🔎
Rank
Retriever
Type
Recall@1
Recall@5
Recall@20
1… See the full description on the dataset page: https://huggingface.co/datasets/weaviate/IRPAPERS.reddit_mental_health_posts
Reddit posts about mental health
files
adhd.csv from r/adhd
aspergers.csv from r/aspergers
depression.csv from r/depression
ocd.csv from r/ocd
ptsd.csv from r/ptsd
fields
author
body
created_utc
id
num_comments
score
subreddit
title
upvote_ratio
url
for more details about theses fields Praw Submission.
scraped-hotel-reviewsELDOR-sample
ELDOR Sample
This repository is a compact sample companion to the full ELDOR dataset:
Full dataset: https://huggingface.co/datasets/IRSC/ELDOR
Paper: https://arxiv.org/abs/2605.15397
It is intended for quick inspection of the imagery, labels, and released patch extraction without downloading the full release.
Citation
@article{cui2026eldor,
title={ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest},
author={Cui, Kangning and Bohara… See the full description on the dataset page: https://huggingface.co/datasets/IRSC/ELDOR-sample.recipe-cleaned
Recipe Cleaned Dataset
Dataset Summary
This dataset is a structured and cleaned collection of recipe data derived from the Food.com Recipes and Interactions dataset. It is designed for ingredient-based personalization, machine learning training, and interactive recommendation systems. The dataset integrates a hierarchical ingredient taxonomy, standardized nutrition information, and categorical metadata (e.g., diet tags, cuisine attributes, region) to support downstream… See the full description on the dataset page: https://huggingface.co/datasets/Iris314/recipe-cleaned.iris-tabularscikit_irisopencorporates_IRS990_EIN_matched_organizationsExternal Dataset Credits: OpenCorporates Limited
Dataset containing all IRS 990 organizations by EIN for 2022-2023 and their matched opencorporates unique identifiers (company_number). This allows researchers to combine IRS990 data with opencorporates state-level company data. If an organization 0 for the match confidence, that EIN was tested and did not match any opencorporates entities. All 990-CN, EZ, and PF organizations for 2022-2023 were tested.
IrtNet-Dataset
Dataset for Learning Compact Representations of LLM Abilities via Item Response Theory
Paper Link:https://arxiv.org/abs/2510.00844
Code Link:https://github.com/JianhaoChen-nju/IrtNet
Dataset Details
Our data is mostly the same as EmbedLLM's. Specially, we applied a majority vote to consolidate multiple answers from a model to the same query.
If the number of 0s and 1s is the same, we prioritize 1.
This step ensures a unique ground truth for each model-query pair… See the full description on the dataset page: https://huggingface.co/datasets/JianhaoNJU/IrtNet-Dataset.customersatisfactionreference-csvIRIS_sts
Work developed as part of Project IRIS.
Thesis: A Semantic Search System for Supremo Tribunal de Justiça
Portuguese Legal Sentences
Collection of Legal Sentences pairs from the Portuguese Supreme Court of Justice
The goal of this dataset was to be used for Semantic Textual Similarity
Values from 0-1: random sentences across documents
Values from 2-4: sentences from the same summary (implying some level of entailment)
Values from 4-5: sentences pairs generated through OpenAi'… See the full description on the dataset page: https://huggingface.co/datasets/stjiris/IRIS_sts.africa-synth-agriculture-irrigation-access-efficiency-africa-all
Africa Synth Agriculture Irrigation Access Efficiency Africa All | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-agriculture-irrigation-access-efficiency-africa-all.Accel-IR
Accel-IR Benchmark: A Gold Standard for Particle Accelerator Physics
This repository contains the Accel-IR Benchmark, a domain-specific Information Retrieval (IR) dataset for particle accelerator physics. It was developed as part of the Master's Thesis "From Dataset to Optimization: A Benchmarking Framework for Information Retrieval in the Particle Accelerator Domain" by Qing Dai (University of Zurich, 2025), in collaboration with the Paul Scherrer Institute (PSI).… See the full description on the dataset page: https://huggingface.co/datasets/qdai/Accel-IR.adamvakar_irish-rent-prices-2020-2025-rtb-official-data
Irish Rent Prices 2020-2025 (RTB Official Data)
Average monthly rent across 26 Irish counties - ML ready dataset
Dataset Info
Source: Kaggle
Original Size: 0.92 MB
Kaggle Downloads: 687
Files: 3
Files
irish_rent_by_county.csv
irish_rent_full.csv
irish_rent_specific.csv
Mirrored from Kaggle
iris-datasetirisiranian-elderly-psychospiritual-interviews
Iranian Elderly Psycho-Spiritual Interviews
A culturally grounded, fully synthetic conversational interview dataset for assessing the mental and spiritual health of Iranian older adults, generated using Large Language Models.
Dataset Summary
This dataset introduces a culturally grounded, fully synthetic conversational interview corpus designed for the assessment and analysis of mental and spiritual health among Iranian older adults. All interviews are conducted in… See the full description on the dataset page: https://huggingface.co/datasets/liamirali/iranian-elderly-psychospiritual-interviews.Iranian-Household-DataThis dataset contains anonymized socio-economic data directly from the Iranian Welfare Embassy. It was collected and prepared for tasks related to socio-economic modeling and preference learning.
The main prediction task is to determine the 'percentile' column, which represents the socio-economic percentile of a household.
This dataset was first introduced in the following paper:
Cold Start Active Preference Learning in Socio-Economic Domains
Mojtaba Fayaz-Bakhsh, Danial Ataee, MohammadAmin… See the full description on the dataset page: https://huggingface.co/datasets/Dan-A2/Iranian-Household-Data.financeArtificial-intelligence-dataset-for-IR-systems
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
information-retrieval
semantic-search
Languages
English
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Adel-Elwan/Artificial-intelligence-dataset-for-IR-systems.iris
Iris Species Dataset
The Iris dataset is a classic dataset in machine learning, originally published by Ronald Fisher. It contains 150 instances of iris flowers, each described by four features (sepal length, sepal width, petal length, and petal width), along with the corresponding species label (setosa, versicolor, or virginica).
It is commonly used as an introductory dataset for classification tasks and for demonstrating basic data exploration and model training workflows.… See the full description on the dataset page: https://huggingface.co/datasets/brjapon/iris.ircam-dglai-dataset
Dataset Card for DGLAi-Augmented
An augmented, quad-lingual electronic parallel dataset based on the Dictionnaire Général de la Langue Amazighe (DGLAi). The original standard source fields (Amazigh, Arabic, French) are curated by IRCAM, supplemented with automated English translations to maximize its utility for modern machine translation, cross-lingual NLP, and LLM applications.
Dataset Details
Dataset Description
The original Dictionnaire Général… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/ircam-dglai-dataset.iononosphereiran_hs_customs_classification
طبقهبندیِ کالا (HS) — تعرفهٔ گمرک ایران
جدولِ طبقهبندیِ نظامِ هماهنگشدهٔ توصیف و کدگذاری کالا (HS) برای تعرفهٔ گمرک ایران، با
سلسلهمراتبِ کامل: فصل (۲ رقم) ← عنوان (۴) ← زیرعنوان (۶، بینالمللی) ← تعرفه (۸ رقم، ملی).
ستون
توضیح
code
کدِ HS
level
فصل / عنوان / زیرعنوان / تعرفه
digits
تعداد رقم (۲/۴/۶/۸)
name_fa
شرحِ فارسیِ کالا
parent_code
کدِ والد (برای تجمیع)
۱۴٬۷۳۷ کد شامل ۸٬۴۰۰ ردیفِ تعرفهٔ هشترقمی. مرجعی برای پیوستن به دادههای تجارت و تحلیلِ… See the full description on the dataset page: https://huggingface.co/datasets/Farmaanaa/iran_hs_customs_classification.ir-to-adm
