datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
messirveJuly 2025 UPDATE: We released version 1.1, adding almost 200k new queries 🎉🎉🎉.
v1.2 further adds the article titles as columns for convenience.
Use with:
country = "full" # "ar", "bo", ...
version = "1.2"
dataset = datasets.load_dataset("spanish-ir/messirve", country, revision=version)
print(dataset)
Dataset Card for MessIRve
MessIRve is a large-scale dataset for Spanish IR, designed to better capture the information needs of Spanish speakers across different countries.… See the full description on the dataset page: https://huggingface.co/datasets/spanish-ir/messirve.prospect-ptms-irt
PROSPECT PTMs - Retention Time Prediction
A mass-spectrometry dataset for applied machine learning in proteomics, processed and split for the task of retention time prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo.
Repository: https://github.com/wilhelm-lab/PROSPECT
Uses
The… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-irt.RFSD
The Russian Financial Statements Database (RFSD)
The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms:
🔓 First open data set with information on every active firm in Russia.
🗂️ First open financial statements data set that includes non-filing firms.
🏛️ Sourced from two official data providers: the Rosstat and the Federal Tax Service.
📅 Covers 2011-2025, will be… See the full description on the dataset page: https://huggingface.co/datasets/irlspbru/RFSD.RPX
RPX: Robot Perception X
RPX is a real-world RGB-D benchmark for measuring robot perception across scene changes. The canonical naren/all release combines the multi-object, egocentric, single-object, VQA, and tracking metadata that previously lived on separate dataset branches.
Code and benchmark toolkit: github.com/IRVLUTD/RPX
Recommended dataset revision: naren/all (pin the commit SHA printed by your download for reproducible results)
License: Creative Commons Attribution 4.0… See the full description on the dataset page: https://huggingface.co/datasets/IRVLUTD/RPX.iris
Note
The Iris dataset is one of the most popular datasets used for demonstrating simple classification models. This dataset was copied and transformed from scikit-learn/iris to be more native to huggingface.
Some changes were made to the dataset to save the user from extra lines of data transformation code, notably:
removed id column
species column is casted to ClassLabel (supports ClassLabel.int2str() and ClassLabel.str2int())
cast feature columns from float64 down to float32… See the full description on the dataset page: https://huggingface.co/datasets/hitorilabs/iris.irs-990-parsed
IRS 990 Parsed Nonprofit Database
Public relational extract of IRS Form 990 / 990-EZ / 990-PF filings, plus the colocated public files we join for address research: CMS NPPES + T-MSIS Medicare spend, FMCSA DOT carriers, OFAC SDN, FEC committees, and the IRS EO BMF.
Generated: 2026-08-17Tables: 34Rows (sum): 459,069,505License: CC0 / public domain — derived from U.S. government recordsHub: https://huggingface.co/datasets/piercewetter3/irs-990-parsed
Layout
Tables… See the full description on the dataset page: https://huggingface.co/datasets/piercewetter3/irs-990-parsed.ipfs_germany_laws_ir
Germany legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_germany_laws (revision 62477e216917df268f186972ef49581b02144156) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Germany prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_germany_laws_ir.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.ipfs_switzerland_laws_ir
Switzerland legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_switzerland_laws (revision 25cd3f9f14d6e6ee5453bb5ab36a9c2d6e8f8c1c) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Switzerland prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_switzerland_laws_ir.struct-ir
SSRB: Direct Natural Language Querying to Massive Heterogeneous Semi-Structured Data
github
We employ LLM-based automatic evaluation and build a large-scale semi-structured retrieval benchmark (SSRB) using LLM generation and filtering, containing 14M structured objects from 99 different schemas across 6 domains, along with 8,485 test queries that combine both exact and fuzzy matching conditions.
This repository contains the data for SSRB.
Data Download
Data can be… See the full description on the dataset page: https://huggingface.co/datasets/vec-ai/struct-ir.natori-irodori-tts-dataset
natori-irodori-tts-dataset
High-quality single-speaker Japanese speech dataset prepared for Irodori-TTS LoRA training from multiple long-form さなちゃんねる videos featuring 名取さな.
Dataset Summary
This repository is a merged export of four curated subsets derived from public YouTube playlists on さなちゃんねる.
Current top-level merged export:
40,995 total utterances
33,398 training examples
7,597 validation examples
43.4277 hours of speech
speaker id: natori_sana
The merged… See the full description on the dataset page: https://huggingface.co/datasets/argo11/natori-irodori-tts-dataset.wikipedia-multilingual-synthetic-ir-query
wikipedia-multilingual-synthetic-ir-query
This dataset contains multilingual Wikipedia-derived synthetic query-document pairs for information retrieval training.
It was created with the query-crafter-multilingual model, which generates search-like queries from Wikipedia text.
The current release contains two different retrieval settings:
short_doc: pairs of (query, short document)
long_doc: pairs of (query, long document)
These two subsets are not generated in the same way… See the full description on the dataset page: https://huggingface.co/datasets/hotchpotch/wikipedia-multilingual-synthetic-ir-query.IA-Bench
IA-bench ( Interaction-Aware Bench)
Human ground-truth annotations of the interacted object for robot manipulation
subtasks. Each sample is one subtask: the full subtask video clip, the gripper
proprioception aligned 1:1 to those frames, the language instruction, and two boxes:
initial_object_box (object on the first frame) and target_object_box
(object on the last frame). Boxes are pixel [x1, y1, x2, y2].
Configs
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/irl-kit/IA-Bench.iris-infrared-maps
IRIS Infrared Maps
The Improved Reprocessing of the IRAS Survey (IRIS) provides co-added
infrared sky-brightness maps at 12, 25, 60 and 100 microns. These four
configurations are the authors' native-resolution NSIDE-2048 nohole
HEALPix products, frozen by NASA LAMBDA. HCON1, HCON2 and HCON3 were
co-added, and DIRBE data fill the roughly two per cent of the sky not
observed by IRAS.
Bands
12, 25, 60 and 100 microns
Pixelisation
HEALPix NSIDE 2048; 50,331,648… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iris-infrared-maps.ipfs_azerbaijan_laws_ir
Azerbaijan legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_azerbaijan_laws (revision 01e4eb269e4de3aa301daaa5ca725c7de218b3b6) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Azerbaijan prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_azerbaijan_laws_ir.ipfs_sweden_laws_ir
Sweden legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_sweden_laws (revision 0a1e752cb178b83a779996f0e60e6f8d305a857e) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Sweden prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_sweden_laws_ir.patent-legal-ir-graphrag
Patent Legal IR GraphRAG
Single-dataset retrieval release at
justicedao/patent-legal-ir-graphrag.
Layout matches Publicus IR GraphRAG releases
(Publicus/cvefixes-security-ir-graphrag, Publicus/skillcenter-ir):
Family
Path
Notes
Corpus
data/corpus/*.parquet
dense document_index, CID keys
BM25 documents
data/bm25/documents/*.parquet
lengths + entry CID
BM25 postings
data/bm25/postings/*.parquet
sorted terms, FTS5 IDF, sparse lists
Vectors
data/vectors/*.parquet… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/patent-legal-ir-graphrag.Hand_Tools
ABC-iRobotics/Hand_Tools: dataset 0000
Metric RGB-D scene dataset generated with the Label Factory workflow.
The files for this capture are stored below the 0000/ directory so
multiple numbered datasets can coexist in this repository.
Training configurations
depth_estimation: RGB input, metric depth target, intrinsics and depth units
instance_segmentation: RGB input, instance-mask target, boxes and annotations
object_pose_estimation: RGB-D input, masks, camera… See the full description on the dataset page: https://huggingface.co/datasets/ABC-iRobotics/Hand_Tools.irish-census
Irish Census 1901 & 1926
Person-level records from the 1901 and 1926 censuses of Ireland, as published by
the National Archives of Ireland — every individual return, in flat CSV.
Year
Rows
Size
Coverage
1901
4,434,939
4.31 GB
All of Ireland (32 counties)
1926
2,973,480
0.56 GB
Saorstát Éireann (26 counties)
Total
7,408,419
4.87 GB
The 1926 census is the first taken by the Irish Free State and was released to
the public in 2026 under the 100-year rule. The… See the full description on the dataset page: https://huggingface.co/datasets/Cianmcnally/irish-census.naive-physics-ironing-v0.2
nAIve physics — Ironing Pilot v0.2 + Interaction Analysis v0.3
Visual Preview
Original RGB demonstration — IRON_009
▶ Watch IRON_009 original RGB demonstration
v0.3 interaction analysis — IRON_009
▶ Watch IRON_009 analysed interaction video
Raw → analysed: the first video is the original RGB demonstration; the second shows the v0.3 garment semantics, tool tracking, and temporal interaction analysis derived from the same episode.
A… See the full description on the dataset page: https://huggingface.co/datasets/CaramelCoffee19/naive-physics-ironing-v0.2.iras-sky-flux-plates
IRAS Sky Flux Plates
IRAS high-resolution hours-confirmed survey intensity and statistical-weight plates in the four survey bands.
Data structure
Each represented .INTE or .STAT archive member is one source-named configuration. Raw FITS axes and any whole-pixel truncation are retained.
The holding has 4,827 source-named configurations from 212 served archives containing 4,832 members.
Raw values, source order, FITS image axes and structured provenance are… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/iras-sky-flux-plates.LSD_mPro_Fink_2023_noncovalent_Screen2_ZINC22ipfs_bangladesh_laws_ir
Bangladesh legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_bangladesh_laws (revision 16782096c126f7342b3cfeaa312c437c9fa2de73) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Bangladesh prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_bangladesh_laws_ir.ipfs_angola_laws_ir
Angola legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_angola_laws (revision ebcd1a38594e8b382ce462cd2aa6c9e664980754) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Angola prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_angola_laws_ir.skillcenter-ir
SkillCenter Intent IR retrieval corpus
This is the CID-keyed retrieval release at
Publicus/skillcenter-ir. It
converts the complete
Tommysha/skillcenter-bundles
SkillCenter corpus and its local retrieval artifacts into
thin-client-friendly, Zstandard-compressed Parquet. It is bound to upstream
revision f9dd4fec3c86d85ebf116c7408ac5ce602c418a1 and contains:
216,972 canonical skills keyed by entry_cid;
3,776,520 BM25 terms and
107,971,682 document-term postings;
434,135 graph… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/skillcenter-ir.iran_tedpix_stock_index_daily
شاخص کل بورس تهران (TEDPIX) — روزانه
سریِ روزانهٔ شاخص کل بورس تهران (TEDPIX) بر پایهٔ دادههای رسمیِ سازمان بورس.
پوشش: 1387-09-14 → 1405-04-31 (4,249 روزِ معاملاتی) · تناوب: روزانه · سطح: ملی
داده
هر ردیف شاملِ مقدارِ پایانیِ شاخص در آن روز است؛ بیشترین و کمترینِ روز نیز در کنارِ آن ارائه میشود.
فایلِ bourse.parquet نمای کامل (تاریخِ میلادی و شمسی، پایانی/بیشترین/کمترین) را دارد.
منبع: سازمان بورس و اوراق بهادار تهران
روششناسی: farmaanaa.ir
ipfs_albania_laws_ir
Albania legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_albania_laws (revision f293e236b21786c58101a8c96c0ac5b71ef00960) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Albania prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_albania_laws_ir.ipfs_cyprus_laws_ir
Cyprus legislation IR (CID-keyed sparse GraphRAG)
Research retrieval release of endomorphosis/ipfs_cyprus_laws (revision c46312ea702229228f76c228a26f54058f3482fd) packaged as
country-laws-ir-graphrag/v1 (layout family skillcenter-huggingface-release/v3 / publicus-ir).
Not legal advice. This is a research snapshot. The official gazette /
authentic source of Cyprus prevails over this corpus. Retrieved documents
and graph edges are retrieval evidence only. No legal text was… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/ipfs_cyprus_laws_ir.
