datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
iris
Iris Species Dataset
The Iris dataset was used in R.A. Fisher's classic 1936 paper, The Use of Multiple Measurements in Taxonomic Problems, and can also be found on the UCI Machine Learning Repository.
It includes three iris species with 50 samples each as well as some properties about each flower. One flower species is linearly separable from the other two, but the other two are not linearly separable from each other.
The dataset is taken from UCI Machine Learning Repository's… See the full description on the dataset page: https://huggingface.co/datasets/scikit-learn/iris.IROS-2025-Challenge-Manip
IROS-2025-Challenge-Manip
Dataset Summary 📖
This dataset contains the IROS Challenge - Manipulation Track benchmark, organized into pretrain, train, and validation splits.
Pretrain split: ~20,000 single pick-and-place trajectories, packaged into tar files (each containing ~1,000 trajectories).
Train split: task-specific demonstrations, with ~100 trajectories provided per task.
Validation split: includes the test-time scenes and object assets in USD format.
Each… See the full description on the dataset page: https://huggingface.co/datasets/InternRobotics/IROS-2025-Challenge-Manip.splatatlas-ir
SplatAtlas-IR: Task-Aligned Evaluation of 3DGS Inverse Rendering for Robot Contact
Release status: Paper under review · Code to be released · Experiment artifacts available in this repository
The numbers, task definitions, and figure on this card match the submitted manuscript. Dated snapshot packages in the repository retain earlier experiment archives for provenance; they should not be read as a replacement for the protocol below.
TL;DR
SplatAtlas-IR… See the full description on the dataset page: https://huggingface.co/datasets/KCBtheone/splatatlas-ir.nfcorpusirish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B.
meta13sphere_IRS_DCE_Topological_Dynamics__Boundary_Dissolution_Physics
Resonance Resonance / IRS-DCE
MASTER README (FULL EXTENDED VERSION)
If you need the other data or pdf check on [https://huggingface.co/datasets/meta13sphere/phaseShift_shell_result_pdf]
[2026-09-22 Update]
MASK_BDP BBRCM: 조건부 운용의 통합 구조
배경·경계·분기·관측·계량을 분해하고 다시 조립하면서, 무엇이 보존되고 어떤 조건에서 결과가 달라지는지 정리한 연구입니다. 상보성 쌍과 제타/RH 관련 변환뿐 아니라 BBRCM 자체에도 같은 변형·스트레스 테스트를 적용했습니다.
주요 성과는 다음과 같습니다.
R32: 원래 함수공간과 정확한 직교사영 조건 아래에서 첫 셀의 잔차를 전체 극한 잔차와… See the full description on the dataset page: https://huggingface.co/datasets/meta13sphere/meta13sphere_IRS_DCE_Topological_Dynamics__Boundary_Dissolution_Physics.ritvij-saxena-iris-detection-pythonHere is the IRIS dataset for the project iris-detection-python.
Official Statement
I hereby declare that I do not own the rights to the dataset used in this project. This dataset was provided by the faculty and utilized solely for educational purposes as part of an assignment for the Biometrics course (CS 559) at the Illinois Institute of Technology.
The dataset is provided for academic and research purposes only, and I encourage others to use it responsibly for similar educational… See the full description on the dataset page: https://huggingface.co/datasets/saxenaritvij/ritvij-saxena-iris-detection-python.ltlomessirveJuly 2025 UPDATE: We released version 1.1, adding almost 200k new queries 🎉🎉🎉.
v1.2 further adds the article titles as columns for convenience.
Use with:
country = "full" # "ar", "bo", ...
version = "1.2"
dataset = datasets.load_dataset("spanish-ir/messirve", country, revision=version)
print(dataset)
Dataset Card for MessIRve
MessIRve is a large-scale dataset for Spanish IR, designed to better capture the information needs of Spanish speakers across different countries.… See the full description on the dataset page: https://huggingface.co/datasets/spanish-ir/messirve.koprospect-ptms-irt
PROSPECT PTMs - Retention Time Prediction
A mass-spectrometry dataset for applied machine learning in proteomics, processed and split for the task of retention time prediction.
Dataset Details
Curated by: Wilhelmlab - Technical University of Munich - School of Life Sciences - Germany
License: CC-BY4.0
Dataset Sources
The data is based on the PROSPECT PTMs datasets hosted in Zenodo.
Repository: https://github.com/wilhelm-lab/PROSPECT
Uses
The… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelmlab/prospect-ptms-irt.RFSD
The Russian Financial Statements Database (RFSD)
The Russian Financial Statements Database (RFSD) is an open, harmonized collection of annual unconsolidated financial statements of the universe of Russian firms:
🔓 First open data set with information on every active firm in Russia.
🗂️ First open financial statements data set that includes non-filing firms.
🏛️ Sourced from two official data providers: the Rosstat and the Federal Tax Service.
📅 Covers 2011-2025, will be… See the full description on the dataset page: https://huggingface.co/datasets/irlspbru/RFSD.ipfs_iran_laws
Laws of Iran
Research snapshot of official legislation collected from DOTIC dotic.ir.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-21
Coverage
catalog-backed incomplete
Source
DOTIC dotic.ir
Collector
scrapers/collect_ir.py
Laws / instruments
5844
Articles
11374
Language
fa
Jurisdiction
Iran
License
ir-dotic
Contents
data/laws.parquet —… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_iran_laws.IRASimnih-chest-xray-14-flatqrVCog-Bench
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
Description:
VCog-Bench is a publicly available zero-shot abstract visual reasoning (AVR) benchmark designed to evaluate Multimodal Large Language Models (MLLMs). This benchmark integrates two well-known AVR datasets from the AI community and includes a newly proposed MaRs-VQA dataset. The findings in VCog-Bench show that current state-of-the-art MLLMs and Vision-Language Models (VLMs), such as GPT-4o… See the full description on the dataset page: https://huggingface.co/datasets/IrohXu/VCog-Bench.iroiro_data■■LECO&DIFF置き場■■
主にXLで使用するLECOが格納されています。
作成者の都合上、数としてはhakushiMix_v14.1 向けのLECO関連が一番充実しています
※2026/6/10 更新
anima_baseV10向けのLECO作成開始しました。
従来のSDXL向け程ピンポイント効果は低いかもですが、プロンプト効果の補助としては有用なようです。
SDXLの頃よりeraseとしての効果は強いため、むしろerase目的で使うのが良いかもしれない……
※簡易な使い方説明は下位フォルダ内txt参照の事
xlam-irrelevance-7.5k
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.RPX
RPX: Robot Perception X
RPX is a real-world RGB-D benchmark for measuring robot perception across scene changes. The canonical naren/all release combines the multi-object, egocentric, single-object, VQA, and tracking metadata that previously lived on separate dataset branches.
Code and benchmark toolkit: github.com/IRVLUTD/RPX
Recommended dataset revision: naren/all (pin the commit SHA printed by your download for reproducible results)
License: Creative Commons Attribution 4.0… See the full description on the dataset page: https://huggingface.co/datasets/IRVLUTD/RPX.iris
Note
The Iris dataset is one of the most popular datasets used for demonstrating simple classification models. This dataset was copied and transformed from scikit-learn/iris to be more native to huggingface.
Some changes were made to the dataset to save the user from extra lines of data transformation code, notably:
removed id column
species column is casted to ClassLabel (supports ClassLabel.int2str() and ClassLabel.str2int())
cast feature columns from float64 down to float32… See the full description on the dataset page: https://huggingface.co/datasets/hitorilabs/iris.malfunction-image-datasetiros2026-ikea-assembly
IKEA Assembly Robot Simulation Dataset
数据集简介
本数据集为机器人挑战赛提供的桌面组装仿真数据,任务为 AssembleTableTask,共 300 条演示轨迹(HDF5 格式)。
数据规模
子集
数量
data
300 条 (.hdf5)
文件结构
data/AssembleTableTask_*.hdf5:单条演示轨迹数据
meta/download_manifest.json:采集记录清单
状态
本数据集当前为私有(private),发布前将补充完整的 schema 说明和加载示例代码。
许可证
CC BY 4.0 — 允许学术和商业使用,需署名。
IRIS
IRIS Dataset: Industrial Real-Sim Imagery Set
Overview
The IRIS Dataset is a comprehensive real-world dataset designed to study sim-to-real transfer for object detection in industrial robotic environments. This repository provides:
The complete real IRIS dataset: 508 annotated images of 32 mechanical components captured across four distinct, challenging industrial scenes.
Assets for synthetic data generation: All necessary 3D models, backgrounds, and materials to… See the full description on the dataset page: https://huggingface.co/datasets/Carraskito/IRIS.kv_retrievalIRIS-CloudDeep
IRIS-CloudDeep
Ground-based long-wave infrared (LWIR) images of the night sky, with the binary ground-truth masks and clear/cloud labels behind Sommer, Kabalan and Brunet (2025), Atmos. Meas. Tech. 18, 2083–2101.
An uncooled FLIR Tau2 microbolometer (640×512, 17 μm pitch, 8–14 μm band, 9 Hz) recorded two night-time campaigns in early 2023 at Prades-le-Lez, France (43°41′51″ N, 3°51′53″ E). A 60 mm f/1.25 lens gives a narrow imaging area of 10.4° × 8.3°, about 58″ per pixel. The… See the full description on the dataset page: https://huggingface.co/datasets/ASKabalan/IRIS-CloudDeep.IRRISIGHT
IRRISIGHT
IRRISIGHT is a large-scale multimodal dataset to address water availability problems in agriculture. It is designed to support supervised and semi-supervised learning tasks related to agricultural water use monitoring.
Due to the space constraints, we uploaded the files across multiple repositories as follows:
To download Pennsylvania and Maryland, use the current repository (OBH30/IRRISIGHT).
To download Arizona, Arkansas, Florida, Georgia, New Jersey, North Carolina… See the full description on the dataset page: https://huggingface.co/datasets/OBH30/IRRISIGHT.farsi-asr-iran-international-raw
Iran International raw Farsi audio archive
This repository preserves 11,993 individually addressable FLAC source files for incremental ASR relabeling and reproducible restoration.
Repository file layout
The first 9,990 FLAC files are stored at the repository root. The remaining 2,003 FLAC files are stored individually under overflow/ to respect Hugging Face's 10,000-entry-per-directory limit.
REMOTE_PATHS.jsonl records every source filename, remote path, byte size… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/farsi-asr-iran-international-raw.AI2_Alphabot_2_sort_angle_iron
AI2_Alphabot_2_sort_angle_iron
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 295
Total Frames: 256772
FPS: 30
Dataset Size: 7.92 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_sort_angle_iron.
