datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gaia2_filesystem
GAIA2 Filesystem
This is a dataset containing files for the GAIA2 benchmark. You should not use this dataset on its own, but instead use the Meta Agents Research Environments framework to execute scenarios from that GAIA2 dataset.
Dataset Link
https://huggingface.co/datasets/meta-agents-research-environments/gaia2
Contact Details
Publishing POC: Meta AI Research Team
Affiliation: Meta Platforms, Inc.
Website:… See the full description on the dataset page: https://huggingface.co/datasets/meta-agents-research-environments/gaia2_filesystem.OmniRooms
UniSHARP:
Universal Sharp Monocular View Synthesis
Meixi Song1 ·
Dizhe Zhang1,* ·
Hao Ren1 ·
Ruiyang Zhang1 ·
Bo Du2 ·
Ming-Hsuan Yang3 ·
Lu Qi1,2,*
1Insta360 Research · 2Wuhan University · 3University of California, Merced
UniSHARP extends SHARP-style photorealistic monocular view synthesis to universal camera systems. Given a single image from a perspective, wide-FoV, fisheye, or panoramic camera, UniSHARP predicts a 3D Gaussian representation and… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/OmniRooms.kaz-vision-50kconceptual_captions
Dataset Card for Conceptual Captions
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.OmegaUse-OfficeVal
OmegaUse-OfficeVal
Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon,
real-world office-suite tasks that span word-processing documents, spreadsheets,
presentations, and cross-file productivity workflows. Tasks are derived from
authentic office requests proposed by practitioners and drawn from freelance
platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.GenIRbankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.Defactify_Image_Dataset
Defactify_Image_Dataset
This dataset is associated with the paper A Comprehensive Dataset for Human vs. AI Generated Image Detection.
📝 Dataset Description
Dataset Summary
The Defactify_Image_Dataset (A Comprehensive Dataset for Human vs. AI Generated Image Detection) is a high-quality collection of 96,000 images and associated metadata designed to benchmark models for detecting and identifying the source of artificially generated content. Built using the MS… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Defactify_Image_Dataset.relaion2B-en-research-safeVAREX
VAREX: A Benchmark for Multi-Modal Structured Extraction from Documents
VAREX (VARied-schema EXtraction) is a benchmark for evaluating multimodal foundation models on structured data extraction from government forms. It comprises 1,777 documents with 1,771 unique schemas across three structural categories, each provided in four input modalities. Ground truth is deterministic — generated via a Reverse Annotation pipeline that programmatically fills PDF templates with synthetic values… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/VAREX.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.cif-dataset
Cracks in the Foundation
A civil-infrastructure visual inspection dataset for instance segmentation with 6 defect/condition categories:
Algae · Crack · Net-Crack · Crack with Precipitation · Rust · Spalling
Each sample is either a full-resolution inspection image or a 1024×1024 tile derived from one.
Tiled samples carry extra fields (tile_row, tile_col, file_name_original, …) that are None for full-resolution samples.
Splits
Each split is its own parquet shard and… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/cif-dataset.relaion2B-multi-research-safelca-bug-localization
🏟️ Long Code Arena (Bug localization)
This is the benchmark for the Bug localization task as part of the
🏟️ Long Code Arena benchmark.
The bug localization problem can be formulated as follows: given an issue with a bug description and a repository snapshot in a state where the bug is reproducible, identify the files within the repository that need to be modified to address the reported bug.
The dataset provides all the required components for evaluation of bug localization… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/lca-bug-localization.hilti-trimble-slam-challenge-2026
Hilti x Trimble SLAM Challenge 2026
The Hilti x Trimble SLAM Challenge 2026 dataset is a real-world robotics benchmark for evaluating visual-inertial SLAM and localization systems on active construction sites.
The dataset combines synchronized dual-fisheye imagery and inertial measurements with building floor plan priors and LiDAR-derived reference trajectories. It was created through a collaboration between Hilti, Trimble, and the Dynamic Robot Systems Group at the University… See the full description on the dataset page: https://huggingface.co/datasets/Hilti-Research/hilti-trimble-slam-challenge-2026.Matterport3D_polished
Matterport3D_polished
Matterport3D_Polished is a panoramic dataset derived from Matterport3D, which was introduced in DiT360.
This dataset contains 10,000+ high-resolution (2048 x 1024) indoor panoramic images along with corresponding prompts.
Compared with the original dataset, it removes the blurred artifacts at both ends, providing clearer and sharper visual details.
Which tasks will benefit from our dataset?
Text-to-Panorama Generation
⚙️ Getting… See the full description on the dataset page: https://huggingface.co/datasets/Insta360-Research/Matterport3D_polished.relaion2B-multi-researchFinMTM
Fin Benchmark v1
统一整理的金融多模态 benchmark,共 11,133 条。原始成品文件保持不变,本目录为独立发布副本。
数据构成
类别
文件
数量
单选题
objective/single_choice.jsonl
1,982
多选题
objective/multiple_choice.jsonl
1,982
L1 理解/空间感知
open_ended/L1.jsonl
2,082
L2 多步数值计算
open_ended/L2.jsonl
1,893
L3 自我纠错
open_ended/L3.jsonl
1,210
L4 多页 Memory
open_ended/L4.jsonl
984
Financial Agent
agent/agent.jsonl
1,000
总计
11,133
关键说明
单选和多选逐行对应同一批 1,982 个题目与图片。
L3 中 605/1,210 条含显式错误… See the full description on the dataset page: https://huggingface.co/datasets/HiThink-Research/FinMTM.VDR_ibm-research_REAL-MM-RAG
VDR_ibm-research_REAL-MM-RAG - Overview
Dataset Summary
VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG).
Dataset Creation
This dataset is a merge and shuffle of the following datasets in the VDR format:
ibm-research/REAL-MM-RAG_TechSlides
ibm-research/REAL-MM-RAG_TechReport
ibm-research/REAL-MM-RAG_FinTabTrainSet
ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.relaion2B-en-researchDAP_dataSDS-KoPub-VDR-Benchmark
📘 Dataset Summary
SDS KoPub-VDR is a benchmark dataset for Visual Document Retrieval (VDR) in the context of
Korean public documents. It contains real-world government document images paired with natural-language
queries, corresponding answer pages, and ground-truth answers. The dataset is designed to evaluate AI models that
go beyond simple text matching, requiring comprehensive understanding of visual layouts, tables, graphs, and images
to accurately locate relevant… See the full description on the dataset page: https://huggingface.co/datasets/SamsungSDS-Research/SDS-KoPub-VDR-Benchmark.relaion1b-nolang-research-safeGEMRec-PromptBook
GEMRec-18k -- Prompt Book
This is the official image dataset for the paper Towards Personalized Prompt-Model Retrieval for Generative Recommendation.
Dataset Intro
GEMRec-18K is a prompt-model interaction dataset with 18K images generated by 200 publicly-available generative models paired with a diverse set of 90 textual prompts. We randomly sampled a subset of 197 models from the full set of models (all finetuned from Stable Diffusion) on Civitai according to the… See the full description on the dataset page: https://huggingface.co/datasets/MAPS-research/GEMRec-PromptBook.OpenMedReason
This data is the Open-PMC derived data part that we discuss in the paper
Medical VQA with Reasoning Traces (Anonymous)
A multiple-choice medical visual question answering benchmark with
chain-of-thought reasoning traces. Each example consists of a medical image
(radiology, pathology, clinical photograph, etc.), a multiple-choice question
with labeled options, a reasoning trace, and the correct answer letter.
This dataset is released anonymously in support of a… See the full description on the dataset page: https://huggingface.co/datasets/researcher2026/OpenMedReason.asset-alignment-reference-views
Asset Alignment Reference Views
Companion dataset for the paper "Rigid 3D Object Alignment: Optimization vs. Feed-Forward Prediction".
Multi-view renderings of correctly assembled source–target pairs: each row
shows one asset already aligned onto its target object, rendered from 12
orbiting viewpoints with RGB and depth.
Where asset-alignment-pairs-905k
shows the asset misaligned and supplies the transformation that fixes it, this
dataset shows the ground-truth assembled result.… See the full description on the dataset page: https://huggingface.co/datasets/macpaw-research/asset-alignment-reference-views.HumanRef-CoT-45k
🦖🧠 Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning 🦖🧠
We propose Rex-Thinker, a Chain-of-Thought (CoT) reasoning model for object referring that addresses two key challenges: lack of interpretability and inability to reject unmatched expressions. Instead of directly predicting bounding boxes, Rex-Thinker reasons step-by-step over candidate objects to determine which, if any, match a given expression.… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-Research/HumanRef-CoT-45k.Funcdex-MT-Function-Calling
Funcdex-MT-Function-Calling Dataset
Funcdex-MT-Function-Calling is a multi-turn function calling dataset designed for training language models to interact with real-world tools and APIs. The dataset contains 1,787 conversations covering 10 individual toolkits and 5 multi-toolkit bundles, with comprehensive system prompts and realistic multi-turn interactions.The code used to generate the dataset can be found here.
Models trained on this dataset have excellent… See the full description on the dataset page: https://huggingface.co/datasets/prem-research/Funcdex-MT-Function-Calling.MetaPKLot-Dataset
MetaPKLot
A Large-Scale Benchmark for Vision-Based Parking Lot Management
2,265,974 labeled samples · 1,366,185 new annotations · 3 research challenges · COCO-style annotations
MetaPKLot is a large-scale, harmonized dataset designed for research on vision-based parking lot management.
It extends and standardizes three existing parking datasets:
PKLot
CNRPark-EXT
PLds
MetaPKLot introduces new annotations, revises existing parking-space annotations, standardizes… See the full description on the dataset page: https://huggingface.co/datasets/DSBD-Research/MetaPKLot-Dataset.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.
