datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AVGen-Bench
AVGen-Bench Generated Videos Data Card
Overview
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.IMAGE_UNDERSTANDINGA key question for understanding multimodal performance is analyzing the ability for a model to have basic
vs. detailed understanding of images. These capabilities are needed for models to be used in
real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection
and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting.
The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.cats_vs_dogs
Dataset Card for Cats Vs. Dogs
Dataset Summary
A large set of images of cats and dogs. There are 1738 corrupted images that are dropped. This dataset is part of a now-closed Kaggle competition and represents a subset of the so-called Asirra dataset.
From the competition page:
The Asirra data set
Web services are often protected with a challenge that's supposed to be easy for people to solve, but difficult for computers. Such a challenge is often called a CAPTCHA… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/cats_vs_dogs.Orchard
Orchard Dataset
Overview
Orchard is the trajectory release accompanying the paper "Orchard: An Open-Source Agentic Modeling Framework" (Peng et al., 2026). It bundles two parallel agentic-modeling datasets distilled from strong teacher models, both produced inside the same Orchard Env sandbox infrastructure:
swe — 107,185 multi-turn software-engineering trajectories across 2,788 GitHub repositories, each labeled with whether the agent's final patch passed the… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Orchard.SciFormaData-700KSciFormaData-700K: Training Data for Scientific Diagram Generation
SciFormaData-700K is the official training dataset for
SciForma. It contains scientific
methodology-diagram records collected from arXiv papers spanning January
2015–December 2025, structured generation prompts, multi-resolution training
targets, and axis-specific editing triplets.
Features
🧩 Structure-aware prompts. Detailed descriptions organize diagram
components… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SciFormaData-700K.microsoftexcelWebSTAR
WebSTAR: WebVoyager Step-Level Trajectories with Augmented Reasoning
Dataset Description
WebSTAR (WebVoyager Step-Level Trajectories with Augmented Reasoning) is a large-scale dataset for training and evaluating computer use agents with step-level quality scores. This dataset is part of the research presented in "Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering" (He et al., 2025).
Unlike traditional trajectory-level filtering approaches… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WebSTAR.Do-You-See-Me
DoYouSeeMe
Overview
The DoYouSeeMe benchmark is a comprehensive evaluation framework designed to assess visual perception capabilities in Machine Learning Language Models (MLLMs). This fully automated test suite dynamically generates both visual stimuli and perception-focused questions (VPQA) with incremental difficulty levels, enabling a graded evaluation of MLLM performance across multiple perceptual dimensions. Our benchmark consists of both 2D and 3D… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Do-You-See-Me.VISION_LANGUAGEA key question for understanding multimodal vs. language capabilities of models is what is
the relative strength of the spatial reasoning and understanding in each modality, as spatial understanding is
expected to be a strength for multimodality? To test this we created a procedurally generatable, synthetic dataset
to testing spatial reasoning, navigation, and counting. These datasets are challenging and by
being procedurally generated new versions can easily be created to combat the effects… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/VISION_LANGUAGE.PEACE
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
[Code] [Paper] [Data]
Introduction
We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table.
Property
Description
Source
USGS(English)
CGS(Chinese)
Content
Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.CUAVerifierBench
CUAVerifierBench: A Human-Annotated Benchmark for Computer-Using-Agent Verifiers
Universal Verifier paper: The Art of Building Verifiers for Computer Use Agents
Dataset Summary
CUAVerifierBench is an evaluation benchmark for verifiers of computer-using agents (CUAs) — i.e. judges that read an agent's trajectory (screenshots + actions + final answer) and decide whether the task was completed correctly. Where benchmarks like WebTailBench measure agents… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CUAVerifierBench.echelon-original-eef-main_right_wristThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "ur3e",
"total_episodes": 193,
"total_frames": 5829,
"total_tasks": 13,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:193"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/echelon-original-eef-main_right_wrist.hnm-search-data
HnM Search Dataset Created from Recommendations Dataset
This synthetic data-set is created using the recommendations dataset:
https://huggingface.co/datasets/einrafh/hnm-fashion-recommendations-data (Use of this dataset is subject to the terms and conditions set forth on the original distribution page. This dataset is intended for non-commercial and research use.)
https://www.kaggle.com/competitions/h-and-m-personalized-fashion-recommendations/data (DATA ACCESS AND USE:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/hnm-search-data.BiomedParseData
BiomedParseData
This is the official dataset repository for "A foundation model for joint segmentation, detection and recognition of biomedical objects across nine modalities".
[Code] [Paper] [Demo] [Model] [Data]
We processed from the below public segmentation datasets, and host a subset of our processed datasets as ZIP files here. Each instance include a 1024x1024 PNG image, a list of textual description for the segmentation target, and a binary groundtruth mask also in 1024x1024… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/BiomedParseData.echelon-original-ja-main_and_right_wristThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "ur3e",
"total_episodes": 193,
"total_frames": 5829,
"total_tasks": 13,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 5,
"splits": {
"train": "0:193"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/echelon-original-ja-main_and_right_wrist.SciFormaBenchSciFormaBench-2K: Structure-Faithful Evaluation of Scientific Diagram Generation
SciFormaBench-2K is the official benchmark for
SciForma. It evaluates whether a
generated scientific diagram faithfully presents the visual entities,
relationships, and text specified by its prompt. The benchmark contains 2,000
samples across three difficulty levels: Simple (500), Medium (900),
and Hard (600).
Features
🎯 Structure-faithful evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SciFormaBench.microsoft-fluentui-emoji-512-whitebg
Dataset Card for "microsoft-fluentui-emoji-512-whitebg"
svg and their file names were converted to images and text from Microsoft's fluentui-emoji repo
microsoft-fluentui-emoji-768
Dataset Card for "microsoft-fluentui-emoji-768"
svg and their file names were converted to images and text from Microsoft's fluentui-emoji repo
Microsoft-ai-agents-for-beginnersMicrosoft100
