datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.llmail-inject-challenge
Dataset Summary
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge.
We first describe the details of the challenge, and then we provide a documentation of the dataset
For the accompanying code, check out: https://github.com/microsoft/llmail-inject-challenge.
Citation
@article{abdelnabi2025,
title = {LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/llmail-inject-challenge.RHELM
RHELM: Beyond Static Dialogues
Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants.
Unlike benchmarks built around static dialogues, RHELM provides realistic,
heterogeneous, and temporally evolving memory sources, together with
challenging questions that require multi-hop reasoning, temporal synthesis, and
hallucination detection.
⚠️ All characters, events, and personal… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RHELM.vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.kitab
Overview
🕮 KITAB is a challenging dataset and a dynamic data collection approach for testing abilities of Large Language Models (LLMs) in answering information retrieval queries with constraint filters. A filtering query with constraints can be of the form "List all books written by Toni Morrison that were published between 1970-1980". The dataset was originally contributed by the paper "KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval" Marah I Abdin… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/kitab.XL-DocBench
XL-DocBench
Evidence-grounded reasoning across hundreds or thousands of pages.
Fully verified by 194 human experts.
Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡,
Bei Liu2,*, Yifan Yang2, Qi Dai2,
Ruichun Ma2, Kai Qiu2, Yunsheng Li2,
Dongdong Chen2, Chong Luo2,
Zhenzhong Chen1, Baining Guo2
1Wuhan University 2Microsoft
†Equal contribution ‡Work done during an internship at MSRA
*Project leader
Project Page ·
Paper ·
Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.ba-calendarMuseVLA-dataset
MuseVLA Dataset
Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic,
thermal, and radar streams. Released as two parts (dataset_01/,
dataset_02/) sharing the same per-episode layout. Together they cover
~1400 episodes across 11 instructions (towel / clothes / box / item / drink
manipulation).
Per-episode contents
{episode_name}/
├── video.mp4 # RGB, 1280×720, 30 fps
├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.delegate52
DELEGATE52
Overview
DELEGATE52 is a benchmark dataset for evaluating LLMs on long-horizon delegated document editing across 52 professional document domains (crystallography files, music notation, accounting ledgers, Python source code, etc.). The dataset was developed to study the readiness of AI systems for delegated workflows, a new interaction paradigm where knowledge workers instruct LLMs to edit documents on their behalf over long sessions.
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delegate52.mediflow
MediFlow
A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents.
t-SNE 2D Plot of MediFlow Embeddings by Task Types
Dataset Splits
mediflow: 2.5M instruction data for SFT alignment.
mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment.
Main Columns
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.WildFeedback
Dataset Card for WildFeedback
WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.FStarDataSet
Proof Oriented Programming with AI (PoPAI) - FStarDataSet
This dataset contains programs and proofs in F* proof-oriented programming language.
The data, proposed in Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming,
is an archive of source code, build artifacts, and metadata assembled from eight different F⋆-based open source projects on GitHub.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet.SchGen_dataset
SchGen
SchGen is a dataset of approximately 8K paired natural-language requests and Python-based schematic generation code for research on LLM-driven PCB schematic generation.
The generated Python code can be rendered into KiCad schematic designs, enabling research on hardware generation from natural-language descriptions.
➡️ Paper:
@misc{luo2026schgenpcbschematicgeneration,
title={SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations}… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SchGen_dataset.templatic_generation_tasks
Dataset Card for Active/Passive/Logical Transforms
Dataset Summary
This dataset is a synthetic dataset containing a set of templatic generation tasks using both English and random 2-letter words.
Supported Tasks and Leaderboards
[TBD]
Languages
All data is in English or random 2-letter words.
Dataset Structure
The dataset consists of several subsets, or tasks. Each task contains a train split, a dev split, and a
test split, and multiple… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/templatic_generation_tasks.EpiCoder-func-380k
Dataset Card for EpiCoder-func-380k
Dataset Description
Dataset Summary
The EpiCoder-func-380k is a dataset containing 380k function-level instances of instruction-output pairs. This dataset is designed to fine-tune large language models (LLMs) to improve their code generation capabilities. Each instance includes a detailed programming instruction and a corresponding Python output code that aligns with the instruction.
This dataset is synthesized using methods… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/EpiCoder-func-380k.PatientSafetyBench
Disclaimer
The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models.
⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.AgentRx
AgentRx Benchmark
1. Dataset Summary
Name: AgentRx (Agent Root Cause Attribution Benchmark)
Purpose:AgentRx is designed to support research on diagnosing failures in multi-agent LLM systems. The dataset contains failed agent trajectories annotated with step-level failure categories and a designated root cause failure. It enables research on root cause localization, agent debugging, trajectory-level reasoning, and constraint-based supervision
Domains:
tau_retail… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AgentRx.TDC_clearance_microsome_azWorld-R1
World-R1 Prompt Dataset
World-R1 is a prompt-only dataset for text-to-video world simulation. It accompanies World-R1: Reinforcing 3D Constraints for Text-to-Video Generation, where reinforcement learning is used to improve 3D consistency while preserving visual quality and motion diversity in generated videos.The dataset contains English prompts that describe static environments, dynamic scenes, and camera-aware video generation scenarios. It is designed for post-training… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/World-R1.FStarDataSet-V2This dataset is the Version 2.0 of microsoft/FStarDataSet.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming with AI (PoPAI, in short). Given a specification of a program and proof in F*,
the objective of a AI model is to synthesize the implemantation (see below for details about the usage of this dataset, including the input and output).
Data Format
Each of the examples in this dataset are organized as dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet-V2.openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample
A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.OB-Inference-Microtasks
Inference Microtasks
29 synthetic microtasks with reference answers across meeting-notes lookup,
support-ticket triage, and contract-terms extraction. The public set accompanies
the OpenBenchmarks Inference Benchmark,
which measures single-user delay on short, deliberately easy structured tasks.
Dataset contents
Configuration
Rows
Task
contract-terms-extraction
10
Extract commercial terms from a technology contract excerpt.
meeting-notes-lookup
13… See the full description on the dataset page: https://huggingface.co/datasets/openbenchmarks/OB-Inference-Microtasks.CancerGUIDE
Dataset Card for CancerGUIDE Synthetic Patient Data
Dataset Summary
CancerGUIDE Synthetic Patient Data contains synthetically generated oncology patient profiles paired with recommended treatments. The dataset was created using GPT-4.1 following the methodology described in the CancerGUIDE paper, employing both structured and unstructured generation approaches. The resulting data serves as a benchmark for evaluating and training large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CancerGUIDE.pii-masking-micro-100k
PII Masking Micro: Multilingual Sample
A micro-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.microduck-policy-golden-vectors
Microduck policy golden vectors
Observation → action pairs recorded from Pollen Robotics' trained
Microduck policies, so that anybody writing their
own runner can check it against the same numbers instead of against a video.
This is a conformance fixture, not a model and not a dataset to train on. It contains no
weights. If you want the networks, they are Pollen's, in
pollen-robotics/microduck and
pollen-robotics/microduck_rl.
What is in it
golden_policies.json… See the full description on the dataset page: https://huggingface.co/datasets/craigm26/microduck-policy-golden-vectors.SYNUR
Dataset Card: SYNUR (Synthetic Nursing Observation Dataset)
1. Dataset Summary
Name: SYNUR
Full name / acronym: SYnthetic NURsing Observation Extraction
Purpose / use case:SYNUR is intended to support research in structuring nurse dictation transcripts by extracting clinical observations that can feed into flowsheet-style EHR entries.
It is designed to reduce documentation burden by enabling automated conversion from spoken nurse assessments to structured… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SYNUR.Microsoft_Learnprototypical-hai-collaborationsPaper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
LICENSE: ODC-BY
Contact: Sheshera Mysore, Bahar Sarrafzadeh
Introduction
The repository releases code and data for the paper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. The dataset release only contains the public WildChat-1M dataset annotated with labels used for the analysis in the paper.
Dataset contents
wildchat1m_en3u-task_anns.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/prototypical-hai-collaborations.Parameter-Golf-V10-Critical-Memory-FineWeb-MicroMix
Parameter-Golf-V10 Critical Memory FineWeb MicroMix
Version: v10.0.0Language: EnglishFormat: JSONL + critical memory cards + validation scriptsPrimary research objective: build a small, high-signal V10 auxiliary dataset that helps a Parameter Golf workflow move toward an ambitious 0.8 BPB target while preventing false record claims, stale-state drift, and metric mistakes.
Extended description
Parameter-Golf-V10 Critical Memory FineWeb MicroMix is a compact research… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/Parameter-Golf-V10-Critical-Memory-FineWeb-MicroMix.
