datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.llmail-inject-challenge
Dataset Summary
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge.
We first describe the details of the challenge, and then we provide a documentation of the dataset
For the accompanying code, check out: https://github.com/microsoft/llmail-inject-challenge.
Citation
@article{abdelnabi2025,
title = {LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/llmail-inject-challenge.RHELM
RHELM: Beyond Static Dialogues
Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants.
Unlike benchmarks built around static dialogues, RHELM provides realistic,
heterogeneous, and temporally evolving memory sources, together with
challenging questions that require multi-hop reasoning, temporal synthesis, and
hallucination detection.
⚠️ All characters, events, and personal… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RHELM.XL-DocBench
XL-DocBench
Evidence-grounded reasoning across hundreds or thousands of pages.
Fully verified by 194 human experts.
Hongchen Wei1,†,‡, Yuanzhe Wang2,†,‡,
Bei Liu2,*, Yifan Yang2, Qi Dai2,
Ruichun Ma2, Kai Qiu2, Yunsheng Li2,
Dongdong Chen2, Chong Luo2,
Zhenzhong Chen1, Baining Guo2
1Wuhan University 2Microsoft
†Equal contribution ‡Work done during an internship at MSRA
*Project leader
Project Page ·
Paper ·
Live Leaderboard… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/XL-DocBench.kitab
Overview
🕮 KITAB is a challenging dataset and a dynamic data collection approach for testing abilities of Large Language Models (LLMs) in answering information retrieval queries with constraint filters. A filtering query with constraints can be of the form "List all books written by Toni Morrison that were published between 1970-1980". The dataset was originally contributed by the paper "KITAB: Evaluating LLMs on Constraint Satisfaction for Information Retrieval" Marah I Abdin… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/kitab.ba-calendarMuseVLA-dataset
MuseVLA Dataset
Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic,
thermal, and radar streams. Released as two parts (dataset_01/,
dataset_02/) sharing the same per-episode layout. Together they cover
~1400 episodes across 11 instructions (towel / clothes / box / item / drink
manipulation).
Per-episode contents
{episode_name}/
├── video.mp4 # RGB, 1280×720, 30 fps
├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.delegate52
DELEGATE52
Overview
DELEGATE52 is a benchmark dataset for evaluating LLMs on long-horizon delegated document editing across 52 professional document domains (crystallography files, music notation, accounting ledgers, Python source code, etc.). The dataset was developed to study the readiness of AI systems for delegated workflows, a new interaction paradigm where knowledge workers instruct LLMs to edit documents on their behalf over long sessions.
A detailed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delegate52.mediflow
MediFlow
A large-scale synthetic instruction dataset of 2.5M rows (~700k unique instructions) for clinical natural language processing covering 14 task types and 98 fine-grained input clinical documents.
t-SNE 2D Plot of MediFlow Embeddings by Task Types
Dataset Splits
mediflow: 2.5M instruction data for SFT alignment.
mediflow_dpo: ~135k top-quality instructions with GPT-4o generated rejected_output for DPO alignment.
Main Columns
instruction:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/mediflow.SchGen_dataset
SchGen
SchGen is a dataset of approximately 8K paired natural-language requests and Python-based schematic generation code for research on LLM-driven PCB schematic generation.
The generated Python code can be rendered into KiCad schematic designs, enabling research on hardware generation from natural-language descriptions.
➡️ Paper:
@misc{luo2026schgenpcbschematicgeneration,
title={SchGen: PCB Schematic Generation with Semantic-Grounded Code Representations}… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SchGen_dataset.WildFeedback
Dataset Card for WildFeedback
WildFeedback is a preference dataset constructed from real-world user interactions with ChatGPT. Unlike synthetic datasets that rely solely on AI-generated rankings, WildFeedback captures authentic human preferences through naturally occurring user feedback signals in conversation. The dataset is designed to improve the alignment of large language models (LLMs) with actual human values by leveraging direct user input.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WildFeedback.FStarDataSet
Proof Oriented Programming with AI (PoPAI) - FStarDataSet
This dataset contains programs and proofs in F* proof-oriented programming language.
The data, proposed in Towards Neural Synthesis for SMT-Assisted Proof-Oriented Programming,
is an archive of source code, build artifacts, and metadata assembled from eight different F⋆-based open source projects on GitHub.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet.templatic_generation_tasks
Dataset Card for Active/Passive/Logical Transforms
Dataset Summary
This dataset is a synthetic dataset containing a set of templatic generation tasks using both English and random 2-letter words.
Supported Tasks and Leaderboards
[TBD]
Languages
All data is in English or random 2-letter words.
Dataset Structure
The dataset consists of several subsets, or tasks. Each task contains a train split, a dev split, and a
test split, and multiple… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/templatic_generation_tasks.EpiCoder-func-380k
Dataset Card for EpiCoder-func-380k
Dataset Description
Dataset Summary
The EpiCoder-func-380k is a dataset containing 380k function-level instances of instruction-output pairs. This dataset is designed to fine-tune large language models (LLMs) to improve their code generation capabilities. Each instance includes a detailed programming instruction and a corresponding Python output code that aligns with the instruction.
This dataset is synthesized using methods… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/EpiCoder-func-380k.PatientSafetyBench
Disclaimer
The synthetic prompts may contain offensive, discriminatory, or harmful language. These fake prompts also mention topics that are not based on the scientific consensus at all.These prompts are included solely for the purpose of evaluating safety behavior of language models.
⚠️ Disclaimer: The presence of such prompts does not reflect the views, values, or positions of the authors, their institutions, or any affiliated organizations. They are provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PatientSafetyBench.AgentRx
AgentRx Benchmark
1. Dataset Summary
Name: AgentRx (Agent Root Cause Attribution Benchmark)
Purpose:AgentRx is designed to support research on diagnosing failures in multi-agent LLM systems. The dataset contains failed agent trajectories annotated with step-level failure categories and a designated root cause failure. It enables research on root cause localization, agent debugging, trajectory-level reasoning, and constraint-based supervision
Domains:
tau_retail… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AgentRx.World-R1
World-R1 Prompt Dataset
World-R1 is a prompt-only dataset for text-to-video world simulation. It accompanies World-R1: Reinforcing 3D Constraints for Text-to-Video Generation, where reinforcement learning is used to improve 3D consistency while preserving visual quality and motion diversity in generated videos.The dataset contains English prompts that describe static environments, dynamic scenes, and camera-aware video generation scenarios. It is designed for post-training… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/World-R1.FStarDataSet-V2This dataset is the Version 2.0 of microsoft/FStarDataSet.
Primary-Objective
This dataset's primary objective is to train and evaluate Proof-oriented Programming with AI (PoPAI, in short). Given a specification of a program and proof in F*,
the objective of a AI model is to synthesize the implemantation (see below for details about the usage of this dataset, including the input and output).
Data Format
Each of the examples in this dataset are organized as dictionaries… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/FStarDataSet-V2.CancerGUIDE
Dataset Card for CancerGUIDE Synthetic Patient Data
Dataset Summary
CancerGUIDE Synthetic Patient Data contains synthetically generated oncology patient profiles paired with recommended treatments. The dataset was created using GPT-4.1 following the methodology described in the CancerGUIDE paper, employing both structured and unstructured generation approaches. The resulting data serves as a benchmark for evaluating and training large language models (LLMs) in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CancerGUIDE.Microsoft_LearnSYNUR
Dataset Card: SYNUR (Synthetic Nursing Observation Dataset)
1. Dataset Summary
Name: SYNUR
Full name / acronym: SYnthetic NURsing Observation Extraction
Purpose / use case:SYNUR is intended to support research in structuring nurse dictation transcripts by extracting clinical observations that can feed into flowsheet-style EHR entries.
It is designed to reduce documentation burden by enabling automated conversion from spoken nurse assessments to structured… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SYNUR.prototypical-hai-collaborationsPaper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
LICENSE: ODC-BY
Contact: Sheshera Mysore, Bahar Sarrafzadeh
Introduction
The repository releases code and data for the paper: Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild. The dataset release only contains the public WildChat-1M dataset annotated with labels used for the analysis in the paper.
Dataset contents
wildchat1m_en3u-task_anns.jsonl:… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/prototypical-hai-collaborations.microsoft-Phi-3-mini-4k-instruct-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
microsoft__phi-4-details
Dataset Card for Evaluation run of microsoft/phi-4
Dataset automatically created during the evaluation run of model microsoft/phi-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-4-details.MM-WebGen-Bench
MM-WebGen-Bench: A Benchmark for Multimodal Webpage Generation
MM-WebGen-Bench is a multi-level evaluation benchmark for multimodal webpage generation, proposed in MM-WebAgent. It contains 120 curated webpage design prompts covering 11 scene categories, 11 visual styles, and diverse multimodal compositions (4 video types, 8 image types, and 17 chart types).
Links
Project Page: aka.ms/mm-webagent
GitHub: microsoft/MM-webagent
Paper: MM-WebAgent: A Hierarchical… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MM-WebGen-Bench.microsoft__phi-2-details
Dataset Card for Evaluation run of microsoft/phi-2
Dataset automatically created during the evaluation run of model microsoft/phi-2
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__phi-2-details.BizGenEval
BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation
BizGenEval is a benchmark for evaluating image generation models on real-world commercial design tasks. It covers 5 document types × 4 capability dimensions = 20 evaluation tasks, with 400 curated prompts and 8,000 checklist questions (4,000 easy + 4,000 hard).
Overview
Document Types
Domain
Description
Count
Slides
Presentation slides used in reports, lectures, and business… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/BizGenEval.microsoft__Phi-3-mini-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-mini-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-mini-4k-instruct
The dataset is composed of 73 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 6 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-mini-4k-instruct-details.microsoft__Phi-3-medium-4k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-medium-4k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-medium-4k-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-medium-4k-instruct-details.cleand_microsoft_rStar-Coder元データ: https://huggingface.co/datasets/microsoft/rStar-Coder
データ件数: 269,863
平均トークン数: 11674
最大トークン数: 31,184
合計トークン数: 3,150,447,484
ファイル形式: JSONL
ファイルサイズ: 不明
加工内容
synthetic_sftを使用
トークン処理が重たいので、文字数でフィルター
seed_question < 6000
generation < 80000
thinkタグ除去 が中途半端なものを除外
トークナイズ処理(速度向上アップデート
繰り返し除去
