datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ms_marco
Dataset Card for "ms_marco"
Dataset Summary
Starting with a paper released at NIPS 2016, MS MARCO is a collection of datasets focused on deep learning in search.
The first dataset was a question answering dataset featuring 100,000 real Bing questions and a human generated answer.
Since then we released a 1,000,000 question dataset, a natural langauge generation dataset, a passage ranking dataset,
keyphrase extraction dataset, crawling dataset, and a conversational search.… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/ms_marco.orca-math-word-problems-200k
Dataset Card
This dataset contains ~200K grade school math word problems. All the answers in this dataset is generated using Azure GPT4-Turbo. Please refer to Orca-Math: Unlocking the potential of
SLMs in Grade School Math for details about the dataset construction.
Dataset Sources
Repository: microsoft/orca-math-word-problems-200k
Paper: Orca-Math: Unlocking the potential of
SLMs in Grade School Math
Direct Use
This dataset has been designed to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-math-word-problems-200k.NOTSOFAR
Introduction
Welcome to the "NOTSOFAR-1: Distant Meeting Transcription with a Single Device" Challenge.
This repo contains the baseline system code for the NOTSOFAR-1 Challenge.
For more information about NOTSOFAR, visit CHiME's official challenge website
Register to participate.
Baseline system description.
Contact us: join the chime-8-notsofar channel on the CHiME Slack, or open a GitHub issue.
📊 Baseline Results on NOTSOFAR dev-set-1
Values are presented in… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/NOTSOFAR.rStar-Coder
rStar-Coder Dataset
Project GitHub | Paper
Dataset Description
rStar-Coder is a large-scale competitive code problem dataset containing 418K programming problems, 580K long-reasoning solutions, and rich test cases of varying difficulty levels. This dataset aims to enhance code reasoning capabilities in large language models, particularly in handling competitive code problems.
Experiments on Qwen models (1.5B-14B) across various code reasoning benchmarks demonstrate… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/rStar-Coder.microscope-dataUpdesh_beta
📢 Updesh: Synthetic Multilingual Instruction Tuning Dataset for 13 Indic Languages
NOTE: This is an initial $\beta$-release. We plan to release subsequent versions of Updesh with expanded coverage and enhanced quality control. Future iterations will include larger datasets, improved filtering pipelines.
Updesh is a large-scale synthetic dataset designed to advance post-training of LLMs for Indic languages. It integrates translated reasoning data and synthesized open-domain… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Updesh_beta.timewarp
Timewarp datasets
This dataset contains molecular dynamics simulation data that was used to train the neural networks in the NeurIPS 2023 paper Timewarp: Transferable Acceleration of Molecular Dynamics by Learning Time-Coarsened Dynamics by Leon Klein, Andrew Y. K. Foong, Tor Erlend Fjelde, Bruno Mlodozeniec, Marc Brockschmidt, Sebastian Nowozin, Frank Noé, and Ryota Tomioka.
Please see the accompanying GitHub repository.
This dataset consists of many molecular dynamics trajectories… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/timewarp.libero-microwave-grm-rollouts
LIBERO Microwave GRM Rollouts
Dense-reward-annotated rollout dataset from GRPO training of OpenVLA-OFT on LIBERO-10 Task 9 ("put the yellow and white mug in the microwave and close it").
Dataset
Stat
Value
Episodes
~2,100 (T >= 5 steps)
Format
LeRobot (parquet + images)
Task
put the yellow and white mug in the microwave and close it
Size
41 GB
Reward model
Robo-Dopamine GRM-3B
Policy
OpenVLA-OFT (LoRA, GRPO-trained)
Simulator
LIBERO… See the full description on the dataset page: https://huggingface.co/datasets/Auryal/libero-microwave-grm-rollouts.wiki_qa
Dataset Card for "wiki_qa"
Dataset Summary
Wiki Question Answering corpus from Microsoft.
The WikiQA corpus is a publicly available set of question and sentence pairs, collected and annotated for research on open-domain question answering.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
default
Size of downloaded dataset files: 7.10 MB
Size… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/wiki_qa.AVGen-Bench
AVGen-Bench Generated Videos Data Card
Overview
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/AVGen-Bench.webgym_tasks
WebGym Tasks Dataset
Dataset Description
This dataset contains web navigation tasks for training and evaluating autonomous web agents. Each task consists of a natural language instruction that describes an action to be performed on a specific website, along with evaluation criteria and metadata.
Dataset Summary
Total Training Tasks: 292,092
Total Test Tasks: 1,167
Domains: Multiple domains including Lifestyle & Leisure, Sports & Fitness, and more
Source… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/webgym_tasks.microwakewordThis dataset contains spectrogram features in an mmap ninja format intended to use for microWakeWord training. The features are generated using TensorFlow's microfrontend with the following settings:
sample_rate=16000,
window_size=30,
window_step=20,
num_channels=40,
upper_band_limit=7500,
lower_band_limit=125,
enable_pcan=True,
min_signal_remaining=0.05,
out_scale=1… See the full description on the dataset page: https://huggingface.co/datasets/kahrendt/microwakeword.microsolvated_peptides
Microsolvated Peptide Ensembles
This repository provides explicit-water microsolvated structures for a subset of peptides initialized from ManyPeptidesMD conformers.
Overview
Path
Model
Use
trajectories/v1/
Amber14 in TIP3P, cropped to the nearest 128 waters
Peptide-solvent sampling
webdatasets/v1/train_300K/
Streamable samples from the 300 K state
Model training
webdatasets/v1/train_300-500K/
Streamable samples from all REMD states… See the full description on the dataset page: https://huggingface.co/datasets/niklastr/microsolvated_peptides.IMAGE_UNDERSTANDINGA key question for understanding multimodal performance is analyzing the ability for a model to have basic
vs. detailed understanding of images. These capabilities are needed for models to be used in
real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection
and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting.
The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.cats_vs_dogs
Dataset Card for Cats Vs. Dogs
Dataset Summary
A large set of images of cats and dogs. There are 1738 corrupted images that are dropped. This dataset is part of a now-closed Kaggle competition and represents a subset of the so-called Asirra dataset.
From the competition page:
The Asirra data set
Web services are often protected with a challenge that's supposed to be easy for people to solve, but difficult for computers. Such a challenge is often called a CAPTCHA… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/cats_vs_dogs.RESOURCE2SKILL
Resource2Skill: Executable Agent Skill Libraries
This is the official Microsoft dataset release for
Resource2Skill, a system that
distills human-created multimodal resources into reusable executable skills for
software agents.
Project page: https://microsoft.github.io/Resource2Skill/
Paper: https://arxiv.org/abs/2606.29538
Code: https://github.com/microsoft/Resource2Skill
Contents
skills_wiki/ Structured skill entries used for discovery and inspection… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RESOURCE2SKILL.llmail-inject-challenge
Dataset Summary
This dataset contains a large number of attack prompts collected as part of the now closed LLMail-Inject: Adaptive Prompt Injection Challenge.
We first describe the details of the challenge, and then we provide a documentation of the dataset
For the accompanying code, check out: https://github.com/microsoft/llmail-inject-challenge.
Citation
@article{abdelnabi2025,
title = {LLMail-Inject: A Dataset from a Realistic Adaptive Prompt Injection… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/llmail-inject-challenge.orca-agentinstruct-1M-v1
Dataset Card
This dataset is a fully synthetic set of instruction pairs where both the prompts and the responses have been synthetically generated, using the AgentInstruct framework.
AgentInstruct is an extensible agentic framework for synthetic data generation.
This dataset contains ~1 million instruction pairs generated by the AgentInstruct, using only raw text content publicly avialble on the Web as seeds. The data covers different capabilities, such as text editing, creative… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/orca-agentinstruct-1M-v1.Lurcher_10x
Lurcher 10x Microscopy Dataset
Dataset overview
This dataset consists of 2-D microscopy images of histologically stained 3-D structures in tissue sections through the cerebellum of 21 mouse brains. Animals are grouped into wild-type controls (n = 10) and Lurcher mutant mice (n = 11). The classification task is to distinguish Lurcher mutant mice from wild-type controls.
All images were captured at low magnification (10x) and stained with Cresyl violet, a general… See the full description on the dataset page: https://huggingface.co/datasets/USF-CS-Microscopy-Image-Analysis/Lurcher_10x.SciFormaData-700KSciFormaData-700K: Training Data for Scientific Diagram Generation
SciFormaData-700K is the official training dataset for
SciForma. It contains scientific
methodology-diagram records collected from arXiv papers spanning January
2015–December 2025, structured generation prompts, multi-resolution training
targets, and axis-specific editing triplets.
Features
🧩 Structure-aware prompts. Detailed descriptions organize diagram
components… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SciFormaData-700K.SCBench
SCBench
[Paper]
[Code]
[Project Page]
SCBench (SharedContextBench) is a comprehensive benchmark to evaluate efficient long-context methods in a KV cache-centric perspective, analyzing their performance across the full KV cache lifecycle (generation, compression, retrieval, and loading) in real-world scenarios where context memory (KV cache) is shared and reused across multiple requests.
🎯 Quick Start
Load Data
You can download and load the SCBench data… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/SCBench.CABRA
CABRA: Code Understanding is a Bottleneck for Coding Agents
CABRA (Coding Ability Blueprint for Rigorous Agent evaluation) is a synthetic
code-editing evaluation framework for studying agentic coding capabilities in controlled settings.
This dataset contains generated tasks, test cases, and experiment artifacts.
For code, experiment configurations, and analysis, see the project's GitHub repo:
microsoft/CABRA.
Contents
Directory
Contents
dag/
Generated… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/CABRA.Orchard
Orchard Dataset
Overview
Orchard is the trajectory release accompanying the paper "Orchard: An Open-Source Agentic Modeling Framework" (Peng et al., 2026). It bundles two parallel agentic-modeling datasets distilled from strong teacher models, both produced inside the same Orchard Env sandbox infrastructure:
swe — 107,185 multi-turn software-engineering trajectories across 2,788 GitHub repositories, each labeled with whether the agent's final patch passed the… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Orchard.fineweb-edu-micro
FineWeb-Edu Micro
This dataset is a subset of the FineWeb-Edu Sample-10BT, which contains passages that are at least 1000 tokens long, totalling about 1 Million tokens .
This dataset was primarily made to evaluate different RAG Chunking mechanisms in Chonkie
AI2_Alphabot_2_open_microwave
AI2_Alphabot_2_open_microwave
Dataset Description
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Task Preview
View Video Directly
Overview
Total Episodes: 392
Total Frames: 337551
FPS: 30
Dataset Size: 7.91 GB
Robot Name: AI2_Alphabot_2
End-Effector Type: two_finger_end_effector
Teleoperation Type: vr_controller
Sensors: cam_front_chest_rgb,
cam_front_head_rgb,
cam_left_wrist_rgb… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AI2_Alphabot_2_open_microwave.Dayhoff
Dataset Card for Dayhoff
Dayhoff is an Atlas of both protein sequence data and generative language models — a centralized resource that brings together 3.34 billion protein sequences across 1.7 billion clusters of metagenomic and natural protein sequences (GigaRef), 46 million structure-derived synthetic sequences (BackboneRef), and 16 million multiple sequence alignments (OpenProteinSet). These models can natively predict zero-shot mutation effects on fitness, scaffold structural… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Dayhoff.RHELM
RHELM: Beyond Static Dialogues
Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants.
Unlike benchmarks built around static dialogues, RHELM provides realistic,
heterogeneous, and temporally evolving memory sources, together with
challenging questions that require multi-hop reasoning, temporal synthesis, and
hallucination detection.
⚠️ All characters, events, and personal… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/RHELM.MMLU-CF
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
[📜 Paper] •
[🤗 HF Dataset] •
[🐱 GitHub]
MMLU-CF is a contamination-free and more challenging multiple-choice question benchmark. This dataset contains 10K questions each for the validation set and test set, covering various disciplines.
1. The Motivation of MMLU-CF
The open-source nature of these benchmarks and the broad sources of training data for LLMs have inevitably led to… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MMLU-CF.Taskbench
TaskBench: Benchmarking Large Language Models for Task Automation
Introduction
TaskBench is a benchmark for evaluating large language models (LLMs) on task automation. Task automation can be formulated into three critical stages: task decomposition, tool invocation, and parameter prediction. This complexity makes data collection and evaluation more challenging compared to common NLP tasks. To address this challenge, we propose a comprehensive evaluation framework… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/Taskbench.vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.
