datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
delulu-fim-benchmarkDelulu — Fill-in-the-Middle Code Hallucination Benchmark
A verified multilingual benchmark for code-completion hallucinations.
Every golden completion compiles. Every hallucination provably doesn't.
📄 Read the preprint on arXiv →
Every Delulu sample ships as a self-contained Docker image. The viewer above lets you browse the dataset, pull a sample's verifier, and re-run verify golden / verify hallucinated / verify patch <your-completion> with one… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/delulu-fim-benchmark.WorkflowPerturb
WorkflowPerturb — Dataset Artifact
Companion data for the EMNLP 2026 Industry Track paper
“WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics.”
Canonical location: https://huggingface.co/datasets/microsoft/WorkflowPerturbPaper: https://arxiv.org/abs/2602.17990
This release is the complete WorkflowPerturb benchmark plus documentation. It is
self-contained: the CSVs carry every golden workflow, every perturbed variant, and all
shipped pre-computed… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/WorkflowPerturb.Burmese-Microbiology-1K
Burmese-Microbiology-1K
Min Si Thu, min@globalmagicko.com
Microbiology 1K QA pairs in Burmese Language
Purpose
Before this Burmese Clinical Microbiology 1K dataset, the open-source resources to train the Burmese Large Language Model in Medical fields were rare.
Thus, the high-quality dataset needs to be curated to cover medical knowledge for the development of LLM in the Burmese language
Motivation
I found an old notebook in my box. The book was… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Burmese-Microbiology-1K.
