datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ESL-Bench
ESL-bench
ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework.
⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.MedHall-Bench
MedHall-Bench
MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework.
⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.MedHarm-Bench
MedHarm-Bench
MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework.
⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.MiroVerse-v0.1
MiroVerse: A Reproducible, Full-Trajectory, Ever-Growing Deep Research Dataset
🔥 News & Updates
MiroVerse v0.1 has been released. This dataset can be used with our training framework, MiroTrain. In MiroVerse v0.1, we provide both SFT and DPO data, making it easy to reproduce MiroThinker-v0.1’s benchmark performance on Qwen3. Give it a try!
The initial release of MiroVerse (v0.1) is coming this Friday—stay tuned!
🔥 First Batch of MiroVerse… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroVerse-v0.1.miroeval-benchmark-2026
MiroEval Benchmark 2026
Description
MiroEval Benchmark 2026 is a benchmark for evaluating deep research agents on long-form research tasks. It contains 100 tasks, including 70 text-only tasks and 30 multimodal tasks with accompanying attachments such as PDFs, documents, images, and structured files.
The benchmark is designed to evaluate three complementary aspects of deep research systems:
Synthesis Quality: whether the final report is comprehensive, insightful… See the full description on the dataset page: https://huggingface.co/datasets/anon-ed2026/miroeval-benchmark-2026.LingxiDiag-16K
LingxiDiag-16K
A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support
Overview
LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues.
The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/LingxiDiag-16K.openseeker-miroverse-mix-full
OpenSeeker + MiroVerse SFT mix (full)
All-data union of two deep-research agent SFT datasets in the OpenDR-eval agent wire
format (OpenAI-native messages, tools = search/visit, final answer wrapped in
<answer>...</answer>).
split
rows
composition
train
28691
4885 OpenSeeker + 23806 MiroVerse
validation
320
held-out
Columns: messages, tools, question, answer, n_tool_calls, source.
Why "full" rather than 1:1-by-rows
The earlier… See the full description on the dataset page: https://huggingface.co/datasets/Zephyr271828/openseeker-miroverse-mix-full.openseeker-miroverse-mix-1to1
OpenSeeker + MiroVerse 1:1 SFT mix
A 1:1 (by row count) mix of two deep-research agent SFT datasets, in the OpenDR-eval agent
wire format (OpenAI-native messages, tools = search/visit, final answer in <answer>...</answer>).
split
rows
composition
train
9,770
4,885 OpenSeeker + 4,885 MiroVerse
validation
128
64 + 64
Columns: messages, tools, question, answer, n_tool_calls, source.
OpenSeeker half: from Zephyr271828/openseeker_v1_sft.
MiroVerse half: HotpotQA /… See the full description on the dataset page: https://huggingface.co/datasets/Zephyr271828/openseeker-miroverse-mix-1to1.formatted_miromind-1000firstdataset
