datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ESL-Bench
ESL-bench
ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework.
⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.MedHall-Bench
MedHall-Bench
MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework.
⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.MedHarm-Bench
MedHarm-Bench
MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework.
⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker
Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914
Mirod simulation subset: three tasks, approximately 3x real frames
This release contains simulation data only, selected for mixed training with
the local Real_3tasks release. Real recordings are not included.
Folder
Task
Real reference frames
Simulation episodes
Simulation frames
Ratio
task1/dataset
Stack the small box on the other box (叠盒子)
13,785
171
41,358
3.0002
task2/dataset
Put the cup into the tray (杯子入盘)
12,361
192
37,081
2.9998
task3/dataset
Take the box… See the full description on the dataset page: https://huggingface.co/datasets/liujiting/Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914.MiROIRMiroEval-data
MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome
MiroEval is a comprehensive evaluation framework for Deep Research systems, providing automated task generation and assessment across three complementary dimensions: Factual correctness, Point-wise quality, and Process quality.
Quick Start
1. Setup
All three evaluation modules share a single Python environment managed by uv at the repo root:
uv sync
If you use… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroEval-data.MiroVerse-v0.1tts-vc-miro_ar-SA
tts-vc-miro_ar-SA
This is a single-speaker speech dataset for Arabic (Saudi Arabia). It carries the Miro voice, a male voice. Source audio is donor Arabic speech for the Saudi Arabia dialect. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 8189 recordings.
Related links
Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets
phoonnx: https://github.com/TigreGotico/phoonnx… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-vc-miro_ar-SA.MiroMind-M1-SFT-719K
MiroMind-M1
🧾 Overview
Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25.
MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-SFT-719K.MiroVerse-v0.1
MiroVerse: A Reproducible, Full-Trajectory, Ever-Growing Deep Research Dataset
🔥 News & Updates
MiroVerse v0.1 has been released. This dataset can be used with our training framework, MiroTrain. In MiroVerse v0.1, we provide both SFT and DPO data, making it easy to reproduce MiroThinker-v0.1’s benchmark performance on Qwen3. Give it a try!
The initial release of MiroVerse (v0.1) is coming this Friday—stay tuned!
🔥 First Batch of MiroVerse… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroVerse-v0.1.MiroMind-M1-RL-62K
MiroMind-M1
🧾 Overview
Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25.
MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-RL-62K.MiroMind-SFTVideo-To-Dataset-Orchard-Apple
Video-To-Dataset-Orchard-Apple Dataset
Dataset Description
This dataset provides a collection of representative image frames extracted from drone video sequences captured in various agricultural environments, with a particular focus on apple orchards. It also includes scenes from grasslands and panoramic views. The frame extraction and selection process followed the "Video-To-Dataset" methodology developed by the primary author.
The dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/miroslavjaros/Video-To-Dataset-Orchard-Apple.miroeval-benchmark-2026
MiroEval Benchmark 2026
Description
MiroEval Benchmark 2026 is a benchmark for evaluating deep research agents on long-form research tasks. It contains 100 tasks, including 70 text-only tasks and 30 multimodal tasks with accompanying attachments such as PDFs, documents, images, and structured files.
The benchmark is designed to evaluate three complementary aspects of deep research systems:
Synthesis Quality: whether the final report is comprehensive, insightful… See the full description on the dataset page: https://huggingface.co/datasets/anon-ed2026/miroeval-benchmark-2026.tts-vc-kabyle-22khz-kab-dz-miro
tts-vc-kabyle-22khz-kab-dz-miro
This is a single-speaker speech dataset for Kabyle. It carries the Miro voice, a male voice. Source audio comes from boffire/kabyle-piper-22khz, in Algerian (Beber) Kabyle. The audio was converted to the Miro voice identity using voice-conversion. After filtering, the dataset contains 57383 recordings.
Ownership and licensing
No license has been declared upstream, the original data at boffire/kabyle-piper-22khz has no license label.… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/tts-vc-kabyle-22khz-kab-dz-miro.LingxiDiag-16K
LingxiDiag-16K
A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support
Overview
LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues.
The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/LingxiDiag-16K.tts-train-synthetic-miro_en-GB
tts-train-synthetic-miro_en-GB
This is a single-speaker synthetic speech dataset for British English. It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for British English. The dataset contains 1150 recordings, with a metadata.csv transcript file and one WAV file per line.
Related links
Models trained on this dataset:
OpenVoiceOS/pipertts_en-GB_miro
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_en-GB.MiroMind-M1-SFT-719K-transformedeeg-veritas-resultsMiroRL-GenQA
MiroRL-GenQA
A curated dataset for reinforcement learning (RL) training within the MiroRL framework.
Overview
Source: Provided by MiroMind AI as part of the MiroRL project.
Format & Size: Contains ~13.1k examples in Parquet format for efficient loading and processing.
License: Released under CC-BY-NC-4.0 for non-commercial use.
Purpose: Designed to serve as high-quality input for RL fine-tuning in the MiroRL pipeline.
Dataset Structure
Each record… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroRL-GenQA.tts_vc_sabela_gl-ES_miro
tts_vc_sabela_gl-ES_miro
This is a single-speaker speech dataset for Galician. It carries the Miro voice, a male voice. Source audio is the donor voice "Sabela", a Galician speech corpus. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 9999 recordings.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_gl-ES_miro_unicode
Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_sabela_gl-ES_miro.agent-reliability-corpustts-train-synthetic-miro_ar-SA
tts-train-synthetic-miro_ar-SA
This is a single-speaker synthetic speech dataset for Arabic (Saudi Arabia). It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for Arabic (Saudi Arabia). The dataset contains 1099 recordings, with a metadata.csv transcript file and one WAV file per line.
Related links
Models trained on this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_ar-SA.tts_vc_Vaani_marwari_mwr_miro
tts_vc_Vaani_marwari_mwr_miro
This is a single-speaker speech dataset for Marwari. It carries the Miro voice, a male voice. Source audio comes from the Vaani speech corpus, Marwari language. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 4247 recordings.
Related links
Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets
phoonnx: https://github.com/TigreGotico/phoonnx… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_Vaani_marwari_mwr_miro.social-prediction-market-sim
MiroShark Social + Prediction Market Simulation
Agent decisions from MiroShark simulations (GitHub). In each simulation, LLM agents with distinct personas (companies, founders, communities, regulators, commentators) share a Twitter/Reddit-style feed and a Polymarket-style prediction market. Every round, each agent reads the feed (or its portfolio and the open markets) and decides what to do: post, comment, quote, like, follow, buy or sell shares, or do nothing.
Each row is one… See the full description on the dataset page: https://huggingface.co/datasets/MiroShark/social-prediction-market-sim.mirokai_1rgb_popcorn_real_v0.1tts-vc-mcv-scripted-v24.0-fy-nl-miro
tts-vc-mcv-scripted-v24.0-fy-nl-miro
This is a single-speaker speech dataset for West Frisian. It carries the Miro voice, a male voice. Source audio comes from Mozilla Common Voice scripted-speech prompts (release 24.0), read in Dutch/West Frisian. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 9676 recordings.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_fy-NL_miro_unicode
Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-vc-mcv-scripted-v24.0-fy-nl-miro.tts_vc_mcv-scripted-v23.0_an_miro
tts_vc_mcv-scripted-v23.0_an_miro
This is a single-speaker speech dataset for Aragonese. It carries the Miro voice, a male voice. Source audio comes from Mozilla Common Voice scripted-speech prompts (release 23.0). The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 7592 recordings.
Related links
Models trained on this dataset:
OpenVoiceOS/phoonnx_an_miro_unicode
Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_mcv-scripted-v23.0_an_miro.tts-train-synthetic-miro_hi-IN
tts-train-synthetic-miro_hi-IN
This is a single-speaker synthetic speech dataset for Hindi. It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for Hindi. The dataset contains 1100 recordings, with a metadata.csv transcript file and one WAV file per line.
Related links
Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_hi-IN.
