CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mirobody /ESL-Bench ESL-bench ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework. ⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.textquestion-answering1K<n<10K18 likes6.8k downloads19d agoHugging Face02mirobody /MedHall-Bench MedHall-Bench MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework. ⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.textquestion-answeringn<1K6 likes4.5k downloads1mo agoHugging Face03mirobody /MedHarm-Bench MedHarm-Bench MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework. ⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.textquestion-answeringn<1K2 likes4.3k downloads1mo agoHugging Face04miromind-ai /MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker 7 likes3.9k downloads9mo agoHugging Face05liujiting /Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914 Mirod simulation subset: three tasks, approximately 3x real frames This release contains simulation data only, selected for mixed training with the local Real_3tasks release. Real recordings are not included. Folder Task Real reference frames Simulation episodes Simulation frames Ratio task1/dataset Stack the small box on the other box (叠盒子) 13,785 171 41,358 3.0002 task2/dataset Put the cup into the tray (杯子入盘) 12,361 192 37,081 2.9998 task3/dataset Take the box… See the full description on the dataset page: https://huggingface.co/datasets/liujiting/Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914.video1K<n<10K0 likes1.1k downloads12d agoHugging Face062018haha /MiROIR0 likes372 downloads1y agoHugging Face07miromind-ai /MiroEval-data MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome MiroEval is a comprehensive evaluation framework for Deep Research systems, providing automated task generation and assessment across three complementary dimensions: Factual correctness, Point-wise quality, and Process quality. Quick Start 1. Setup All three evaluation modules share a single Python environment managed by uv at the repo root: uv sync If you use… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroEval-data.documentn<1K1 likes329 downloads6mo agoHugging Face08ZhuofengLi /MiroVerse-v0.1tabular100K<n<1M0 likes325 downloads8mo agoHugging Face09TigreGotico /tts-vc-miro_ar-SA tts-vc-miro_ar-SA This is a single-speaker speech dataset for Arabic (Saudi Arabia). It carries the Miro voice, a male voice. Source audio is donor Arabic speech for the Saudi Arabia dialect. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 8189 recordings. Related links Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets phoonnx: https://github.com/TigreGotico/phoonnx… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-vc-miro_ar-SA.audio-to-audio0 likes306 downloads2mo agoHugging Face10miromind-ai /MiroMind-M1-SFT-719K MiroMind-M1 🧾 Overview Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25. MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-SFT-719K.text100K<n<1M21 likes284 downloads1y agoHugging Face11miromind-ai /MiroVerse-v0.1gated MiroVerse: A Reproducible, Full-Trajectory, Ever-Growing Deep Research Dataset 🔥 News & Updates MiroVerse v0.1 has been released. This dataset can be used with our training framework, MiroTrain. In MiroVerse v0.1, we provide both SFT and DPO data, making it easy to reproduce MiroThinker-v0.1’s benchmark performance on Qwen3. Give it a try! The initial release of MiroVerse (v0.1) is coming this Friday—stay tuned! 🔥 First Batch of MiroVerse… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroVerse-v0.1.textquestion-answering100K<n<1M241 likes223 downloads8mo agoHugging Face12miromind-ai /MiroMind-M1-RL-62K MiroMind-M1 🧾 Overview Training performance of MiroMind-M1-RL-7B on AIME24 and AIME25. MiroMind-M1 is a fully open-source series of reasoning language models built on Qwen-2.5, focused on advancing mathematical reasoning. It is trained through supervised fine-tuning (SFT) on 719K curated problems and reinforcement learning with verifiable rewards (RLVR) on 62K challenging examples, using a context-aware multi-stage policy optimization method… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroMind-M1-RL-62K.text10K<n<100K12 likes170 downloads1y agoHugging Face13LLMcompe-Team-Watanabe /MiroMind-SFTtext100K<n<1M0 likes169 downloads1y agoHugging Face14miroslavjaros /Video-To-Dataset-Orchard-Apple Video-To-Dataset-Orchard-Apple Dataset Dataset Description This dataset provides a collection of representative image frames extracted from drone video sequences captured in various agricultural environments, with a particular focus on apple orchards. It also includes scenes from grasslands and panoramic views. The frame extraction and selection process followed the "Video-To-Dataset" methodology developed by the primary author. The dataset is designed to support research… See the full description on the dataset page: https://huggingface.co/datasets/miroslavjaros/Video-To-Dataset-Orchard-Apple.imagen<1K1 likes161 downloads1y agoHugging Face15anon-ed2026 /miroeval-benchmark-2026 MiroEval Benchmark 2026 Description MiroEval Benchmark 2026 is a benchmark for evaluating deep research agents on long-form research tasks. It contains 100 tasks, including 70 text-only tasks and 30 multimodal tasks with accompanying attachments such as PDFs, documents, images, and structured files. The benchmark is designed to evaluate three complementary aspects of deep research systems: Synthesis Quality: whether the final report is comprehensive, insightful… See the full description on the dataset page: https://huggingface.co/datasets/anon-ed2026/miroeval-benchmark-2026.documentquestion-answeringn<1K0 likes143 downloads2mo agoHugging Face16fdemelo /tts-vc-kabyle-22khz-kab-dz-miro tts-vc-kabyle-22khz-kab-dz-miro This is a single-speaker speech dataset for Kabyle. It carries the Miro voice, a male voice. Source audio comes from boffire/kabyle-piper-22khz, in Algerian (Beber) Kabyle. The audio was converted to the Miro voice identity using voice-conversion. After filtering, the dataset contains 57383 recordings. Ownership and licensing No license has been declared upstream, the original data at boffire/kabyle-piper-22khz has no license label.… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/tts-vc-kabyle-22khz-kab-dz-miro.audio-to-audio0 likes133 downloads12d agoHugging Face17mirobody /LingxiDiag-16K LingxiDiag-16K A Large-Scale Synthetic Psychiatric Dialogue Dataset for Diagnostic Decision Support Overview LingxiDiag-16K is a synthetic psychiatric dialogue dataset containing approximately 16,000 electronic medical records (EMRs) and doctor-patient consultation dialogues. The dataset is designed for evaluating and training LLM-based psychiatric diagnostic decision support systems, with demographically aligned distributions reflecting real-world clinical… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/LingxiDiag-16K.texttext-classification10K<n<100K3 likes116 downloads1mo agoHugging Face18TigreGotico /tts-train-synthetic-miro_en-GB tts-train-synthetic-miro_en-GB This is a single-speaker synthetic speech dataset for British English. It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for British English. The dataset contains 1150 recordings, with a metadata.csv transcript file and one WAV file per line. Related links Models trained on this dataset: OpenVoiceOS/pipertts_en-GB_miro Dataset… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_en-GB.text-to-speech0 likes109 downloads2mo agoHugging Face19LLMcompe-Team-Watanabe /MiroMind-M1-SFT-719K-transformedtext100K<n<1M0 likes106 downloads1y agoHugging Face20mirom11 /eeg-veritas-results1 likes101 downloads2mo agoHugging Face21miromind-ai /MiroRL-GenQA MiroRL-GenQA A curated dataset for reinforcement learning (RL) training within the MiroRL framework. Overview Source: Provided by MiroMind AI as part of the MiroRL project. Format & Size: Contains ~13.1k examples in Parquet format for efficient loading and processing. License: Released under CC-BY-NC-4.0 for non-commercial use. Purpose: Designed to serve as high-quality input for RL fine-tuning in the MiroRL pipeline. Dataset Structure Each record… See the full description on the dataset page: https://huggingface.co/datasets/miromind-ai/MiroRL-GenQA.text10K<n<100K14 likes98 downloads1y agoHugging Face22TigreGotico /tts_vc_sabela_gl-ES_miro tts_vc_sabela_gl-ES_miro This is a single-speaker speech dataset for Galician. It carries the Miro voice, a male voice. Source audio is the donor voice "Sabela", a Galician speech corpus. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 9999 recordings. Related links Models trained on this dataset: OpenVoiceOS/phoonnx_gl-ES_miro_unicode Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_sabela_gl-ES_miro.audio-to-audio0 likes63 downloads2mo agoHugging Face23mirotomasik /agent-reliability-corpustabular10K<n<100K0 likes62 downloads1mo agoHugging Face24TigreGotico /tts-train-synthetic-miro_ar-SA tts-train-synthetic-miro_ar-SA This is a single-speaker synthetic speech dataset for Arabic (Saudi Arabia). It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for Arabic (Saudi Arabia). The dataset contains 1099 recordings, with a metadata.csv transcript file and one WAV file per line. Related links Models trained on this dataset:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_ar-SA.text-to-speech0 likes58 downloads2mo agoHugging Face25TigreGotico /tts_vc_Vaani_marwari_mwr_miro tts_vc_Vaani_marwari_mwr_miro This is a single-speaker speech dataset for Marwari. It carries the Miro voice, a male voice. Source audio comes from the Vaani speech corpus, Marwari language. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 4247 recordings. Related links Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets phoonnx: https://github.com/TigreGotico/phoonnx… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_Vaani_marwari_mwr_miro.audio-to-audio0 likes57 downloads2mo agoHugging Face26MiroShark /social-prediction-market-sim MiroShark Social + Prediction Market Simulation Agent decisions from MiroShark simulations (GitHub). In each simulation, LLM agents with distinct personas (companies, founders, communities, regulators, commentators) share a Twitter/Reddit-style feed and a Polymarket-style prediction market. Every round, each agent reads the feed (or its portfolio and the open markets) and decides what to do: post, comment, quote, like, follow, buy or sell shares, or do nothing. Each row is one… See the full description on the dataset page: https://huggingface.co/datasets/MiroShark/social-prediction-market-sim.tabulartext-generation10K<n<100K1 likes54 downloads2d agoHugging Face27eruedas /mirokai_1rgb_popcorn_real_v0.1tabular10K<n<100K0 likes49 downloads3d agoHugging Face28TigreGotico /tts-vc-mcv-scripted-v24.0-fy-nl-miro tts-vc-mcv-scripted-v24.0-fy-nl-miro This is a single-speaker speech dataset for West Frisian. It carries the Miro voice, a male voice. Source audio comes from Mozilla Common Voice scripted-speech prompts (release 24.0), read in Dutch/West Frisian. The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 9676 recordings. Related links Models trained on this dataset: OpenVoiceOS/phoonnx_fy-NL_miro_unicode Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-vc-mcv-scripted-v24.0-fy-nl-miro.audio-to-audio0 likes45 downloads2mo agoHugging Face29TigreGotico /tts_vc_mcv-scripted-v23.0_an_miro tts_vc_mcv-scripted-v23.0_an_miro This is a single-speaker speech dataset for Aragonese. It carries the Miro voice, a male voice. Source audio comes from Mozilla Common Voice scripted-speech prompts (release 23.0). The audio was converted to the Miro voice identity using voice-conversion. The dataset contains 7592 recordings. Related links Models trained on this dataset: OpenVoiceOS/phoonnx_an_miro_unicode Dataset collection:… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts_vc_mcv-scripted-v23.0_an_miro.audio-to-audio0 likes43 downloads2mo agoHugging Face30TigreGotico /tts-train-synthetic-miro_hi-IN tts-train-synthetic-miro_hi-IN This is a single-speaker synthetic speech dataset for Hindi. It carries the Miro voice, a male voice. The audio was synthesized with text-to-speech and adapted to the Miro speaker identity, to train a Miro voice model for Hindi. The dataset contains 1100 recordings, with a metadata.csv transcript file and one WAV file per line. Related links Dataset collection: https://huggingface.co/collections/TigreGotico/synthetic-tts-datasets… See the full description on the dataset page: https://huggingface.co/datasets/TigreGotico/tts-train-synthetic-miro_hi-IN.text-to-speech0 likes41 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.