CoolFace
20 results

miro

mirobody /ESL-Bench ESL-bench ESL-bench (Event-driven Synthetic Longitudinal Benchmark) is a virtual health user dataset for evaluating AI health assistants. Each virtual user contains a complete health profile, event timeline, clinical exam data, and knowledge-graph-grounded evaluation queries, designed for use with the Mirobody-Eval framework. ⚠️ Research use only. Outputs are synthetic and intended for benchmarking AI agents. They should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/ESL-Bench.textquestion-answering1K<n<10K18 likes6.8k downloads19d agoHugging Facemirobody /MedHall-Bench MedHall-Bench MedHall-Bench is a field-grounded hallucination detection benchmark for medical AI assistants. It decomposes each clinical response into verifiable structured fields (dose value, unit, reference range, ICD/LOINC code, entity relation, ...) and evaluates AI outputs via per-field programmatic matching in addition to sentence-level LLM-as-Judge. Designed for use with the HolyEval framework. ⚠️ Research use only. Content is for benchmarking AI agents and should not be… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHall-Bench.textquestion-answeringn<1K6 likes4.5k downloads1mo agoHugging Facemirobody /MedHarm-Bench MedHarm-Bench MedHarm-Bench is a red-team compliance benchmark for health-management AI assistants. It uses natural-sounding patient questions that bait the assistant into crossing medical safety boundaries, then scores each response against compliance red lines. Designed for use with the HolyEval framework. ⚠️ Research use only. Questions are designed to elicit unsafe behavior for benchmarking purposes and should not be used for diagnosis or treatment decisions.… See the full description on the dataset page: https://huggingface.co/datasets/mirobody/MedHarm-Bench.textquestion-answeringn<1K2 likes4.3k downloads1mo agoHugging Facemiromind-ai /MiroFlow-BenchmarksThese are the benchmarking datasets used for MiroFlow Framework. More information: https://github.com/MiroMindAI/MiroThinker 7 likes3.9k downloads9mo agoHugging Faceliujiting /Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914 Mirod simulation subset: three tasks, approximately 3x real frames This release contains simulation data only, selected for mixed training with the local Real_3tasks release. Real recordings are not included. Folder Task Real reference frames Simulation episodes Simulation frames Ratio task1/dataset Stack the small box on the other box (叠盒子) 13,785 171 41,358 3.0002 task2/dataset Put the cup into the tray (杯子入盘) 12,361 192 37,081 2.9998 task3/dataset Take the box… See the full description on the dataset page: https://huggingface.co/datasets/liujiting/Mirod-Sim-3Tasks-HighQuality-3xReal-15FPS-20260914.video1K<n<10K0 likes1.1k downloads12d agoHugging Face2018haha /MiROIR0 likes372 downloads1y agoHugging Face