datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mimic-medical-imaging-qa
MIMIC Medical Imaging QA Dataset
5,207 Bloom's-taxonomy-stratified question--answer pairs derived from 23 medical imaging lectures (RPI BMED 2300). The dataset supports the paper "MIMIC: A Course-Derivation Pipeline and Benchmark for Slide-Anchored Tutoring with a Domain-Adapted Large Language Model" and was used to fine-tune MIMIC-LM, a domain-adapted Llama-3.1-8B-Instruct model for grounded medical imaging instruction.
License
The benchmark annotations, dataset… See the full description on the dataset page: https://huggingface.co/datasets/zabir1996/mimic-medical-imaging-qa.icd_naive_sft_mimic4_top50mimir-coremimicgen-square-d0-light-seed42-1000-opaque-rerender-lerobotmimicgen-square-d0-lerobot-v3
1.62 GB MimicGen HDF5 → multimodal LeRobot v3
Before → after: HDF5 demonstrations with two RGB streams become a validated, multimodal LeRobot v3.0 subset. Convert HDF5 free →
Community conversion produced by ViaCatalyst BYOD. This repository is not an official upstream release and is not affiliated with the MimicGen authors.
This is a compact, provenance-complete conversion of the first 10 episodes from the pinned MimicGen Square D0 core HDF5 file. It provides a… See the full description on the dataset page: https://huggingface.co/datasets/ViaCatalyst/mimicgen-square-d0-lerobot-v3.tr-hi-mimi-encoded
TR↔HI Mimi-Encoded Parallel Speech
Pre-encoded parallel Turkish↔Hindi speech pairs for training speech-to-speech translation models. All audio has been tokenized through the Mimi neural audio codec (8 codebooks, 12.5 Hz, 24kHz) and stored as .pt files with word-level text alignments.
Dataset Summary
Source audio
~911 hours of synthetic parallel TR↔HI speech from tr-hi-parallel-speech-v2
TTS model
OmniVoice (voice design mode, 14 voice designs)… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/tr-hi-mimi-encoded.mimicgen-square-d0-bg25-seed42-1000-opaque-rerender-lerobotmimic-iv-notesmimic-cxr-reports-summarizationDoppelReflEx__MN-12B-Mimicore-GreenSnake-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-GreenSnake
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-GreenSnake
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-GreenSnake-details.MimicAgent_Skill_Library
MimicAgent Skill Library
A library of 100 natural-language motion-skill specifications for quadruped
and wheeled-quadruped robots. Each entry describes a distinct, kinematically
feasible motion as an animation-style behavior (phase-based base and limb
motion over one cycle).
The library is the input to the MimicAgent text-to-trajectory pipeline,
where an LLM agent turns each skill description into executable code that
generates a motion prior in a MuJoCo simulator.… See the full description on the dataset page: https://huggingface.co/datasets/aicognition/MimicAgent_Skill_Library.fleurs-tr-hi-mimi-encoded
fleurs-tr-hi-mimi-encoded
Mimi-encoded Turkish↔Hindi parallel speech pairs for TinyAya Stage 2
speech-to-speech translation training.
Contents
encoded/*.pt — 9212 Mimi-encoded audio pairs (kyutai/mimi, 8 codebooks, 12.5 Hz, 24 kHz).
Each file keys: pair_id, src_lang, tgt_lang, src_text, tgt_text, src_codes[8, T_src], tgt_codes[8, T_tgt].
encoded/*.alignments.json — 18424 Whisper word-level alignment sidecars
(.src.alignments.json / .tgt.alignments.json).… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/fleurs-tr-hi-mimi-encoded.DoppelReflEx__MN-12B-Mimicore-Nocturne-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Nocturne
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Nocturne
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Nocturne-details.icd_naive_sft_mimic4_fullFLOCK-ja32masterthesis_datacleaned-mimic-1to5k-90k-part2icd_naive_sft_mimic4_top50mimiDoppelReflEx__MN-12B-Mimicore-Orochi-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-details.DoppelReflEx__MN-12B-Mimicore-Orochi-v3-Experiment-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi-v3-Experiment
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi-v3-Experiment
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-v3-Experiment-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-4-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-4
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-4
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-4-details.FLOF04FLOF08icd_naive_sft_mimic3_fullDoppelReflEx__MN-12B-Mimicore-WhiteSnake-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-details.DoppelReflEx__MN-12B-Mimicore-Orochi-v4-Experiment-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-Orochi-v4-Experiment
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-Orochi-v4-Experiment
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-Orochi-v4-Experiment-details.DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-3-details
Dataset Card for Evaluation run of DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-3
Dataset automatically created during the evaluation run of model DoppelReflEx/MN-12B-Mimicore-WhiteSnake-v2-Experiment-3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DoppelReflEx__MN-12B-Mimicore-WhiteSnake-v2-Experiment-3-details.en-tokenized-mimiFLOCK-ja08
