datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
muse_textbooksmuse_textbooksmuse_textbooksMUSE
MUSE: A CAD Design Benchmark with Multi-modal Ground Truth and Rubric-based Evaluation
MUSE is a benchmark of 106 CAD design cases for evaluating language and
multi-modal models on engineering-grade 3D design tasks. Each case pairs a
natural-language design specification with multi-view ground-truth artefacts
(2D engineering drawings + 3D rendered images) and a hand-crafted, rubric-style
evaluation guide.
Why this benchmark
Most CAD/3D benchmarks evaluate either pure… See the full description on the dataset page: https://huggingface.co/datasets/dongxiaoyu/MUSE.omr_benchmark
Muse OMR Benchmark
What this is
A small, clean benchmark dataset for OMR (Optical Music Recognition — recognizing music notation from images/PDFs).
It contains 1077 pairs:
a symbolic music score (the “ground truth”, see dataset fields below)
a corresponding PDF rendering with data augmentation applied
All underlying works are Public Domain.
Why it exists
OMR is often evaluated on private or inconsistent datasets. This dataset aims to provide the community… See the full description on the dataset page: https://huggingface.co/datasets/musegroup/omr_benchmark.MuseVLA-dataset
MuseVLA Dataset
Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic,
thermal, and radar streams. Released as two parts (dataset_01/,
dataset_02/) sharing the same per-episode layout. Together they cover
~1400 episodes across 11 instructions (towel / clothes / box / item / drink
manipulation).
Per-episode contents
{episode_name}/
├── video.mp4 # RGB, 1280×720, 30 fps
├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.MuSEAgent-Evalmuse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.Muse_trainSupersessionBench
SupersessionBench
A behavioral-supersession benchmark for long-term LLM agents: 1,000 multi-session
samples evaluating whether systems honor the user's current state instead of
acting on outdated state from earlier in the conversation history.
This is the data-only mirror of SupersessionBench. The full code release
(reproduction scripts, judge prompts, baselines, paper LaTeX) lives in the
companion GitHub repository. See croissant.json at this repo root for the
complete Croissant… See the full description on the dataset page: https://huggingface.co/datasets/muse0123/SupersessionBench.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.museums-open-italy
Museums Open — Italy 2026 Q2
Registry of Italian museums and galleries derived from OpenStreetMap (ODbL-1.0) and enriched with Wikidata (CC0) identifiers. First iteration of the Museums Open index; v2.0 will add MiC accessibility data and the UK + France extension.
Registro di musei e gallerie italiani derivato da OpenStreetMap (ODbL-1.0) e arricchito con gli identificatori Wikidata (CC0). Prima iterazione del dataset Museums Open; v2.0 aggiungerà i dati MiC di accessibilità e… See the full description on the dataset page: https://huggingface.co/datasets/FedCal/museums-open-italy.Muse_trainMuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B.
The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning.
We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details.
If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.corpus_dominio_museistico_patrimonio
Corpus museístico-patrimonio
Descripción general
El corpus museístico-patrimonio reúne recursos especializados del ámbito museístico y patrimonial, incluyendo tesauros terminológicos, catálogos museísticos y colecciones descriptivas vinculadas al patrimonio cultural. El conjunto representa un registro técnico y descriptivo propio de la documentación patrimonial, la catalogación de bienes culturales y la organización conceptual del conocimiento museístico.
El… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_museistico_patrimonio.tetris-4lpc-survival-krylov
Tetris 4-line PC MDP — death-in-N survival curve (Krylov)
Death-in-N survival curve for the optimal policy of the exact 4-line Perfect Clear
Tetris MDP (see muse918/tetris-4lpc-mdp-vstar-policy
for the underlying V*/policy tables). Built with a Krylov subspace projection
(order k=14) of the per-PC transition operator restricted to the layer-0 boundary
(817,740 quotient-chain states), rather than a direct Monte Carlo estimate.
Caveat: computed against policy pi_policy_v3 (an… See the full description on the dataset page: https://huggingface.co/datasets/muse918/tetris-4lpc-survival-krylov.lmsys_chat_generated_museMuseCraft-Music
🎵 MuseCraft-Music: Million Lyrics Chat Dataset
Train AI models to generate emotionally-rich, contextually-aware lyrics with 992K+ conversation pairs
🚀 Quick Start • 📊 Dataset Info • 💻 Usage • 🏷️ Citation
🎯 What is MuseCraft-Music?
MuseCraft-Music is a comprehensive dataset containing 992,246 conversation pairs designed specifically for training AI models to generate high-quality, emotionally-aware lyrics. Perfect for fine-tuning language models like… See the full description on the dataset page: https://huggingface.co/datasets/alvanalrakib/MuseCraft-Music.MuseGPT-35k-cleanedMeta_Muse_Spark_Distilled_5k
Meta Muse Spark Distilled — 5K Reasoning Traces
Synthetic distillation dataset designed to mirror the thinking and reasoning style of Meta's newest frontier model, Muse Spark (2026). Created May 2026.
This is NOT copied output from Muse Spark. All examples are programmatically generated to teach step-by-step reasoning: Understand → Plan → Execute → Verify.
Contents
muse_spark_distilled_5k.jsonl — 5,000 unique examples, JSON Lines format
preview.csv — first 20… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Meta_Muse_Spark_Distilled_5k.muse-trajectory-annotations
MUSE trajectory annotations
Judge annotations of coding-agent trajectories.
Subset: commit-hook
Event-sequence annotations of 7,593 transcript windows drawn from 433 complete
trajectories of a coding agent working on a git pre-commit-hook task (E1). For
each window the judge identifies the earliest concrete workaround opportunity,
the earliest rejection of a workaround (labelled normative / instrumental /
mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.RAQUEL-MUSE-News-Paraphrase
RAQUEL MUSE-News knowmem paraphrases
One reworded version of each question in the MUSE-News knowledge-memorization (knowmem) QA sets: 100 forget and
100 retain questions. The reference answer is unchanged, so a paraphrase is scored against the same answer as its
original. MUSE-News ships no paraphrased questions; this set fills that gap for the RAQUEL unlearning evaluation,
mirroring the paraphrased_question field that TOFU releases for its forget and retain sets.
Split… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL-MUSE-News-Paraphrase.MuseGPT-20kmuse_syntheticMuseGPT-43knimoz-museum-objects-corpus
NIMoz Museum Objects Corpus — 9 Muzeów (wmuzeach.pl)
Korpus obiektów muzealnych z cyfrowej platformy wmuzeach.pl (NIMoz — Narodowy Instytut Muzealnictwa), wygenerowany z AJAX endpoint /pioro/myajaxlist/objects_list/getcontent/1/ z paginacją i filtrowaniem per sub-kolekcja.
Statystyki
Metryka
Wartość
Muzea
9
Sub-kolekcje
312
Rekordy
10,649
Unikalne tagi
5,694
Znaki
14,716,670
Słowa
1,968,273
Tokeny (szac.)
~2,558,754
Rozmiar (ZIP)
7.3 MB… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/nimoz-museum-objects-corpus.sam-paech__Darkest-muse-v1-details
Dataset Card for Evaluation run of sam-paech/Darkest-muse-v1
Dataset automatically created during the evaluation run of model sam-paech/Darkest-muse-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sam-paech__Darkest-muse-v1-details.muse-datamuse_textbooks
