CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amongglue /muse_textbookstext1M<n<10M3 likes12k downloads3y agoHugging Face02SkySyrup /muse_textbookstext100K<n<1M1 likes2.4k downloads3y agoHugging Face03ndurkee /muse_textbookstext100K<n<1M0 likes2.2k downloads3y agoHugging Face04dongxiaoyu /MUSE MUSE: A CAD Design Benchmark with Multi-modal Ground Truth and Rubric-based Evaluation MUSE is a benchmark of 106 CAD design cases for evaluating language and multi-modal models on engineering-grade 3D design tasks. Each case pairs a natural-language design specification with multi-view ground-truth artefacts (2D engineering drawings + 3D rendered images) and a hand-crafted, rubric-style evaluation guide. Why this benchmark Most CAD/3D benchmarks evaluate either pure… See the full description on the dataset page: https://huggingface.co/datasets/dongxiaoyu/MUSE.imageimage-to-textn<1K3 likes893 downloads4mo agoHugging Face05musegroup /omr_benchmark Muse OMR Benchmark What this is A small, clean benchmark dataset for OMR (Optical Music Recognition — recognizing music notation from images/PDFs). It contains 1077 pairs: a symbolic music score (the “ground truth”, see dataset fields below) a corresponding PDF rendering with data augmentation applied All underlying works are Public Domain. Why it exists OMR is often evaluated on private or inconsistent datasets. This dataset aims to provide the community… See the full description on the dataset page: https://huggingface.co/datasets/musegroup/omr_benchmark.documentimage-feature-extractionn<1K1 likes582 downloads9mo agoHugging Face06microsoft /MuseVLA-dataset MuseVLA Dataset Multi-modal robot manipulation dataset with synchronized RGB, depth, acoustic, thermal, and radar streams. Released as two parts (dataset_01/, dataset_02/) sharing the same per-episode layout. Together they cover ~1400 episodes across 11 instructions (towel / clothes / box / item / drink manipulation). Per-episode contents {episode_name}/ ├── video.mp4 # RGB, 1280×720, 30 fps ├── mask/video.mp4 #… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/MuseVLA-dataset.tabularrobotics1K<n<10K2 likes478 downloads2mo agoHugging Face07ShijianW01 /MuSEAgent-Evalimage1K<n<10K0 likes217 downloads6mo agoHugging Face08txgsync /muse-k2-vision-pilot-20260910 Muse → K2 bridge: first training experiment Prepared September 10, 2026. This experiment tests whether training a connector lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual encoder. It does not retrain the vision encoder or K2, and it does not establish general screenshot, document, natural-image, or visual reasoning capability. Authorized budget and selected first hardware The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.tabularn<1K0 likes152 downloads16d agoHugging Face09DaoCloud /Muse-Glimmer-OPB-100K Muse Glimmer OPB 100K On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark. Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B. The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths. Reasoning strength Conversations Train-turn rows low 64,997 96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.texttext-generation100K<n<1M4 likes95 downloads2mo agoHugging Face10bolshyC /Muse_traintext10K<n<100K2 likes86 downloads9mo agoHugging Face11muse0123 /SupersessionBench SupersessionBench A behavioral-supersession benchmark for long-term LLM agents: 1,000 multi-session samples evaluating whether systems honor the user's current state instead of acting on outdated state from earlier in the conversation history. This is the data-only mirror of SupersessionBench. The full code release (reproduction scripts, judge prompts, baselines, paper LaTeX) lives in the companion GitHub repository. See croissant.json at this repo root for the complete Croissant… See the full description on the dataset page: https://huggingface.co/datasets/muse0123/SupersessionBench.tabular10K<n<100K0 likes86 downloads5mo agoHugging Face12muset-ai /PALATE PALATE Dataset PALATE contains de-identified human–role-playing-agent conversations, satisfaction annotations, frozen session-level splits, bilingual character cards, and the scoring rubrics used by the PALATE benchmark. Related resources: Code: Zhuyh1139/PALATE Five user-simulator adapters: muset-ai/PALATE-LoRA The dataset stores source annotations rather than ready-to-train examples. Use the processing command in the PALATE GitHub repository to construct role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.tabulartext-generationn<1K1 likes80 downloads2mo agoHugging Face13FedCal /museums-open-italy Museums Open — Italy 2026 Q2 Registry of Italian museums and galleries derived from OpenStreetMap (ODbL-1.0) and enriched with Wikidata (CC0) identifiers. First iteration of the Museums Open index; v2.0 will add MiC accessibility data and the UK + France extension. Registro di musei e gallerie italiani derivato da OpenStreetMap (ODbL-1.0) e arricchito con gli identificatori Wikidata (CC0). Prima iterazione del dataset Museums Open; v2.0 aggiungerà i dati MiC di accessibilità e… See the full description on the dataset page: https://huggingface.co/datasets/FedCal/museums-open-italy.geospatial1K<n<10K1 likes69 downloads3mo agoHugging Face14deehugh /Muse_traintext10K<n<100K0 likes61 downloads13d agoHugging Face15zyx1234 /MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B. The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning. We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details. If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.textquestion-answering10K<n<100K4 likes58 downloads9mo agoHugging Face16proxectonos /corpus_dominio_museistico_patrimonio Corpus museístico-patrimonio Descripción general El corpus museístico-patrimonio reúne recursos especializados del ámbito museístico y patrimonial, incluyendo tesauros terminológicos, catálogos museísticos y colecciones descriptivas vinculadas al patrimonio cultural. El conjunto representa un registro técnico y descriptivo propio de la documentación patrimonial, la catalogación de bienes culturales y la organización conceptual del conocimiento museístico. El… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_museistico_patrimonio.texttext-generation10K<n<100K0 likes48 downloads4mo agoHugging Face17muse918 /tetris-4lpc-survival-krylov Tetris 4-line PC MDP — death-in-N survival curve (Krylov) Death-in-N survival curve for the optimal policy of the exact 4-line Perfect Clear Tetris MDP (see muse918/tetris-4lpc-mdp-vstar-policy for the underlying V*/policy tables). Built with a Krylov subspace projection (order k=14) of the per-PC transition operator restricted to the layer-0 boundary (817,740 quotient-chain states), rather than a direct Monte Carlo estimate. Caveat: computed against policy pi_policy_v3 (an… See the full description on the dataset page: https://huggingface.co/datasets/muse918/tetris-4lpc-survival-krylov.tabularn<1K0 likes39 downloads25d agoHugging Face18AlexHung29629 /lmsys_chat_generated_musetext100K<n<1M0 likes32 downloads1mo agoHugging Face19alvanalrakib /MuseCraft-Music 🎵 MuseCraft-Music: Million Lyrics Chat Dataset Train AI models to generate emotionally-rich, contextually-aware lyrics with 992K+ conversation pairs 🚀 Quick Start • 📊 Dataset Info • 💻 Usage • 🏷️ Citation 🎯 What is MuseCraft-Music? MuseCraft-Music is a comprehensive dataset containing 992,246 conversation pairs designed specifically for training AI models to generate high-quality, emotionally-aware lyrics. Perfect for fine-tuning language models like… See the full description on the dataset page: https://huggingface.co/datasets/alvanalrakib/MuseCraft-Music.text100K<n<1M0 likes30 downloads1y agoHugging Face20oof-baroomf /MuseGPT-35k-cleanedtext10K<n<100K0 likes29 downloads3y agoHugging Face21WithinUsAI /Meta_Muse_Spark_Distilled_5k Meta Muse Spark Distilled — 5K Reasoning Traces Synthetic distillation dataset designed to mirror the thinking and reasoning style of Meta's newest frontier model, Muse Spark (2026). Created May 2026. This is NOT copied output from Muse Spark. All examples are programmatically generated to teach step-by-step reasoning: Understand → Plan → Execute → Verify. Contents muse_spark_distilled_5k.jsonl — 5,000 unique examples, JSON Lines format preview.csv — first 20… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Meta_Muse_Spark_Distilled_5k.text1K<n<10K5 likes29 downloads4mo agoHugging Face22thebajajra /muse-trajectory-annotations MUSE trajectory annotations Judge annotations of coding-agent trajectories. Subset: commit-hook Event-sequence annotations of 7,593 transcript windows drawn from 433 complete trajectories of a coding agent working on a git pre-commit-hook task (E1). For each window the judge identifies the earliest concrete workaround opportunity, the earliest rejection of a workaround (labelled normative / instrumental / mixed / unclear), and the earliest later adoption, with… See the full description on the dataset page: https://huggingface.co/datasets/thebajajra/muse-trajectory-annotations.tabular1K<n<10K0 likes28 downloads1mo agoHugging Face23Hyukkyu /RAQUEL-MUSE-News-Paraphrase RAQUEL MUSE-News knowmem paraphrases One reworded version of each question in the MUSE-News knowledge-memorization (knowmem) QA sets: 100 forget and 100 retain questions. The reference answer is unchanged, so a paraphrase is scored against the same answer as its original. MUSE-News ships no paraphrased questions; this set fills that gap for the RAQUEL unlearning evaluation, mirroring the paraphrased_question field that TOFU releases for its forget and retain sets. Split… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL-MUSE-News-Paraphrase.textquestion-answeringn<1K0 likes28 downloads2d agoHugging Face24oof-baroomf /MuseGPT-20ktext10K<n<100K0 likes13 downloads3y agoHugging Face25ribhu /muse_synthetictext1K<n<10K0 likes13 downloads3y agoHugging Face26oof-baroomf /MuseGPT-43ktext10K<n<100K0 likes11 downloads3y agoHugging Face27PiotrSty /nimoz-museum-objects-corpus NIMoz Museum Objects Corpus — 9 Muzeów (wmuzeach.pl) Korpus obiektów muzealnych z cyfrowej platformy wmuzeach.pl (NIMoz — Narodowy Instytut Muzealnictwa), wygenerowany z AJAX endpoint /pioro/myajaxlist/objects_list/getcontent/1/ z paginacją i filtrowaniem per sub-kolekcja. Statystyki Metryka Wartość Muzea 9 Sub-kolekcje 312 Rekordy 10,649 Unikalne tagi 5,694 Znaki 14,716,670 Słowa 1,968,273 Tokeny (szac.) ~2,558,754 Rozmiar (ZIP) 7.3 MB… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/nimoz-museum-objects-corpus.image10K<n<100K1 likes11 downloads3mo agoHugging Face28open-llm-leaderboard /sam-paech__Darkest-muse-v1-detailsgated Dataset Card for Evaluation run of sam-paech/Darkest-muse-v1 Dataset automatically created during the evaluation run of model sam-paech/Darkest-muse-v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sam-paech__Darkest-muse-v1-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face29yanghuazhi /muse-datatextn<1K0 likes9 downloads2y agoHugging Face30killerx7 /muse_textbookstext1K<n<10K0 likes7 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.