datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepResearch-Bench-II-DatasetDeepResearch-Bench-Dataset
DeepResearch Bench Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the DeepResearch Bench paper. It contains research reports generated by 4 leading deep research AI systems along with detailed human expert annotations evaluating these reports.
DeepResearch Bench is the first comprehensive benchmark for systematically evaluating Deep Research Agents (DRAs) on their ability to handle complex, PhD-level research… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/DeepResearch-Bench-Dataset.Muse-Glimmer-SWE-Gym-2k
Muse-Glimmer-SWE-Gym-2k
Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a
speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and
SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them.
Configs
Config
Rows
Size
What it is
train
1,981
57 MB
One row per trajectory: the full conversation as messages.
raw
159,999
2.7 GB
One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.muse12-nemo-agentic
Muse Spark 1.2 High-Reasoning NeMo Agentic Dataset
A reproducible, verified 24,000-row synthetic agentic dataset generated with Meta Muse Spark 1.2, NeMo Gym, and deterministic task-family verifiers.
The project is a quality-focused successor to r0b0tlab/deepseek-v4-pro-0813-agentic. It keeps rollout prompts separate from reference trajectories and offline-training views, records usage and provenance, and does not publish private chain-of-thought.
[!IMPORTANT]
Status:… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/muse12-nemo-agentic.smithsonian-data
Smithsonian Open Access Data
Pre-processed data dumps from the Smithsonian Open Access initiative, covering millions of objects across Smithsonian Institution museums and archives.
What is this?
The Smithsonian publishes their Open Access metadata on S3, but the raw data is split across 255 individual .txt files per unit. This dataset consolidates each unit's data into a single .jsonl.gz file for easier downloading and processing.
Files
Each file corresponds to… See the full description on the dataset page: https://huggingface.co/datasets/museado/smithsonian-data.UncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.Muse-Glimmer-OPB-100K
Muse Glimmer OPB 100K
On-policy OpenPerfectBlend training data used for DaoCloud/Muse-Glimmer-30B-DSpark.
Prompts are sampled from mlabonne/open-perfectblend, and assistant turns are regenerated on-policy with Muse Glimmer 30B.
The dataset contains 99,984 successfully generated conversations and 148,900 train-turn rows. Responses were regenerated with Muse Glimmer 30B at four reasoning strengths.
Reasoning strength
Conversations
Train-turn rows
low
64,997
96,765… See the full description on the dataset page: https://huggingface.co/datasets/DaoCloud/Muse-Glimmer-OPB-100K.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.ifparse-v1.0
ifparse v1.0
IFParse is a benchmark for structured extraction from real developer logs: can a model turn a raw, unstructured log line into JSON that satisfies a fixed schema, with no code fence, no commentary, and no type errors. Scoring is binary and covers only compliance, not extraction creativity, so the signal is isolated to whether the output would actually parse in a production pipeline.
Each prompt gives the model a single raw log line, either an Apache access log record… See the full description on the dataset page: https://huggingface.co/datasets/Muse-research/ifparse-v1.0.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.MUSE-benchmark
MUSE: Measuring Uncertainty Source Discrimination
MUSE is a behavioral benchmark designed to evaluate how LLMs distinguish between Epistemic (knowledge gaps) and Aleatoric (stochasticity) uncertainty.
Dataset Summary
This dataset contains 200 items across four dimensions:
E-Type: Pure knowledge gaps.
A-Type: Purely stochastic outcomes.
PA (Pseudo-Aleatoric): Deterministic but complex facts (where the "Trap" occurs).
S (Sycophancy): Adversarial social pressure items.
corpus_dominio_museistico_patrimonio
Corpus museístico-patrimonio
Descripción general
El corpus museístico-patrimonio reúne recursos especializados del ámbito museístico y patrimonial, incluyendo tesauros terminológicos, catálogos museísticos y colecciones descriptivas vinculadas al patrimonio cultural. El conjunto representa un registro técnico y descriptivo propio de la documentación patrimonial, la catalogación de bienes culturales y la organización conceptual del conocimiento museístico.
El… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/corpus_dominio_museistico_patrimonio.es-gl_museo_virtual_usc_descricions_parallel
Spanish-Galician USC Virtual Museum Descriptions Parallel Corpus
Dataset Description
es-gl_museo_virtual_usc_descricions_parallel is a Spanish-Galician parallel corpus containing aligned cultural heritage descriptions associated with the Museo Virtual da USC.
The dataset is intended to support machine translation, domain adaptation, and the development of Galician language resources in the cultural heritage and museum domains.
Dataset Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/es-gl_museo_virtual_usc_descricions_parallel.musenet-chunk
