datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.MUSE
MUSE
Multimodal evaluation data.
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
Contents
1,800 test questions and 1,174 referenced images.
Task
Questions
Activity Localization
200
Culture Identification
200
Activity Description
200
Affective Computing
200
Jigsaw Puzzle
200
Object Count
200
Relative Position
200
Remote Interaction
200
Scene Classification
200
Affective Computing consists of four tasks: Object Classification, Emotion… See the full description on the dataset page: https://huggingface.co/datasets/Cyn7hia-Z/MUSE.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.MuSeR_GPT_OSS_120B_DistillationThis dataset contains ~100k synthetic medical queries and corresponding responses distilled from GPT-OSS-120B.
The generation of synthetic medical queries follows an attribute-conditioned generation method proposed in paper Enhancing the Medical Context-Awareness Ability of LLMs via Multifaceted Self-Refinement Learning.
We found that supervised fine-tuning on this dataset can substantially improve LLMs' medical conversational capabilities. See our paper and project page for more details.
If… See the full description on the dataset page: https://huggingface.co/datasets/zyx1234/MuSeR_GPT_OSS_120B_Distillation.MUSE-Bench
MUSE-Bench: Memory Utilization Evaluation Benchmark
Official dataset for the paper "Beyond Memorization: Benchmarking Memory
Utilization in Conversational LLM Agents."
Anonymous release. This repository is an anonymized copy provided for
double-blind peer review. Author and affiliation information is withheld
until the review process concludes.
Motivation
LLM agents increasingly rely on persistent cross-session memory to support
long-horizon and personalized… See the full description on the dataset page: https://huggingface.co/datasets/anonymous111111111/MUSE-Bench.RAQUEL-MUSE-News-Paraphrase
RAQUEL MUSE-News knowmem paraphrases
One reworded version of each question in the MUSE-News knowledge-memorization (knowmem) QA sets: 100 forget and
100 retain questions. The reference answer is unchanged, so a paraphrase is scored against the same answer as its
original. MUSE-News ships no paraphrased questions; this set fills that gap for the RAQUEL unlearning evaluation,
mirroring the paraphrased_question field that TOFU releases for its forget and retain sets.
Split… See the full description on the dataset page: https://huggingface.co/datasets/Hyukkyu/RAQUEL-MUSE-News-Paraphrase.
