datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UncertaintyGym
UncertaintyGym
A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression
Abstract
UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating.
Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.uncertainty-vlm-llama-emnlp_stage
uncertainty-vlm-llama-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-gemma-emnlp_stage
uncertainty-vlm-gemma-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen3-emnlp_stage
uncertainty-vlm-qwen3-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen2p5-emnlp_stage
uncertainty-vlm-qwen2p5-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
test_multidomain
Dataset Card for "test_multidomain"
More Information needed
mikasa-robo-vla-medium-lerobot
mikasa-robo-vla-medium-lerobot
LeRobotDataset v3 datasets for the 5 "Medium"-horizon MIKASA-Robo-VLA memory envs. One
dataset per env, as a top-level snake_case subfolder in this repo (same layout as
mikasa-robo/mikasa-robo-vla-lerobot).
Env code, install instructions, and an integration example for these 5 envs on top of stock
mikasa_robo_suite live in
UncertaintyVLA/mikasa-robo-vla-medium-envs —
start there if you want to run these envs yourself, not just replay the recorded… See the full description on the dataset page: https://huggingface.co/datasets/UncertaintyVLA/mikasa-robo-vla-medium-lerobot.train_akimbio_gemma2test_multilang
Dataset Card for "test_multilang"
More Information needed
uncertainty-vlm-llama-emnlp_test
uncertainty-vlm-llama-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-prm-datatrain_akimbio_mistraluncertainty-vlm-qwen2p5-emnlp_test
uncertainty-vlm-qwen2p5-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-qwen3-emnlp_test
uncertainty-vlm-qwen3-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-prm-traininguncertainty-vlm-llama-emnlpuncertainty-vlm-gemma-emnlp_test
uncertainty-vlm-gemma-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-qwen2p5-shared4way1k_uncertainty_CoT_data_day0_third_path_multi_turnuncertainty-vlm-qwen2p5-emnlpOpenHermes-headlines-2017-2019-uncertaintyuncertainty-vlm-qwen3-emnlpuncertainty-gated-structural-recovery-64800-cases
Uncertainty-Gated Structural Recovery under Syndrome Ambiguity
This dataset contains the complete 64,800-case simulation record used in
Uncertainty-Gated Structural Recovery under Syndrome Ambiguity across
Canonical Stabilizer Codes by Jaeyun Jeong and HwaYoung Jeong.
The experiments evaluate a fixed two-candidate recovery policy family on the
perfect [[5,1,3]], Steane [[7,1,3]], and Shor [[9,1,3]] stabilizer codes.
No code-specific policy retuning was performed.… See the full description on the dataset page: https://huggingface.co/datasets/nogalee/uncertainty-gated-structural-recovery-64800-cases.uncertainty-voice
Vietnamese uncertain voice dataset
Audio hoàn tất: 23
Segment đã gán: 64
Split: test
Parquet có đúng ba cột: audio, label_types, segments.
audio nhúng bytes để phát trực tiếp trong Dataset Viewer.
segments chứa timestamp, label, transcription và metadata chi tiết.
Ý nghĩa các label
Label
Khi nào sử dụng
clear-voice
Có giọng nói và nội dung nghe rõ, không có nhiễu đáng kể ảnh hưởng đến việc nghe hoặc chép lời.
clear-voice-noise
Có giọng nói vẫn nghe… See the full description on the dataset page: https://huggingface.co/datasets/luvox-ai/uncertainty-voice.uncertainty-vlm-llamauncertainty-vlm-qwen3-officialgranular-uncertainty-quantification-dataset
Granular UQ Contrastive Prompts
This dataset contains synthetic contrastive prompt pairs for testing whether a model
or baseline can identify the primary cause of high uncertainty from the prompt alone.
Each row has a first_prompt that is more uncertain with respect to
primary_uncertainty_cause and a second_prompt that is less uncertain with respect
to that same cause.
Schema
Columns: id, primary_uncertainty_cause, domain, first_prompt, second_prompt, subtypes… See the full description on the dataset page: https://huggingface.co/datasets/myyycroft/granular-uncertainty-quantification-dataset.uncertainty-vlm-gemma1k_uncertainty_CoT_data_day0_third_path_multi_turnuncertainty-vlm-llava13b
