datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
uncertainty-vlm-llama-emnlp_stage
uncertainty-vlm-llama-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-gemma-emnlp_stage
uncertainty-vlm-gemma-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen3-emnlp_stage
uncertainty-vlm-qwen3-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
uncertainty-vlm-qwen2p5-emnlp_stage
uncertainty-vlm-qwen2p5-official EMNLP stage-wise train/dev/test splits
This dataset contains deterministic stratified 80/10/10 train/dev/test assignments for the four diagram stages. Each split uses the stage_label column as the binary target for that stage.
VIVA_Benchmark_EMNLP24
VIVA: A Benchmark for Vision-Grounded Decision-Making with Human Values
Zhe Hu1,
Yixiao Ren1,
Jing Li1,
Yu Yin2
1The Hong Kong Polytechnic University
2Case Western Reserve University
EMNLP 2024 (main)
📄 Paper
🌎 Website
💻 Code
Introduction
This is the official huggingface repo providing the benchmark of our EMNLP'24 paper:
VIVA: A Benchmark for… See the full description on the dataset page: https://huggingface.co/datasets/zhehuderek/VIVA_Benchmark_EMNLP24.uncertainty-vlm-llama-emnlp_test
uncertainty-vlm-llama-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
seqBench
SeqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMs
Description
SeqBench is a programmatically generated benchmark designed to rigorously evaluate and analyze the sequential reasoning capabilities of language models. Task instances involve pathfinding in 2D grid environments, requiring models to perform multi-step inference over a combination of relevant and distracting textual facts.
The benchmark allows for fine-grained, orthogonal control over… See the full description on the dataset page: https://huggingface.co/datasets/emnlp-submission/seqBench.uncertainty-vlm-qwen3-emnlp_test
uncertainty-vlm-qwen3-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-gemma-emnlp_test
uncertainty-vlm-gemma-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-qwen2p5-emnlp_test
uncertainty-vlm-qwen2p5-official EMNLP stage-wise test splits
This dataset contains only the test split from deterministic stratified
80/10/10 train/dev/test assignments for the four diagram stages.
uncertainty-vlm-llama-emnlpuncertainty-vlm-qwen2p5-emnlpuncertainty-vlm-qwen3-emnlpuncertainty-vlm-gemma-emnlpEMNLP-NLLP-CODEswitch to private once done with paper contents, add the files along with rest on github :)
DITrans-EMNLP
