datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chat2Workflow-Evaluation
Chat2Workflow
Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions.
Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language
Repository: zjunlp/Chat2Workflow
Overview
Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.SciMDR-EvalMTabVQA-Eval
Dataset Card for MTabVQA
Paper
Dataset Description
Dataset Summary
MTabVQA (Multi-Tabular Visual Question Answering) is a novel benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to perform multi-hop reasoning over multiple tables presented as images. This scenario is common in real-world documents like web pages and PDFs but is critically under-represented in existing benchmarks.
The dataset consists of two main parts:
MTabVQA-Eval:… See the full description on the dataset page: https://huggingface.co/datasets/mtabvqa/MTabVQA-Eval.C4-Eval
C4-Eval
C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation.
221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures.
1,105 evaluation instances: five task formulations for every base item.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/ACSE-Eval/ACSE-Eval.Medical_Multimodal_Evaluation_Data
Evaluation Guide
This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks.
To get started:
Download the dataset and extract the images.zip file.
Find evaluation code on our GitHub: HuatuoGPT-Vision.
This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.OpenClaw-EvalMix
OpenClaw EvalMix
OpenClaw EvalMix is a Harbor-format collection of 360 agent-evaluation tasks across four task families. Each task directory includes task.toml, an instruction, an environment definition, and verifier tests.
Composition
Family
Tasks
Local payload
clawbench
19
0.00 GiB
liveclawbench
134
0.07 GiB
pinchbench
147
0.02 GiB
wildclawbench
60
14.05 GiB
The repository preserves each family at the root so task paths remain direct.… See the full description on the dataset page: https://huggingface.co/datasets/zt1106/OpenClaw-EvalMix.R4R-Auto-Eval
R4R Auto Eval
持续开发中的多视角机器人任务成功判定 benchmark 与评测 pipeline。
队友请先阅读 PROJECT_STATUS.md 和
INDEX.md,然后按需查看:
benchmarks/:固定的视频输入、来源记录和分层标签;
pipelines/:判定方法及冻结配置;
runs/:不可覆盖的实验记录;
reports/:工作日志、方法分析和结果限制;
registry/:benchmark、pipeline 和 run 的机器可读索引。
当前范围
multiscene30 是 pipeline 开发集,不是干净的留出测试集;
reassemble40 是来自两个长录像的接触密集型校准集;
当前标签为来源数据提供方标签,尚未全部完成独立人工裁决;
Codex 会话内结果是可行性/协议试验,不等价于独立 API 盲测;
在完成逐来源许可证核查前,本仓库应保持 private。
当前发布版本:0.1.0。
UGround-Offline-EvaluationTexOCR-eval[ACL 2026]
TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction
TexOCR_eval
TexOCR: Benchmarking and Advancing Document OCR Models for Compilable Page-to-LaTeX Reconstruction [ACL 2026 Main]
This repository provides the evaluation dataset for TexOCR, designed to benchmark document OCR models on the task of compilable page-to-LaTeX reconstruction.
Overview
Dataset: TexOCR_eval
Type: Evaluation Set
Task: Document OCR → LaTeX generation… See the full description on the dataset page: https://huggingface.co/datasets/chengyewang/TexOCR-eval.MuSEAgent-EvalACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each… See the full description on the dataset page: https://huggingface.co/datasets/pravin112/ACSE-Eval.cambench_binary_eval
CameraBench Binary Evaluation Dataset
A balanced VQA dataset for evaluating camera motion understanding in videos.
📊 Dataset Statistics
Total Questions: 384
Unique Videos: 119
Unique Questions: 31
Yes Answers: 192 (50.0%)
No Answers: 192 (50.0%)
Balance Ratio: 1.00
Total Size: 126.16 MB (0.12 GB)
Average Video Size: 1.06 MB
🎯 Task Categories
This dataset covers various camera motion tasks including:
Static: 42 questions
Move In: 29 questions
Pan Left: 24… See the full description on the dataset page: https://huggingface.co/datasets/tuhink/cambench_binary_eval.ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/acrever/ACSE-Eval.ACSE-Eval
ACSE-Eval Dataset
This repository contains a comprehensive collection of AWS deployment scenarios and their threat-models used for determining LLMs' threat-modeling capabilities.
Dataset Overview
The dataset consists of 100+ different AWS architecture scenarios, each containing:
Architecture diagrams (architecture.png)
Diagram source code (diagram.py)
Generated CDK infrastructure code
Security threat models and analysis
Directory Structure
Each scenario is… See the full description on the dataset page: https://huggingface.co/datasets/miotac/ACSE-Eval.ViDA-MIPI-dataset
Dataset Card for ViDA-MIPI-dataset
Dataset Structure
Train
Meta Data
train_metadata.json: The meta data file contains image information: MOS, image resolution and distortion attributes.
ViDA-D(ViDA-Description: reasoning quality description )
assess.jsonl: overall quality analysis.
brief_assess.jsonl: structured overall quality analysis.
dist_assess.jsonl: individual distortion assessment.
ViDA-G(ViDA-Grounding: fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/DY-Evalab/ViDA-MIPI-dataset.turkish-medical-vqa-evaluatedMUMU-Eval-6000
MUMU Eval 6000
This repository contains the 6,000-image source-data evaluation set used for
the Florence-2 and LFM2.5-VL-450M baselines in the MUMU evaluation repository.
It is an independently prepared research split, not an official MUMU Challenge
release.
Splits
Split
Images
Ground truth in manifest
validation
1,000
Yes
test
5,000
Yes
The split contains 2,001 Task A samples, 2,000 Task B samples, and 1,999 Task C
samples. All 6,000 image… See the full description on the dataset page: https://huggingface.co/datasets/JinyuLiu/MUMU-Eval-6000.EvalMuseMiniUniGUI-Bench
UniGUI-Bench
A comprehensive benchmark for evaluating GUI agents across 11 capability dimensions.
Dataset Structure
System Evaluation
trajectory: 1,200 complete GUI task trajectories with step-by-step screenshots
gui_system: 7,108 individual step evaluation samples
Process Evaluation (11 Abilities, 200 samples each)
Ability
Description
element_grounding
Locating UI elements given instructions
action_understanding… See the full description on the dataset page: https://huggingface.co/datasets/AGI-Eval/UniGUI-Bench.mmstar-qwen-sidecar-eval
MMStar Qwen Sidecar Eval
Local packaging of Lin-Chen/MMStar used for Qwen3-VL sidecar reproduction.
Standard splits:
validation.jsonl: full 1,500-sample split.
test.jsonl: first 1,000 samples used by the sidecar MMStar 55-point reproduction.
Compatibility files:
mmstar_val.jsonl: same content as validation.jsonl.
mmstar_eval_1k.jsonl: same content as test.jsonl.
images/: image files referenced by the JSONL image field.
The JSONL image paths are relative to the repository… See the full description on the dataset page: https://huggingface.co/datasets/huaXiaKyrie/mmstar-qwen-sidecar-eval.PhysicalAI-US-Evaluation
PhysicalAI-US-Evaluation
A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.PhysicalAI-DE-Evaluation
PhysicalAI-DE-Evaluation
A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective.
Provenance
Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner:
the base VLA pretraining mix,
the reasoning supervised fine-tuning (SFT) stage, and
the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.vla-evaluation-v3franka-insert-siemens-lid-eval-normal-50eval3_TOY_video_images
Eval3 TOY Video Images
Image-question-answer grounding dataset extracted from the Eval 3 TOY celebrity permutation videos.
Each episode contributes three frames from the first 6 seconds:
start frame
middle frame
end frame
The labels use the user-provided episode block ground truth:
episodes 0-29: Taylor Swift
episodes 30-59: Barack Obama
episodes 60-89: Yann LeCun
within each 30-episode block: first 10 left, next 10 middle, last 10 right
Files:
all.jsonl: all 270 examples… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval3_TOY_video_images.aiconf-butterfly-qwen-evalQwen evaluation results for butterfly detection.
Generated at: 2026-04-20 01:14:42 UTC
Rows: 356
franka-insert-siemens-lid-eval-ood-2-20franka-insert-siemens-lid-eval-ood-3-5
