datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Terminal-Bench-Hard
Terminal-Bench Hard
Terminal-Bench Hard is a set of 100 challenging terminal-based agent tasks.
The tasks cover software engineering, debugging, data processing, system
administration, security, scientific computing, and related command-line
workflows.
Contents
tasks/: runnable tasks in Harbor format.
metadata/tasks.parquet: searchable task metadata and instructions.
Each task directory contains task.toml, instruction.md, an
environment/ directory, and verifier… See the full description on the dataset page: https://huggingface.co/datasets/Zhongzhi1228/Terminal-Bench-Hard.imagebedPreferImg-Benchminecraft-skins-1.1m-prerenderedshowui-worldmodel-results
ShowUI + WorldModel: Training Results & Artifacts
This dataset contains evaluation results, world models, and training data from the ShowUI + WorldModel integration project.
📦 Contents
1. Evaluation Results (results/miniwob_predictions/)
Size: ~460MB
Format: JSONL files with episode-level predictions
Tasks: 9 MiniWoB++ tasks evaluated with ShowUI agent
Includes:
Task success/failure outcomes
Action predictions and execution traces
World model… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/showui-worldmodel-results.BLINK_RotationQA_90degreesd-datasetbag_grocery_humanA-Survey-on-AI-Agent-Harness
Awesome Agent Harness 🛠️
A curated list of pioneering research papers, tools, and resources on the Agent Harness — the systematic execution layer that transforms raw model capability into sustained, long-horizon autonomy.
A Survey on AI Agent Harness
Agent = Model (Stochastic Intelligence) + Harness (Deterministic Infrastructure)
The survey proposes a Unified Architectural Taxonomy that organizes the Agent Harness as a four-layered stack:
Layer 1: Execution &… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/A-Survey-on-AI-Agent-Harness.xl-genHOIVG-Bench
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
Donghao Zhou1,*, Guisheng Liu2,*, Hao Yang2, Jiatong Li2,†, Jingyu Lin3, Xiaohu Huang4,
Yichen Liu2, Xin Gao2, Cunjian Chen3, Shilei Wen2,§, Chi-Wing Fu1, Pheng-Ann Heng1,§
1The Chinese University of Hong Kong, 2ByteDance, 3Monash University, 4The University of Hong Kong
*Equal contribution, †Project lead, §Corresponding author
🌍 Useful Links
Project Page:… See the full description on the dataset page: https://huggingface.co/datasets/donghao-zhou/HOIVG-Bench.XianKa-Cat-4D-Data
XianKa-Cat-4D-Data
两位抖音博主的视频数据集 + SAM 2.1 语义分割结果,按博主(抖音号/UID)分别打包。
本数据集不包含 4D 重建点云——请使用 4RC 对视频自行运行 4D 重建。
内容
文件/目录
说明
28680959431.tar
博主"显卡从"(抖音号 28680959431):1342 个视频 + metadata.json
28680959431_seg.tar
该博主的 SAM 2.1 分割结果
96187538465.tar
博主"史官"(抖音号 96187538465):569 个视频 + metadata.json
96187538465_seg.tar
该博主的 SAM 2.1 分割结果
examples/
每位博主 10 个样例(视频 + 对应分割),平铺可直接在线浏览
manifest.json
打包清单与统计
视频文件名为 序号_作品ID.mp4;metadata.json 中 aweme_id 与作品 ID… See the full description on the dataset page: https://huggingface.co/datasets/XianKa-Zhong/XianKa-Cat-4D-Data.ScannetppCLoT-Oogiri-GO
Oogiri-GO Dataset Card
Project Page | Paper | Code | Model
Data discription: Oogiri-GO is a multimodal and multilingual humor dataset, and contains more than 130,000 Oogiri samples in English (en.jsonl), Chinese (cn.jsonl), and Japanese (jp.jsonl). Notably, in Oogiri-GO, 77.95% of samples are annotated with human preferences, namely the number of likes, indicating the popularity of a response. As illustrated in Fig. 1, Oogiri-GO contains three types of Oogiri games… See the full description on the dataset page: https://huggingface.co/datasets/zhongshsh/CLoT-Oogiri-GO.ColorBench
🎨 ColorBench
📖 Paper | 💻 GitHub
ColorBench is a multimodal dataset to comprehensively assess capabilities of VLMs in color understanding, including color perception, reasoning, and robustness, introduced in "ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness".
It provides:
More than 5,800 image-text questions covering diverse application scenarios and practical challenges for VLMs evaluation.
3… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/ColorBench.RotationQACoSTAR
🎨 CoSTA* Dataset
CoSTA* is a multimodal dataset for multi-turn image-to-image transformation tasks, designed to accompany the CoSTA* agent presented in CoSTA*: Cost-Sensitive Toolpath Agent for Multi-turn Image Editing. It provides:
High-quality images for various image editing tasks.
Detailed text prompts describing the desired transformations.
Multimodal tasks including inpainting, object recoloring, object segmentation, object replacement, text replacement, and more.
This… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/CoSTAR.PreferThinker-Datasetlanduse_50000BLINK_PositionQA_augmentedportable-hagridv2-mediapipe-hand
portable_hagridv2_mediapipe_hand
这是一个从 HaGRIDv2 派生出来的便携版 MediaPipe Hand 测试数据集,用于 palm detection 和 hand landmark 验证。
它是更大的 hagridv2_512px_mediapipe_hand 生成数据集的过滤子集,体量更小,方便随 ONNX / Ascend 310B 运行时验证工程一起迁移。
数据集包含两类 MediaPipe 风格的目标:
palm_detection:单一 palm 类,包含 palm_bbox_xyxy 和 palm7_keypoints
hand_landmark:同一只手的 full21_keypoints
数据集也保留了原始 HaGRIDv2 手势标签,作为辅助来源标签。这个手势标签不是 palm detector 的类别。
数据集概览
图片总数:9754
数据划分:train、valid、test
划分数量:train=7246,valid=845,test=1663… See the full description on the dataset page: https://huggingface.co/datasets/zhouxzh/portable-hagridv2-mediapipe-hand.ReplicaChartAlignBench
📊 ChartAlignBench
📖 Paper | 💻 GitHub
ChartAlignBench is a multi-modal benchmark designed to evaluate vision-language models (VLMs) on dense-level chart grounding and multi-chart alignment to comprehensively assess fine-grained chart understanding in VLMs.
🌐 Overview
ChartAlignBench contains 9K+ instances, divided into three evaluation subsets:-
Data Grounding & Alignment: Paired charts differ in underlying data values visualized by the chart.
Attribute Grounding &… See the full description on the dataset page: https://huggingface.co/datasets/umd-zhou-lab/ChartAlignBench.AIPhoneTest
🧠 Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images
Welcome to Visual-TableQA, a project designed to generate high-quality synthetic question-answer datasets associated to images of tables. This resource is ideal for training and evaluating models on visually-grounded table understanding tasks such as document QA, table parsing, and multimodal reasoning.
🚀 Latest Update
We have refreshed the dataset with newly generated QA pairs created by… See the full description on the dataset page: https://huggingface.co/datasets/zhousoulpowerII/AIPhoneTest.retinaface_widerface
retinaface_widerface
这个目录用于生成可直接上传到 Hugging Face Datasets 的 WiderFace 教学版数据。
目标格式
转换完成后会生成两个 split:
train
val
对应文件路径默认是:
train/train-00000-of-00001.parquet
val/val-00000-of-00001.parquet
数据来源
train 来自 data/widerface/train/label.txt 和 data/widerface/train/images
val 来自 data/widerface/val/images 和 widerface_evaluate/ground_truth 下的 4 个 mat 文件
运行方式
在仓库根目录执行:
conda run -n retinaface python data/retinaface_widerface/build_parquet.py
如果环境里还没有… See the full description on the dataset page: https://huggingface.co/datasets/zhouxzh/retinaface_widerface.SenseBench
SenseBench
A benchmark for remote sensing low-level visual perception and description in large vision-language models.
🏠 github | 🤗 Hugging Face Subset
Overview
SenseBench is a remote sensing benchmark for evaluating low-level visual perception and description in large vision-language models.
Supported Tasks
Visual question answering
Text generation
Language
English
Data format
Each example contains image paths, a question, an… See the full description on the dataset page: https://huggingface.co/datasets/Zhongchenchen/SenseBench.stream3d-sampleZhongJing-OMNI
ZhongJing-OMNI: The First Multimodal Benchmark for Evaluating Traditional Chinese Medicine
ZhongJing-OMNI is the first multimodal benchmark dataset designed to evaluate Traditional Chinese Medicine (TCM) knowledge in large language models. This dataset provides a diverse array of questions and multimodal data, combining visual and textual information to assess the model’s ability to reason through complex TCM diagnostic and therapeutic scenarios. The unique combination of TCM… See the full description on the dataset page: https://huggingface.co/datasets/CMLM/ZhongJing-OMNI.SenseBench_subset
SenseBench
A benchmark for remote sensing low-level visual perception and description in large vision-language models.
🏠 github | 🤗 Hugging Face Fullset
Overview
SenseBench is a remote sensing benchmark for evaluating low-level visual perception and description in large vision-language models.
Supported Tasks
Visual question answering
Text generation
Language
English
Data format
Each example contains image paths, a question, an… See the full description on the dataset page: https://huggingface.co/datasets/Zhongchenchen/SenseBench_subset.optical-flow-benchmark-samples
KITTI-FC and GoPro-FC Representative Samples
This repository contains a representative mini subset accompanying the
KITTI-FC and GoPro-FC optical-flow robustness benchmarks.
Contents
The samples.zip archive contains 2,824 PNG files, including:
two representative KITTI frame pairs;
one representative pair from each of five GoPro sequences;
clean inputs and the available corruption-severity combinations; and
samples for checking the directory structure and visual… See the full description on the dataset page: https://huggingface.co/datasets/Zhonghua/optical-flow-benchmark-samples.
