datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crello
Dataset Card for Crello
Dataset Description
The Crello dataset is a collection of raster graphic designs originally compiled for the study of vector graphic documents. It contains document meta-data such as canvas size and pre-rendered elements such as images or text boxes. The original templates were collected from crello.com (now create.vista.com) and converted to a low-resolution format suitable for machine learning analysis. More recently, it has been used for… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/crello.mvscps
Neural Multi-View Self-Calibrated Photometric Stereo without Photometric Stereo Cues
This dataset contains multi-view One-Light-at-a-Time (OLAT) images captured for multi-view 3D reconstruction and inverse rendering tasks.
It is first introduced and used for qualitative evaluation in our ICCV 2025 paper.
Dataset Structure
The dataset contains 6 scenes. Each scene directory is organized as follows:
{capture_date}_{material}_{name}/
├── ARW/
├── JPG/
├── mask/
├── CAM/
├──… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/mvscps.cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.MultiPriv
🔐MultiPriv: A Multilingual & Multimodal Dataset of PII Entities and Prompts for LLM Privacy Risk Research
多语言多模态 PII 实体与 Prompt 数据集 —— MultiPriv 数据集(面向大模型的隐私风险研究)
❗Due to the limitations of open-source certificates, attribute-level VLM images cannot be directly published in the repository. We will provide links to each image used in our dataset
由于开源证书限制,属性级 VLM 图像无法直接公布在仓库里,我们会整理我们数据集用到的每一张图片链接
🎉 News
[2026.06] 🎊 Our MultiPriv has been accepted to ICML 2026… See the full description on the dataset page: https://huggingface.co/datasets/CyberChangAn/MultiPriv.VLG-Loc-Dataset
Vision-Language Global Localization (VLG-Loc) Dataset
This dataset is for evaluation of Vision-Language Global Localization (VLG-Loc).
Dataset Structure
Each dataset directory contains the following files. Note: All camera images (.png) are pre-corrected for lens distortion.
left_camera_image.png: Image from the rear-left camera of the robot.
center_camera_image.png: Image from the front-facing camera of the robot.
right_camera_image.png: Image from the rear-right camera… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/VLG-Loc-Dataset.Cabin-Human-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
核心特点:
丰富的场景多样性: 涵盖不同年龄、性别、种族和衣着风格的虚拟人模型,以及多种驾驶与乘坐行为(如使用手机、喝水、疲劳、手势)和面部表情。
专为座舱感知优化: 数据集可直接用于智能座舱端侧视觉模型,尤其是 DMS/OMS 算法的训练、微调与验证,帮助模型精准理解座舱内复杂的交互与状态。… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-Behavior-Dataset.OTR
OTR: Overlay Text Removal Dataset
OTR (Overlay Text Removal) is a synthetic benchmark dataset designed to advance research of text removal from images.It features complex, object-aware text overlays with clean, artifact-free ground truth images, enabling more challenging evaluation scenarios beyond traditional scene text datasets.
📦 Dataset Overview
Subset
Source Dataset
Content Type
# SamplesNotes
OTR-easy (test set)
MS-COCO
Simple backgrounds (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/OTR.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.cyberpunk-gameplay-data
赛博朋克
This public dataset repository contains local gameplay data uploaded from F:\赛博朋克.
Contents
Files: 473
Total local size: 363.44 GB
Generated: 2026-06-09 08:27:59 UTC
File Types
.jsonl: 144
.json: 110
.png: 108
.mkv: 36
.txt: 36
.parquet: 36
.exe: 2
.mp4: 1
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/cyberpunk-gameplay-data.cyberpunkedgerunners
Bangumi Image Base of Cyberpunk: Edgerunners
This is the image base of bangumi Cyberpunk: Edgerunners, we detected 21 characters, 1227 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/cyberpunkedgerunners.packingbench-assets
PackingBench Assets
Simulation assets for the PackingBench robot-manipulation benchmark
(Cybernetic Labs). All scene files are OpenUSD, Z-up, meters.
Structure
backgrounds/
packing_area_simple/ Self-contained warehouse packing-area scene in
scene.usd NVIDIA Isaac Sim's Simple_Warehouse shell
(~24x39 m hall). Flattened stage: warehouse
shell, packing bench, dressing props… See the full description on the dataset page: https://huggingface.co/datasets/Cybernetic-Labs/packingbench-assets.cyberpunk-2077-gameplay-data
赛博朋克2077
This public dataset repository contains local gameplay data uploaded from F:\赛博朋克2077.
Contents
Files: 363
Total local size: 130.15 GB
Generated: 2026-06-06 01:12:32 UTC
File Types
.jsonl: 112
.json: 84
.png: 83
.txt: 28
.mkv: 28
.parquet: 27
.jpg: 1
Notes
This repository may contain gameplay video, images, Parquet files, JSON/JSONL metadata, and keyboard/mouse event logs.
The license is marked as other; review game… See the full description on the dataset page: https://huggingface.co/datasets/xiaoluo11/cyberpunk-2077-gameplay-data.IBCBench
IBCBench: Image Bundle Composition Benchmark
IBCBench is the benchmark introduced in Weaving Visual Narratives: Agentic Image Bundle Composition Beyond Atomic Visual Matching, accepted to the EMNLP 2026 Main Conference.
Paper | GitHub
Overview
Image Bundle Composition (IBC) shifts image retrieval from independently ranking images to dynamically composing a compact, cohesive bundle whose images jointly satisfy relational, temporal, spatial, or narrative… See the full description on the dataset page: https://huggingface.co/datasets/CyberDancer/IBCBench.Cabin-Human-ABNORMAL-Behavior-Dataset
全球最大的智能座舱多模态开源高质量数据集来啦!
一. 数据集摘要 (Dataset Summary)
「CyberData塞塔」智能座舱用户行为数据集是一个专为加速智能座舱感知算法开发而设计的高质量、程序化生成的图像数据集。随着 C-NCAP、EU GSR 等全球汽车安全法规对驾驶员监控系统 (DMS) 和乘客监控系统 (OMS) 提出更高要求,安全、合规、多样化的训练数据变得至关重要。本数据集通过合成方式,旨在解决真实世界数据采集面临的隐私风险、高昂成本和长尾场景覆盖不足等核心挑战。
该数据集包含 5,000 张 由 XAI Lab 自主研发的数据集生成引擎合成的高保真座舱内用户行为图像,每张图像都附带丰富的、100% 精确的标注信息。
数据格式
数据集以JSON格式提供,包含以下字段:
image_id: 图像ID
image_path: 图像路径
category: 行为类别
tags: 行为标签
behaviors: 包含左右乘客行为描述的对象
left_passenger: 左侧乘客行为描述… See the full description on the dataset page: https://huggingface.co/datasets/XAILab-CyberSpark/Cabin-Human-ABNORMAL-Behavior-Dataset.cyberpunk_2077_recordings_01
赛博朋克2077 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ba08ec7bd5c6cd5a0c3bc002f6c5cfdf
Collection: general (泛数据)
Recordings: 248
Layout: recordings/<recording_id>/<raw component>
rl-game-traces-cyberpunk-2077
赛博朋克2077
This public dataset repository contains gameplay trace data uploaded from F:\赛博朋克2077.
Contents
Files: 382
Total local size: 148.78 GB
Generated: 2026-06-14T06:08:28+00:00
File Types
.jsonl: 116
.png: 90
.json: 86
.parquet: 32
.mkv: 28
.txt: 28
.mp4: 1
.jpg: 1
Notes
This repository may contain gameplay video, Parquet files, JSON/JSONL metadata, and input event logs.
The license is marked as other; review game footage… See the full description on the dataset page: https://huggingface.co/datasets/yinhuankuang/rl-game-traces-cyberpunk-2077.meta-cyberseceval-imagesguided_cyberpunk_2077_recordings_01
赛博朋克2077 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_ba08ec7bd5c6cd5a0c3bc002f6c5cfdf
Collection: guided (精数据)
Recordings: 518
Layout: recordings/<recording_id>/<raw component>
gui_and_terminal_operation_recording
GUI/Terminal Operation Recording Dataset
General
This dataset is created by Cybergod Web GUI/Terminal Recorder.
The zipfile "./record-2025-8-9.7z" is of the same structure as the folder "./record".
Recording Format
Common file contents:
description.txt: Description of the recording. Pure text.
begin_recording.txt: Timestamp of when the recording started. JSON format.
Example: {"timestamp": <timestamp: float>, "event": "begin_recording"}
stop_recording.txt:… See the full description on the dataset page: https://huggingface.co/datasets/cybergod-agi/gui_and_terminal_operation_recording.HARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.kovidore-v2-cybersecurity-beirKoViDoRe v2 : Cybersecurity
This dataset, Cybersecurity, is a corpus of technical reports on cyber threat trends and security incident responses in Korea, intended for complex-document understanding tasks. It is one of the 4 corpora comprising the KoViDoRe v2 Benchmark.
Links
Github: https://github.com/whybe-choi/kovidore-benchmark
Collection: https://huggingface.co/collections/whybe-choi/kovidore-benchmark-beir-v2
Data Generation Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-cybersecurity-beir.crello-animation
Crello Animation
Dataset Summary
The Crello Animation dataset is a collection of animated graphics. Animated graphics are videos composed of multiple sprites, with each sprite rendered by animating a texture. The textures are static images, and the animations involve time-varying affine warping and opacity. The original templates were collected from create.vista.com and converted to a low-resolution format suitable for machine learning analysis.
Usage… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/crello-animation.Deepfake_dataset_cybersentinal
Dataset Card for OpenFake
OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/vjjoshi23/Deepfake_dataset_cybersentinal.BannerBench
BannerBench: Benchmarking Vision Language Models for Multi-Ad Selection with Human Preferences
Dataset Summary
The BannerBench is designed to evaluate the ability of VLMs to identify the banner that best matches human preferences from a set of candidates.
Dataset Structure
The structure of the raw dataset is as follows:
{
"train": Dataset({
"features": [
'LPimage', 'image1', 'image2', 'image3', 'image4', 'image5'… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/BannerBench.captcha_100kCFV-Dataset
CFV Dataset: Fine-Grained Predictions of Car Orientation from Images
Description
Repository: CFV Dataset Repository
Paper: CFV Dataset Paper
Dataset Summary
We present the CFV Dataset for estimating the car's orientation from images. Our dataset was obtained by recording cars while walking around them and annotating the frames with the pitch angle value in a semi-automatic manner. All images have the license plates anonymized.
Data Instance
One… See the full description on the dataset page: https://huggingface.co/datasets/fort-cyber/CFV-Dataset.cyber_drug_dataset
Digital Forensic Investigation Scenario Dataset: Online Drug Trafficking
This dataset is a comprehensive collection of digital artifacts and investigative reports designed for forensic research and education. It simulates a sophisticated Online Drug Trafficking scenario, covering the entire investigation lifecycle from initial intelligence gathering to suspect arrest and financial analysis.
Dataset Structure
The dataset is indexed via a standardized 7-column metadata… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/cyber_drug_dataset.in-store-visual-localizationDataset for monocular visual relocalization on a COLMAP 3D reconstruction model.
This dataset was collected at EZOHUB Tokyo with the cooperation of SATUDORA HOLDINGS CO.,LTD.
Dataset structure
train : Images used to construct the COLMAP model.
model : The COLMAP model files.
test : Test images to run relocalization.
intrinsic.xml : Intrinsic camera parameters.
kovidore-v2-cybersecurity-mteb
KoVidore2CybersecurityRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions. This dataset, Cybersecurity, is a corpus of technical reports on cyber threat trends and security incident responses in Korea, intended for complex-document understanding tasks.
Task category
t2i
Domains
Social
Reference
https://github.com/whybe-choi/kovidore-data-generator
Source datasets:
whybe-choi/kovidore-v2-cybersecurity-beir… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-cybersecurity-mteb.Screen2Coord
Screen2Coord_denorm_extend Dataset
Screen2Coord is a dataset for training models that take a screenshot, screen dimensions, and a textual action description as input and output the coordinates of the target bounding box on the screen. This dataset is intended for image-text-to-text LLMs applied to user interface interactions.
Dataset Structure
New feature! Windows, MacOS, Linux-Ubuntu subsets!
Data Instances
A typical data instance in Screen2Coord… See the full description on the dataset page: https://huggingface.co/datasets/cybertruck32489/Screen2Coord.
