datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Lora_Cloud_Dataset_Test
VLM Safety Inspector (2B / 4B / 8B) Mac 端评测与闭环套件
VLM Safety Inspector (2B / 4B / 8B) Mac 端闭环评测包
本目录是一个完全自包含(Self-Contained)的独立评测套件,专门适配您的 Mac(Apple Silicon / MPS)目录布局。
本目录是一个完全独立、自包含(Self-Contained)的评测套件,专为在 Mac (Apple Silicon / MPS) 上运行。
一、Mac 端文件布局自动识别(针对您的 iild 结构)
一、核心架构与流水线
评测脚本已内置针对您 Mac 端 iild/ 目录结构的全自动路径解析器:
在本次评测中,整条上行与闭环流水线严格遵循您的设想:
上游双塔一致性(In-Domain Consistency):
输入给 Planner 和 Inspector 的 150 个任务安全规则,已在 PC 端由纯 Legacy… See the full description on the dataset page: https://huggingface.co/datasets/lvesucces/Lora_Cloud_Dataset_Test.OmniVideo-Test
OmniVideo-Test
Official repository for OmniVideo-Test, the human-verified test set introduced in our paper: "OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains".
This repository includes:
videos/: Raw video files.
test_505.jsonl: The test set containing 505 multiple-choice QA pairs, complete with task taxonomies, ground-truth answers, and options.
OmniVideo-Test serves as the evaluation companion to the OmniVideo-100K… See the full description on the dataset page: https://huggingface.co/datasets/MiG-NJU/OmniVideo-Test.RSCC-RSEdit-Test-Split
RSCC-RSEdit-Test-Split
This directory contains the test split for RSCC-RSEdit dataset.
Directory Structure
RSCC-RSEdit-Test-Split/
├── images/ # Original images (676 PNG files)
├── masks/ # Original grayscale masks (338 PNG files)
│ └── [mask files with pixel values 0,1,2,3,4]
├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files)
│ └── [same filenames as masks/, but in RGBA format with colors]
├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.TestingDataset
SciReC: Diagnostic Evaluation of Relational Reasoning in Multimodal Scientific Conversations with Adaptive Interaction
This dataset contains multimodal question-answering examples grounded in
textbook figures. Records in the figure-grounded configurations are filtered to
include only examples whose referenced image files are present in this release.
Configurations
visual: 13791 figure-grounded visual questions with resolved images.
knowledge: 13501 caption/text-grounded… See the full description on the dataset page: https://huggingface.co/datasets/Naga1289/TestingDataset.SpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset files:… See the full description on the dataset page: https://huggingface.co/datasets/manycore-research/SpatialGen-Testset.ParseBench_test
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/kramp/ParseBench_test.bopask-test
BOPASK-Test
Human-verified evaluation benchmark for the BOPASK spatial-reasoning VQA dataset.
Contains 934 question-answer pairs across two testsets:
core — BOPASK-Core: three BOP-Challenge families (HANDAL, HOPE, YCB-V).
lab — BOPASK-Lab : an in-the-wild set of "home / lab" scenes.
Contents at a glance
Split
Family
Records
RGB images
Depth maps
Masks
core
handal
251
43
41
138
core
hope
189
50
29
231
core
ycbv
248
48
48
153
lab
home
246
21
12 (⚠)
52… See the full description on the dataset page: https://huggingface.co/datasets/bhatvineet/bopask-test.zalo-ai-2025-public-test-data-v2
Zalo AI Challenge 2025 - RoadBuddy Public Test Data V2set
This dataset contains public test data v2 for the RoadBuddy – Understanding the Road through Dashcam AI challenge from Zalo AI Challenge 2025.
Dataset Description
The challenge aims to build a driving assistant capable of understanding video content from dashcams to quickly answer questions about traffic signs, signals, and driving instructions in Vietnam.
Dataset Structure
Files
frames/:… See the full description on the dataset page: https://huggingface.co/datasets/OpenHay/zalo-ai-2025-public-test-data-v2.cs2-action-inference-test
CS2 战术 Action 推理测试集
本测试集用于 WAN I2V 的战术动作定性测试。每个小类只保留 1 张真实比赛 POV 第一帧,以及两种英文文本条件;本版不提供 GT 视频。第一帧来源依据 parse-dem 的 events.csv、game_events.csv 或逐 tick 状态对齐到 opencs2_matches* 视频。
数据约定
共 45 个 case、9 个大类。
每个 case 只有一张 832x480 的 first_frame.png,作为 WAN I2V 条件图;不裁剪或复制 GT clip。首帧优先选择正常持械、水平视角、无遮挡且较开阔的画面。
prompt.txt 是完整英文 prompt,包含首帧可见环境、初始持械状态、画面保持要求和整段唯一动作变化。
chunk_prompts.json 固定包含 5 个英文 prompt,依次描述期望生成视频的 0-1、1-2、2-3、3-4、4-5 秒。
metadata.json… See the full description on the dataset page: https://huggingface.co/datasets/mikusama99/cs2-action-inference-test.SQUADDS_test_clone
THIS IS A CLONE AND IS NOT THE OFFICIAL SQUADDS DB.
SQuADDS_DB - a Superconducting Qubit And Device Design and Simulation Database
The SQuADDS (Superconducting Qubit And Device Design and Simulation) Database Project is an open-source resource aimed at advancing research in superconducting quantum device designs. It provides a robust workflow for generating and simulating superconducting quantum device designs, facilitating the accurate prediction of Hamiltonian… See the full description on the dataset page: https://huggingface.co/datasets/elizabethkunz/SQUADDS_test_clone.nutonic-sft-init-upload-test
NU-TONIC raw SFT init
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning LFM-VL (leap-finetune vlm_sft format). JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (default HF source: stochastic/random_streetview_images_pano_v0.0.2) via download_geoguessr_poi_imagery.py.
Optical: multispectral optical COGs from a public STAC catalog… See the full description on the dataset page: https://huggingface.co/datasets/Tonic/nutonic-sft-init-upload-test.misc-cfo-testing-cfo-analysis
misc-cfo-testing CFO analysis
This folder is the self-contained analysis output for:
C:\Users\15255\Desktop\Research\CSE237D\morty_data\misc-cfo-testing
The source data and the Weyl pipeline are read-only. All generated scripts,
fingerprints, statistics, logs, PNGs, and SVGs remain in this analysis folder.
Start with RESULTS.md.
Dataset and estimator parameters
Experiments: faraday (4 min), reboot (5 min), reboot-10m (10 min)
Receiver: pluto11
Input sample rate:… See the full description on the dataset page: https://huggingface.co/datasets/Morty0311/misc-cfo-testing-cfo-analysis.Rendered_512_32_TestAndroidControl_testSpatialGen-Testset
SpatialGen Testset
This repository contains the test set for SPATIALGEN: Layout-guided 3D Indoor Scene Generation, a novel multi-view multi-modal diffusion model for generating realistic and semantically consistent 3D indoor scenes.
Project page | Paper | Code
We provide a test set of 48 preprocessed point clouds and their corresponding GT layouts, multi-view images are cropped from the high-resolution panoramic images.
Folder Structure
Outlines of the dataset… See the full description on the dataset page: https://huggingface.co/datasets/BenjaminChai579/SpatialGen-Testset.landslide-prevention-italy-2024-blind-test
Pre-Landslide Risk Assessment 2024 Blind Test
Blind test dataset for evaluating VLM-based pre-landslide risk detection on 2024
Italian landslide events. This dataset contains 29 events from January–December 2024
with no ground truth labels — designed for evaluating model predictions against
future landslide occurrences.
Generated for evaluating fine-tuned models on LiquidAI/LFM2.5-VL-450M.
Dataset Summary
Total samples
29
Split
test (blind)
Image… See the full description on the dataset page: https://huggingface.co/datasets/Sciamlab/landslide-prevention-italy-2024-blind-test.fcos_test_pg_rcfnvlm_tsr_test_1
vlm_tsr_test_1
vlm_tsr_test의 1/4 파트. scene 그룹 50000~50004 포함.
전체 테스트셋은 4개 레포로 나뉘어 있습니다:
vlm_tsr_test_1
vlm_tsr_test_2
vlm_tsr_test_3
vlm_tsr_test_4
코드 및 전체 파이프라인: Lim-Sung-Jun/vlm_training_template
구조
각 샘플은 3개 파일 세트로 구성됩니다:
test/source/T01_C01/{id}.jpg # 테이블 이미지
test/source/T01_C01/{id}.json # 정답 HTML + 메타데이터
test/label/T01_C01/{id}.html # 렌더링용 GT HTML
평가 메트릭
메트릭
설명
TEDS
Tree-Edit Distance 기반 구조 유사도 (0~1)
TEDS-Structure
텍스트 제외 구조만… See the full description on the dataset page: https://huggingface.co/datasets/sungjun12/vlm_tsr_test_1.round_8_test1📊 Task Overview
Metric
Value
Total tasks
128
Total judged
125
Skipped
0
Generation failed
0
Render failed
0
Judge failed
3
🏁 Results
Outcome
Count
-------
----:
Wins
11
Draws
106
Losses
8
📈 Performance Metrics
Win Rate: 8.8%
Margin: 2.4%
Average Generation Time: 19.50 seconds📝 Notes
No tasks were skipped or failed during generation/rendering.
A small number of tasks (3) failed during judging.
The majority of evaluated tasks resulted in draws… See the full description on the dataset page: https://huggingface.co/datasets/yuri1996/round_8_test1.Driving_testnft-females-generated-cli-test-3-clone
nft females generated cli test 3
Automated NFT generation report
Number of characters: 10
Model: midorimae-characters-female
Guidance: 6.0
LORA scale: 0.95
Total time taken: 0h 0m 51s
Average time per prompt: 2.39 seconds
Average time per image: 2.74 seconds
Average time per image: 5.13 seconds
eagle360_test
EAGLE-360 Test Set
Project page: EAGLE-360
Paper: arXiv:2607.02479
EAGLE-360 is a benchmark for embodied active global-to-local exploration in 360-degree panoramic scenes. Given a panoramic image and a target-object query, the model is asked to predict the object's angular position as azimuth and elevation in degrees.
This release contains the public test split only. It includes panoramic images and a annotation file with ground-truth metadata.
Files… See the full description on the dataset page: https://huggingface.co/datasets/Sansjudge/eagle360_test.testing-text-image2laser-vibrations-test
Laser Vibrations
Dataset of laser speckle vibration recordings used to locate objects hidden inside a cardboard box.
A 10×10 grid of lasers shines on the side of a box containing an object; as loudspeakers excite the box,
the speckle patterns shift in proportion to the local surface vibration. Per-sample metadata is in
data/metadata.jsonl; full signal data and media files live in per-sample subdirectories.
Dataset Viewer Columns
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/eturok-weizmann/laser-vibrations-test.testresearch-article-template-editor-copy-test-datair-testdataset-test4
DiffusionDBXL
TODO
wonders_testing_sub_dirstest
