datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
visual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.assistive-ocr-benchmark-results
Assistive OCR — Benchmark Results
Real, reproducible benchmark results for the assistive OCR wearable module (offline, multilingual — English, Bengali+English, Hindi+English). This repository is self-contained: it holds the results, the ground-truth manifest, and the 98 real images they were computed from, so it can be run and demoed directly with no other dataset needed.
What's in this repository
File
What it is
manual100_final.csv
The 99-row… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/assistive-ocr-benchmark-results.WolframRavenwolfs_benchmark_results
Results of WolframRavenwolfs(@wolfram on huggingface) tests in csv form.
1st Score = Correct answers to multiple choice questions (after being given curriculum information)
2nd Score = Correct answers to multiple choice questions (without being given curriculum information beforehand)
OK = Followed instructions to acknowledge all data input with just "OK" consistently
+/- = Followed instructions to answer with just a single letter or more than just a single letter
gpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.image-editing-benchmark-results
Image editing benchmark results
Generated results for the PIE-Bench main set (700 cases) and ICE-Bench Reference Editing set (518 cases). The archives contain edited images only. Each archive keeps its original workspace paths under evaluation/.
Archive
Contents
image_editing_baselines_results.zip
Five baselines on PIE-Bench and five baselines on ICE-Bench Reference Editing
image_editing_ours_results.zip
Seven model/checkpoint variants on each benchmark
Each… See the full description on the dataset page: https://huggingface.co/datasets/yanlinli/image-editing-benchmark-results.UniCoT-benchmark-resultsCompleteMe_Benchmark_Results3d-editing-benchmark-results
3D editing results: training-view uniform8 v3
Results from four checkpoints on 16 edits across 8 scenes. Each ZIP contains the edited 3DGS renders for 8 fixed training cameras per case (128 PNGs per checkpoint), case metadata, and CLIP/PH-Loss metrics. The input uses the final train_pose_casewise_wide_26_24fps_v3 protocol: 26 training-camera video frames, vanilla 3DGS reconstruction from sparse SfM for 30,000 iterations, and score views [0,4,7,11,14,18,21,25]. No… See the full description on the dataset page: https://huggingface.co/datasets/yanlinli/3d-editing-benchmark-results.
