benchmark-results
RSEdit-Benchmark-Results
RSCC-RSEdit-Test-Split
Project Page | Paper | GitHub
This repository contains the test split for RSEdit, a unified framework designed for text-guided image editing in the remote sensing (RS) domain.
Description
RSEdit adapts pre-trained text-to-image diffusion models (including U-Net and DiT architectures) into instruction-following editors for Earth observation imagery. This dataset consists of bi-temporal remote sensing image pairs and corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSEdit-Benchmark-Results.results_public
Dataset Card for "resultspublic"
More Information needed
visual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.video-benchmark-resultstokenizer_benchmark_resultsmobileforge-benchmark-results
MobileForge Benchmark Results
This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization.
It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended for result verification, log inspection, and mapping the public model checkpoints to the exact benchmark artifacts reported in… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-benchmark-results.
