datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RSEdit-Benchmark-Results
RSCC-RSEdit-Test-Split
Project Page | Paper | GitHub
This repository contains the test split for RSEdit, a unified framework designed for text-guided image editing in the remote sensing (RS) domain.
Description
RSEdit adapts pre-trained text-to-image diffusion models (including U-Net and DiT architectures) into instruction-following editors for Earth observation imagery. This dataset consists of bi-temporal remote sensing image pairs and corresponding textual… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSEdit-Benchmark-Results.results_public
Dataset Card for "resultspublic"
More Information needed
visual-reasoning-benchmark-results
Visual Reasoning Benchmark Suite v3.3 · 2005 Tasks · 12 Tracks Equal Weight
本版本以用户最新上传的 visual_reasoning_benchmark_suite_v3_修改 为唯一基础版本,不回退、不覆盖用户已经重绘或修改过的既有数据。完整性比对结果:原基础包中 3283 个既有数据文件全部保持字节级不变。
在此基础上新增并整合:
Nonogram(数织)150 题:45 Easy / 60 Medium / 45 Hard;
Tangram(七巧板)150 题:45 Easy / 60 Medium / 45 Hard;
两个任务的一键生成器、统一生成入口、统一评估入口、雷达图和排行榜支持。
最终总规模:2005 题,12 个 Track。
任务与数量
Task
Count
figure_completion
394
spatial_generation
56
maze_beginner
64… See the full description on the dataset page: https://huggingface.co/datasets/songyiren/visual-reasoning-benchmark-results.video-benchmark-resultstokenizer_benchmark_resultsmobileforge-benchmark-results
MobileForge Benchmark Results
This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization.
It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended for result verification, log inspection, and mapping the public model checkpoints to the exact benchmark artifacts reported in… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-benchmark-results.negation-benchmark-results
Negation Benchmark Results — reviewed v1
All seven existing model configurations now contain only the question-reviewed benchmark pairs.
incorrect_option is removed. Answer judgments are annotations only: no model output is excluded
because it is incorrect, unchanged, or of an unexpected type.
Gemma 3 4B/12B/27B and Qwen 3.5 9B each contain 105,048 pairs (41,957 train, 10,348 validation,
52,743 test). OLMo 3 7B, Llama 3.1 8B and Qwen 3 8B each contain 58,084 text-only pairs
(23… See the full description on the dataset page: https://huggingface.co/datasets/Jongwondd/negation-benchmark-results.assistive-ocr-benchmark-results
Assistive OCR — Benchmark Results
Real, reproducible benchmark results for the assistive OCR wearable module (offline, multilingual — English, Bengali+English, Hindi+English). This repository is self-contained: it holds the results, the ground-truth manifest, and the 98 real images they were computed from, so it can be run and demoed directly with no other dataset needed.
What's in this repository
File
What it is
manual100_final.csv
The 99-row… See the full description on the dataset page: https://huggingface.co/datasets/bhumika-tewari-282006/assistive-ocr-benchmark-results.benchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.spec_gemma_4_benchmark_results
Gemma 4 31B speculative-decoding benchmark results
This repository contains reproducible benchmark artifacts for google/gemma-4-31B-it
on one NVIDIA H100 NVL, comparing baseline decoding, native Gemma MTP, and DFlash.
Raw prompts, generated responses, per-request timings, server logs, commands, metadata,
and acceptance records are retained under results/.
H100 quick start
Use an H100 with at least 80 GB VRAM. Target weights are BF16; the published matrix uses
FP8… See the full description on the dataset page: https://huggingface.co/datasets/evnisv/spec_gemma_4_benchmark_results.pii-masking-benchmark-resultscodeforces-benchmark-resultsbenchmark_resultsMADD_Benchmark_and_results
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: Ministry of Economic
Development of the Russian Federation (IGK
000000C313925P4C0002), agreement No139-15-
2025-010
Shared by [optional]: [More Information Needed]
Language(s) (NLP): English
License: [More Information… See the full description on the dataset page: https://huggingface.co/datasets/ITMO-NSS/MADD_Benchmark_and_results.depth-benchmark-resultsmobileforge-benchmark-results
MobileForge Benchmark Results
Anonymous project: https://mobileforge-anonymous.github.io/Anonymous code: https://github.com/mobileforge-anonymous/MobileForge
This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization.
It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended… See the full description on the dataset page: https://huggingface.co/datasets/mobileforge-anonymous/mobileforge-benchmark-results.bokeh-lf-benchmark-resultsgaia-benchmark-resultsscalelab-benchmark-results
ScaleLab Benchmark Results
Reproducible benchmark results comparing optimization algorithms across industrial process instances.
Benchmark Schema
Each result record contains:
Field
Description
instance_id
Instance identifier
algorithm
Optimization algorithm used
solver
Surrogate/solver backend
configuration
Algorithm configuration
seed
Random seed
runtime_sec
Total runtime
time_to_first_solution
Time to reach target quality
n_experiments… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/scalelab-benchmark-results.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.vllm-hust-benchmark-resultsbenchmark-evaluation-resultsWolframRavenwolfs_benchmark_results
Results of WolframRavenwolfs(@wolfram on huggingface) tests in csv form.
1st Score = Correct answers to multiple choice questions (after being given curriculum information)
2nd Score = Correct answers to multiple choice questions (without being given curriculum information beforehand)
OK = Followed instructions to acknowledge all data input with just "OK" consistently
+/- = Followed instructions to answer with just a single letter or more than just a single letter
speculative-decoding-benchmark-resultsmolt-benchmark-results
Molt · elastic on-device inference measurements
Everything measured while building Molt, a
runtime that moves a running generation onto a smaller model between two
tokens, carrying the KV cache across, so an on-device LLM under memory pressure
is neither reclaimed by the OS nor restarted from the prompt.
Published so the claims can be checked rather than taken on trust. The figures in
the repo README and the results page are generated from these files; nothing is
transcribed by… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/molt-benchmark-results.swa-benchmark-resultsturkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.optimization-os-benchmark-results
Optimization OS — Benchmark Results
Pre-computed benchmark runs comparing baseline, exact, scalable, and robust methods.
Runs: 144Methods: 4 per problem type (24 total)
benchmark_results_merge_eventgpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.
