datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
results_public
Dataset Card for "resultspublic"
More Information needed
video-benchmark-resultsbenchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.depth-benchmark-resultsbokeh-lf-benchmark-resultsgaia-benchmark-resultsopen-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.molt-benchmark-results
Molt · elastic on-device inference measurements
Everything measured while building Molt, a
runtime that moves a running generation onto a smaller model between two
tokens, carrying the KV cache across, so an on-device LLM under memory pressure
is neither reclaimed by the OS nor restarted from the prompt.
Published so the claims can be checked rather than taken on trust. The figures in
the repo README and the results page are generated from these files; nothing is
transcribed by… See the full description on the dataset page: https://huggingface.co/datasets/NagaYu/molt-benchmark-results.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.optimization-os-benchmark-results
Optimization OS — Benchmark Results
Pre-computed benchmark runs comparing baseline, exact, scalable, and robust methods.
Runs: 144Methods: 4 per problem type (24 total)
gpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.topic-benchmark-resultsagents-benchmark-eval-resultsbenchmark-demo-results
Benchmark Demo Results
This public dataset stores demo leaderboard results for the Hugging Face leaderboard MVP.
The main table is leaderboard.csv. In the MVP, the metrics are demo values generated from an example submission. Later this file will be updated by the Gradio Space after real submissions are scored.
imp-act-benchmark-results
IMP-act: Benchmarking MARL for Infrastructure Management Planning at Scale with JAX. Results and Model Checkpoints.
Overview
This directory contains all the model checkpoints stored during training and inference outputs. You can use these files to reproduce training curves, evaluate policies, or kick-start your own experiments.
Repository
The guidelines and instructions are available on the IMP-act GitHub repository.
Licenses
This dataset is released… See the full description on the dataset page: https://huggingface.co/datasets/AI-for-Infrastructure-Management/imp-act-benchmark-results.ocr-benchmark-results
OCR Bench Results: ocr-benchmark-combined
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
clearocr.com/clearocr-api
1837
1801–1882
488
56
1
90%
2
deepseek-ai/DeepSeek-OCR
4B
1639
1611–1671
380
164
1
70%
3
lightonai/LightOnOCR-2-1B
1B
1448
1421–1477
242
302
0
44%
4
rednote-hilab/dots.ocr
1.7B
1384
1353–1413
193
351
0
35%… See the full description on the dataset page: https://huggingface.co/datasets/j4xfu2mm/ocr-benchmark-results.nova-industry-benchmark-results
Nova Industry Benchmark — Results
Model run outputs for the questions in
SparkSupernova/nova-industry-benchmark.
Why results live in their own repository
Results were previously stored as extra splits of the question dataset. Because each
model version wrote a different set of columns, adding the v5 run gave that dataset two
splits with incompatible schemas, and load_dataset failed for every consumer — including
the usage example on the model card.
Questions are a… See the full description on the dataset page: https://huggingface.co/datasets/SparkSupernova/nova-industry-benchmark-results.clearocr-benchmark-fiszki91-v1-results
OCR Bench Results: clearocr-benchmark-fiszki91-v1
VLM-as-judge pairwise evaluation of OCR models. Rankings depend on document type — there is no single best OCR model.
Leaderboard
Rank
Model
Params
ELO
95% CI
Wins
Losses
Ties
Win%
1
clearocr.com/clearocr-api
1777
1734–1824
316
47
0
87%
2
lightonai/LightOnOCR-2-1B
1B
1567
1537–1603
222
141
0
61%
3
deepseek-ai/DeepSeek-OCR
4B
1412
1378–1444
137
225
0
38%
4
FireRedTeam/FireRed-OCR
2.1B
1380
1345–1411
120… See the full description on the dataset page: https://huggingface.co/datasets/Lukaszl/clearocr-benchmark-fiszki91-v1-results.raw_benchmark_results_prunaraw_benchmark_results_hqqQwen3-4B-Instruct-cost-benchmark-resultsresume-gen-benchmark-resultsgpt4o-cosa-benchmark-resultspass-at-k-benchmark-results
ISO-Bench Pass@k GPU Benchmark Results
Agent-generated optimization patches benchmarked on real GPU hardware (NVIDIA H100 80GB).
Scope: Lossfunk/ISO-Bench — 54 tasks (39 vLLM + 15 SGLang)
Patches from: Inferencebench/pass-at-k-samples
Summary
vLLM
SGLang
Total
Tasks benchmarked
29/39
10/15
39/54
Successful benchmarks444
150
594
Agent
vLLM
SGLang
Total
Claude Code (Sonnet 4.5)
230
76
306
Codex CLI (GPT-5)
214
74
288… See the full description on the dataset page: https://huggingface.co/datasets/Inferencebench/pass-at-k-benchmark-results.gpt-3.5-turbo-cosa-benchmark-resultssage-benchmark-resultsresume-gen-benchmark-results-optimisedbenchmark-resultsQwen3-8B-Instruct-cost-benchmark-resultsbenchmark-results-correction-prevention
