datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vulpentestbench-results
VulPentestBench -- agent results & trajectories
Benchmark results of autonomous LLM penetration-testing agents (launched
context-free) on
VulPentestBench: an agent-evaluation
harness that boots vulhub vulnerable targets in
isolated Docker networks, injects a fresh random canary token at a
vulnerability-reachable location per run, and scores provenance-verified milestones
(the token must come back through a tool response of a target-aimed action before a
flag submission counts --… See the full description on the dataset page: https://huggingface.co/datasets/C0rk1/vulpentestbench-results.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.benchhub_plus_results_evaluated
BenchHub Plus Results (Evaluated)
LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores.
Folder Structure
├── vllm_inference_results_en/ # English benchmark results (19 models)
│ ├── {model_name}_{date}.jsonl
│ └── ...
└── vllm_inference_results_ko/ # Korean benchmark results (16 models)
├── {model_name}_{date}.jsonl
└── ...
Column Description
Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.medfailbench-v02-results
MedFailBench v0.2 Model Response Artifacts
This dataset preserves full model responses and automated rule scores generated during MedFailBench development. Each row links a prompt, a model identifier, the response text, available scores, a final label, and its source artifact.
Public data scope
The current file contains 181 rows with 11 distinct values in model_name. These values are run identifiers rather than a normalized model registry. They must not be… See the full description on the dataset page: https://huggingface.co/datasets/goktugozkanmd/medfailbench-v02-results.codebench-results
CodeBench Results
Experimental results from the CodeBench evaluation framework.
Two dataset configs are included.
Configs
h4_security — H4 Security-Adjusted Reliability Experiment
240 rollouts across 3 agents × 10 tasks × 8 rollouts.
Tasks split into 5 standard algorithmic tasks and 5 security-sensitive tasks
(eval, exec, shell, yaml, dynamic import).
Column
Description
task_id
Task identifier (algo-* or sec-*)
agent_name
anote-code… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/codebench-results.math500-bon-weighted-results
MATH-500 Best-of-N Weighted Selection Results
Dataset Description
This dataset contains the results of evaluating Best-of-N weighted selection on a subset of the MATH-500 benchmark. It was created as part of a HuggingFace internship exercise exploring how test-time compute scaling with reward models can improve LLM performance on math problems.
How It Was Constructed
1. Problem Selection
Started from the HuggingFaceH4/MATH-500 dataset (500 problems)… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/math500-bon-weighted-results.imo-proofbench-mr-v3sft-results
IMO ProofBench — Meta-Reasoning v3-SFT eval results
GPT-5.4-judged results of a meta-reasoning SFT model (midtrain v3) on
HerrHruby/imo-proofbench-all-vf
(60 modified-olympiad problems). For each problem the model is sampled 3×
(180 candidates total).
Model & inference
Model: Qwen3.5-9B, midtrain-v3 SFT (checkpoint global_step_4500).
Scaffold: a single model runs the full meta-reasoning loop — MR (propose
exploration directions) → E (execute each direction… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/imo-proofbench-mr-v3sft-results.concept-injection-results
Concept Injection Results
Experimental data and analysis logs from the research "Replicating Introspection on Injected Content in Open-Source Language Models"
Overview
This project enables researchers to inject concept vectors directly into a model's hidden layers during inference, allowing investigation of whether language models can detect and report on artificially induced "thoughts."
This dataset contains the raw responses and processed evaluations produced… See the full description on the dataset page: https://huggingface.co/datasets/vazirani/concept-injection-results.results-multi-temp-seed
Benchpress Evaluation Results
This dataset contains normalized outputs from lm-evaluation-harness
runs for the Benchpress project. It is intended for analysis of
post-training recipe behavior across models, benchmarks, temperatures,
and random seeds.
Coverage
Runs: 110
Model/stage recipes: olmo2-13b:dpo, olmo2-13b:rlvr, olmo2-13b:sft, olmo2-7b:dpo, olmo2-7b:rlvr, olmo2-7b:sft, olmo3-7b:dpo, olmo3-7b:rlvr, olmo3-7b:sft, smollm3-3b:apo, smollm3-3b:sft
Temperatures:… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/results-multi-temp-seed.nova-industry-benchmark-results
Nova Industry Benchmark — Results
Model run outputs for the questions in
SparkSupernova/nova-industry-benchmark.
Why results live in their own repository
Results were previously stored as extra splits of the question dataset. Because each
model version wrote a different set of columns, adding the v5 run gave that dataset two
splits with incompatible schemas, and load_dataset failed for every consumer — including
the usage example on the model card.
Questions are a… See the full description on the dataset page: https://huggingface.co/datasets/SparkSupernova/nova-industry-benchmark-results.sinergi-model-evaluation-results
Sinergi model evaluation outputs
Model answers and generation telemetry used by the Sinergi Table 8-aligned evaluation notebook. The repository contains eight configurations so every system can be loaded independently with the Hugging Face datasets library.
from datasets import load_dataset
data = load_dataset(
"Legal-verse/sinergi-model-evaluation-results",
"qwen-sft-rl-rag",
split="test",
)
Configurations
Configuration
System
Rows
Source file… See the full description on the dataset page: https://huggingface.co/datasets/Legal-verse/sinergi-model-evaluation-results.PREPAIR-reproduction-results
PREPAIR Reproduction Results
This dataset contains the reproduction results of the PREPAIR method.
Source Project: PREPAIR
Target File: results_LLMBar_all_methods.csv
