datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SWE-Bench-Verified-O1-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results
SWE-Bench Verified O1 Dataset
Executive Summary
This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances.
Overview
This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.dementor-complete-experiment-results
Dementor complete experiment results
Audited outputs for the configuration-defined Dementor completion campaign.
Audited scope
Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls).
Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and
primary LLM-judge scores, plus 48 target-reference response sets.
Activation steering: 29 models, seven benchmarks, and two operators (original and fpall),
totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.eval-IFBench-results
IFBench Evaluation Results
This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following.
Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos:
eval-IFBench-results - Model evaluation outputs (this repo)
eval-IFBench-prompts - Test prompts/questions (if separated)
Dataset Structure
Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.mobileforge-benchmark-results
MobileForge Benchmark Results
This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization.
It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended for result verification, log inspection, and mapping the public model checkpoints to the exact benchmark artifacts reported in… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-benchmark-results.corral_lfm_binomial_results
Corral – LFM Binomial IRT Results
Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the fitted parameters of a binomial Item Response Theory (IRT) model estimated from agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results.vulpentestbench-results
VulPentestBench -- agent results & trajectories
Benchmark results of autonomous LLM penetration-testing agents (launched
context-free) on
VulPentestBench: an agent-evaluation
harness that boots vulhub vulnerable targets in
isolated Docker networks, injects a fresh random canary token at a
vulnerability-reachable location per run, and scores provenance-verified milestones
(the token must come back through a tool response of a target-aimed action before a
flag submission counts --… See the full description on the dataset page: https://huggingface.co/datasets/C0rk1/vulpentestbench-results.qwopus-dflash-swe20-runtime-results
Qwopus / DFlash SWE20 Runtime Results
Local RTX 3090 Ti benchmark artifacts for 20 long SWE-bench Lite prompts. The run compares Qwopus 3.6 GGUF variants, llama.cpp MTP speculative decoding, QuinsZouls, and DFlash DDTree configurations at 64K context with q8/q8 KV unless noted.
The quality score is a reproducible proxy rubric over gold-patch signals, not official SWE-bench pass/fail. It checks touched-file matches, identifier overlap, patch-like concreteness, test signal, length… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwopus-dflash-swe20-runtime-results.results
TrueVisLies – Results
This dataset contains all raw outputs, extracted fields, semantic similarity scores, and UMAP projections produced in the paper:
True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies
The paper evaluates 16 LLMs, 15 open-weight vision-language models (VLMs), and GPT-5.4 on their ability to (RQ0) detect misleading data visualizations, (RQ1) identify the visualization rhetoric techniques, and… See the full description on the dataset page: https://huggingface.co/datasets/truevislies/results.cad-bench-results
Parametric CAD Bench — Claude Fable 5 results
Run artifacts for claude-fable-5 driven by the claude-code agent on
gnucleus-ai/cad-bench@v1
(the gNucleus Parametric CAD Bench), graded by the project's own
freecad-validator.
This is an independent third-party submission. The layout follows the
cad-bench-submission
contract exactly, mirroring
gnucleus-ai/cad-gen-freecad-bench.
Headline
Metric
Value
Combined score (mean over 100 tasks)
0.7577
Geometry… See the full description on the dataset page: https://huggingface.co/datasets/EricSpencer00/cad-bench-results.phase_tree_results
PHASE-Tree Evaluation Results
Full evaluation outputs for the PHASE-Tree paper
(Psychology-grounded Hierarchical Attribute-Structured Evolving Tree),
covering 8 character-dialogue datasets, 4 experimental paradigms, and
2 evaluation splits (random test + OOD test).
Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose.
🔗 Resources:
📄 Paper: arXiv:2608.06975
📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.inkslop-results
InkSlop Benchmark Results
Model evaluation results for the InkSlop Benchmark - a vibe-coded benchmark for spatial reasoning with digital ink.
Collection: InkSlop Benchmark
Contents
This dataset contains inference results and evaluation metrics for multiple VLMs across all InkSlop tasks:
overlap_easy / overlap_hard - Overlapped handwriting recognition
autocomplete_easy / autocomplete_hard - Handwriting autocompletion
derender_easy / derender_hard - Ink derendering (image… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-results.fucc-boi-bench-results-v03
fucc boi bench results v0.8.0
This is the public results release for fucc boi bench.
41 models
96 prompts per model
3936 scored answers
one combined leaderboard
The site defaults to General-purpose models, with Wildcards as a secondary view
and All models as an optional combined view.
Access route and billing are metadata; they do not create separate rankings.
Files:
responses.jsonl: sanitized model outputs and run metadata.
grades.jsonl: parsed grades, fuccboi scores… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v03.rubric_rl_results
Rubric RL Evaluation Results
Evaluation data for rubric-based reward modeling experiments. Contains generated rubrics from multiple rubric generators and pairwise scoring results comparing rl-research/DR-Tulu-8B (RL, step_4000) vs rl-research/DR-Tulu-SFT-8B.
Data Structure
rubrics/ — Generated evaluation rubrics
Each JSONL file contains per-question rubrics with fields: prompt_id, question, generated_rubric, generated_rubric_raw, rubric_style, rubric_model.… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/rubric_rl_results.mobileforge-benchmark-results
MobileForge Benchmark Results
Anonymous project: https://mobileforge-anonymous.github.io/Anonymous code: https://github.com/mobileforge-anonymous/MobileForge
This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization.
It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended… See the full description on the dataset page: https://huggingface.co/datasets/mobileforge-anonymous/mobileforge-benchmark-results.idt5-v4-results-final-lora-s123-20260912T013040606815Z
final-lora-s123-20260912T013040606815Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 94.7565543071161,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.oral-args-data-and-results
Oral Arguments Arena
Data repository for AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments (Zhang, Nadeem, Zheng, Stammbach, Henderson, 2026). Refer to the paper for background on the evaluation framework, experimental design, and findings.
Repository Structure
oral-args-arena-annotations/
├── transcript_data/ # SCOTUS oral argument transcripts and case briefs
├── automated_metrics/ # LLM classifier outputs (SQLite… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/oral-args-data-and-results.connections-rl-results
connections-rl: raw evaluation artifacts
Per-puzzle records, bootstrap summaries and analysis outputs backing
connections-rl, a two-scale (Qwen2.5-1.5B / 7B), three-seed study of
what verifiable-reward RL actually transfers.
This is an artifact bundle for auditing published numbers, not a loadable
training dataset, so the dataset viewer is disabled.
Read this before using the numbers
Two conventions in these files are easy to misread. Both have bitten this
project… See the full description on the dataset page: https://huggingface.co/datasets/jacksonlukas/connections-rl-results.idt5-v4-results-final-lora-s2026-20260912T034640190015Z
final-lora-s2026-20260912T034640190015Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 95.88014981273409,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s2026-20260912T034640190015Z.sophia-training-results
Sophia Training Results
All training history results from the sophia-agi repository.
Contents
824 result files across all Sophia four-theme training experiments:
Prosoche — interactive decision-making, behavioral batteries, focus experiments
RunPod training — LoRA training eval ladders (Qwen2.5-3B, multiple seeds/epochs)
Continual learning — sequential curriculum, QA judged results
World model — on-device, adaptive, real-corpus experiments
Coherence reframe — NLI… See the full description on the dataset page: https://huggingface.co/datasets/tomyimkc/sophia-training-results.open-models-benchmark-results
⚡ Local LLM Evaluation Leaderboard
Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs.
💻 Hardware & System Specifications
All evaluations are executed under standardized local cluster environments:
Specification
Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.benchhub_plus_results_evaluated
BenchHub Plus Results (Evaluated)
LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores.
Folder Structure
├── vllm_inference_results_en/ # English benchmark results (19 models)
│ ├── {model_name}_{date}.jsonl
│ └── ...
└── vllm_inference_results_ko/ # Korean benchmark results (16 models)
├── {model_name}_{date}.jsonl
└── ...
Column Description
Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.idt5-v4-results-final-fft-s2026-20260911T063507183162Z
final-fft-s2026-20260911T063507183162Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 88.01498127340824,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s2026-20260911T063507183162Z.catbench-results
CatBench results
Model outputs for CatBench, a small
benchmark that asks a model to draw a cute kitten two ways and looks at what comes
back. Produced by the /catbench command in
blockquant.
Upstream publishes its own results at
Katehuuh.github.io/demos/CatBench/assets.
This dataset holds runs for models that are not in that set. /catbench checks both
and only rents a pod when neither has the model, so the two do not duplicate
each other.
The prompts
Verbatim… See the full description on the dataset page: https://huggingface.co/datasets/Honkware/catbench-results.idt5-v4-results-final-lora-s42-20260912T063343815032Z
final-lora-s42-20260912T063343815032Z
Run artifacts and per-item predictions.
Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables.
See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity.
Metrics
{
"n": 267,
"rule_version": "structural-proxy-v0.4-grounding-separated",
"parse_success_pct": 92.88389513108615,
"bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s42-20260912T063343815032Z.fucc-boi-bench-results-v02
fucc boi bench results v0.2
This is the public results release for fucc boi bench.
13 models
96 prompts per model
1248 scored answers
one combined leaderboard
Files:
responses.jsonl: sanitized model outputs and run metadata.
grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales.
leaderboard.json: the ranked summary.
summary.json: the interactive site's data and case examples.
benchmark_report.md: the full short report.
benchmark_card.md: the compact… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v02.inplace-ttt-results
In-Place TTT - Experimental Results
This dataset contains experimental results and training logs for the In-Place Test-Time Training project.
📦 Dataset Contents
1. Experimental Results (results_20260904.tar.gz)
Size: 191MB (compressed from 1.7GB)
Files: 1,999 files
Contents:
Language model evaluation results
RULER benchmark outputs
ProLong training metrics
Various test configurations
2. WandB Training Logs (wandb_20260904.tar.gz)… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/inplace-ttt-results.turkish-seo-reasoning-benchmark-results
Turkish SEO Reasoning Benchmark Results
Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir.
Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora
Sonuç
Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti.
Mutlak artış: +10,28 puan
Göreli artış: %85,97
Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.HumanAgencyBench_Evaluation_Results
HumanAgencyBench evaluation results
Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants
Code: https://github.com/BenSturgeon/HumanAgencyBench/
Dataset Description
This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.compartmentalized-harm-v1-results
Four-model character-training results
This package contains 29,952 paper-facing response records and their matching visible requests and blind outcome judgments. It covers Qwen3-8B, Qwen3-32B, Mistral Small 3.2 24B, and Gemma 4 31B at L1 and L2. Each model is evaluated with an identical rule-plus-hostile system message in the base and trained arms, and again with no experimental system prompt.
Each lane has requests.jsonl, responses.jsonl, and judgments.jsonl. The public… See the full description on the dataset page: https://huggingface.co/datasets/daios/compartmentalized-harm-v1-results.
