CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlexCuadron /SWE-Bench-Verified-O1-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities on the SWE-Bench Verified dataset, achieving a 28.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code generation through enhanced action-based reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-reasoning-high-results.textquestion-answeringn<1K7 likes6.1k downloads2y agoHugging Face02AlexCuadron /SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results SWE-Bench Verified O1 Dataset Executive Summary This repository contains verified reasoning traces from the O1 model evaluating software engineering tasks. Using OpenHands + CodeAct v2.2, we tested O1's bug-fixing capabilities using their native tool calling capabilities on the SWE-Bench Verified dataset, achieving a 45.8% success rate across 500 test instances. Overview This dataset was generated using the CodeAct framework, which aims to improve code… See the full description on the dataset page: https://huggingface.co/datasets/AlexCuadron/SWE-Bench-Verified-O1-native-tool-calling-reasoning-high-results.textquestion-answeringn<1K4 likes1.9k downloads2y agoHugging Face03dementor-research /dementor-complete-experiment-results Dementor complete experiment results Audited outputs for the configuration-defined Dementor completion campaign. Audited scope Behavioral imitation adapters: 1,104 total (528 SFT, 528 DPO, 48 self-SFT controls). Behavioral-fidelity evaluation: 1,104 adapters on 200 held-out prompts, with embedding and primary LLM-judge scores, plus 48 target-reference response sets. Activation steering: 29 models, seven benchmarks, and two operators (original and fpall), totaling… See the full description on the dataset page: https://huggingface.co/datasets/dementor-research/dementor-complete-experiment-results.text-generation0 likes1.7k downloads1mo agoHugging Face04shisa-ai /eval-IFBench-results IFBench Evaluation Results This dataset contains evaluation results for various language models on IFBench, a challenging benchmark for precise instruction following. Naming Convention: This repo follows the eval-{EVAL}-{type} schema for organizing evaluation datasets. Related repos: eval-IFBench-results - Model evaluation outputs (this repo) eval-IFBench-prompts - Test prompts/questions (if separated) Dataset Structure Results are organized by model name:… See the full description on the dataset page: https://huggingface.co/datasets/shisa-ai/eval-IFBench-results.texttext-generation10K<n<100K0 likes483 downloads3mo agoHugging Face05lgy0404 /mobileforge-benchmark-results MobileForge Benchmark Results This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization. It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended for result verification, log inspection, and mapping the public model checkpoints to the exact benchmark artifacts reported in… See the full description on the dataset page: https://huggingface.co/datasets/lgy0404/mobileforge-benchmark-results.reinforcement-learning1 likes409 downloads3mo agoHugging Face06jablonkagroup /corral_lfm_binomial_results Corral – LFM Binomial IRT Results Fitted parameters of a binomial Item Response Theory model quantifying the contributions of model and scaffold to agent performance across all Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the fitted parameters of a binomial Item Response Theory (IRT) model estimated from agent evaluation… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_lfm_binomial_results.text-generation1K<n<10K0 likes349 downloads5mo agoHugging Face07C0rk1 /vulpentestbench-results VulPentestBench -- agent results & trajectories Benchmark results of autonomous LLM penetration-testing agents (launched context-free) on VulPentestBench: an agent-evaluation harness that boots vulhub vulnerable targets in isolated Docker networks, injects a fresh random canary token at a vulnerability-reachable location per run, and scores provenance-verified milestones (the token must come back through a tool response of a target-aimed action before a flag submission counts --… See the full description on the dataset page: https://huggingface.co/datasets/C0rk1/vulpentestbench-results.tabulartext-generation10K<n<100K1 likes285 downloads12d agoHugging Face08jakeatx /qwopus-dflash-swe20-runtime-results Qwopus / DFlash SWE20 Runtime Results Local RTX 3090 Ti benchmark artifacts for 20 long SWE-bench Lite prompts. The run compares Qwopus 3.6 GGUF variants, llama.cpp MTP speculative decoding, QuinsZouls, and DFlash DDTree configurations at 64K context with q8/q8 KV unless noted. The quality score is a reproducible proxy rubric over gold-patch signals, not official SWE-bench pass/fail. It checks touched-file matches, identifier overlap, patch-like concreteness, test signal, length… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwopus-dflash-swe20-runtime-results.texttext-generation10K<n<100K0 likes262 downloads4mo agoHugging Face09truevislies /results TrueVisLies – Results This dataset contains all raw outputs, extracted fields, semantic similarity scores, and UMAP projections produced in the paper: True (VIS) Lies: Analyzing How Generative AI Recognizes Intentionality, Rhetoric, and Misleadingness in Visualization Lies The paper evaluates 16 LLMs, 15 open-weight vision-language models (VLMs), and GPT-5.4 on their ability to (RQ0) detect misleading data visualizations, (RQ1) identify the visualization rhetoric techniques, and… See the full description on the dataset page: https://huggingface.co/datasets/truevislies/results.textimage-classification10M<n<100M0 likes249 downloads5mo agoHugging Face10EricSpencer00 /cad-bench-results Parametric CAD Bench — Claude Fable 5 results Run artifacts for claude-fable-5 driven by the claude-code agent on gnucleus-ai/cad-bench@v1 (the gNucleus Parametric CAD Bench), graded by the project's own freecad-validator. This is an independent third-party submission. The layout follows the cad-bench-submission contract exactly, mirroring gnucleus-ai/cad-gen-freecad-bench. Headline Metric Value Combined score (mean over 100 tasks) 0.7577 Geometry… See the full description on the dataset page: https://huggingface.co/datasets/EricSpencer00/cad-bench-results.text-generationn<1K1 likes249 downloads2mo agoHugging Face11Mathematics-Yang /phase_tree_results PHASE-Tree Evaluation Results Full evaluation outputs for the PHASE-Tree paper (Psychology-grounded Hierarchical Attribute-Structured Evolving Tree), covering 8 character-dialogue datasets, 4 experimental paradigms, and 2 evaluation splits (random test + OOD test). Please cite this work if you use these results for analysis, comparison, reproduction, or any other research purpose. 🔗 Resources: 📄 Paper: arXiv:2608.06975 📦 GitHub Repository: MemTensor/PHASE-Tree (code… See the full description on the dataset page: https://huggingface.co/datasets/Mathematics-Yang/phase_tree_results.texttext-generation10K<n<100K1 likes228 downloads1mo agoHugging Face12amaksay /inkslop-results InkSlop Benchmark Results Model evaluation results for the InkSlop Benchmark - a vibe-coded benchmark for spatial reasoning with digital ink. Collection: InkSlop Benchmark Contents This dataset contains inference results and evaluation metrics for multiple VLMs across all InkSlop tasks: overlap_easy / overlap_hard - Overlapped handwriting recognition autocomplete_easy / autocomplete_hard - Handwriting autocompletion derender_easy / derender_hard - Ink derendering (image… See the full description on the dataset page: https://huggingface.co/datasets/amaksay/inkslop-results.imagetext-generationn<1K0 likes197 downloads8mo agoHugging Face13patrickleenyc /fucc-boi-bench-results-v03 fucc boi bench results v0.8.0 This is the public results release for fucc boi bench. 41 models 96 prompts per model 3936 scored answers one combined leaderboard The site defaults to General-purpose models, with Wildcards as a secondary view and All models as an optional combined view. Access route and billing are metadata; they do not create separate rankings. Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v03.text-generation0 likes128 downloads20d agoHugging Face14rl-rag /rubric_rl_results Rubric RL Evaluation Results Evaluation data for rubric-based reward modeling experiments. Contains generated rubrics from multiple rubric generators and pairwise scoring results comparing rl-research/DR-Tulu-8B (RL, step_4000) vs rl-research/DR-Tulu-SFT-8B. Data Structure rubrics/ — Generated evaluation rubrics Each JSONL file contains per-question rubrics with fields: prompt_id, question, generated_rubric, generated_rubric_raw, rubric_style, rubric_model.… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/rubric_rl_results.text-generation0 likes126 downloads6mo agoHugging Face15mobileforge-anonymous /mobileforge-benchmark-results MobileForge Benchmark Results Anonymous project: https://mobileforge-anonymous.github.io/Anonymous code: https://github.com/mobileforge-anonymous/MobileForge This dataset contains the evaluation artifacts used by MobileForge: Annotation-Free Adaptation for Mobile GUI Agents with Hierarchical Feedback-Guided Policy Optimization. It includes AndroidWorld and MobileWorld GUI-only evaluation runs for the base agents and their MobileForge-adapted variants. The repository is intended… See the full description on the dataset page: https://huggingface.co/datasets/mobileforge-anonymous/mobileforge-benchmark-results.reinforcement-learning0 likes115 downloads27d agoHugging Face16Firmansyah-Ibrahim /idt5-v4-results-final-lora-s123-20260912T013040606815Z final-lora-s123-20260912T013040606815Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 94.7565543071161, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s123-20260912T013040606815Z.texttext-generationn<1K0 likes112 downloads13d agoHugging Face17ai-law-society-lab /oral-args-data-and-results Oral Arguments Arena Data repository for AI-Assisted Moot Courts: Simulating Justice-Specific Questioning in Oral Arguments (Zhang, Nadeem, Zheng, Stammbach, Henderson, 2026). Refer to the paper for background on the evaluation framework, experimental design, and findings. Repository Structure oral-args-arena-annotations/ ├── transcript_data/ # SCOTUS oral argument transcripts and case briefs ├── automated_metrics/ # LLM classifier outputs (SQLite… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/oral-args-data-and-results.text-classification100K<n<1M1 likes108 downloads6mo agoHugging Face18jacksonlukas /connections-rl-results connections-rl: raw evaluation artifacts Per-puzzle records, bootstrap summaries and analysis outputs backing connections-rl, a two-scale (Qwen2.5-1.5B / 7B), three-seed study of what verifiable-reward RL actually transfers. This is an artifact bundle for auditing published numbers, not a loadable training dataset, so the dataset viewer is disabled. Read this before using the numbers Two conventions in these files are easy to misread. Both have bitten this project… See the full description on the dataset page: https://huggingface.co/datasets/jacksonlukas/connections-rl-results.text-generationn<1K0 likes99 downloads18d agoHugging Face19Firmansyah-Ibrahim /idt5-v4-results-final-lora-s2026-20260912T034640190015Z final-lora-s2026-20260912T034640190015Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 95.88014981273409, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s2026-20260912T034640190015Z.texttext-generationn<1K0 likes89 downloads13d agoHugging Face20tomyimkc /sophia-training-results Sophia Training Results All training history results from the sophia-agi repository. Contents 824 result files across all Sophia four-theme training experiments: Prosoche — interactive decision-making, behavioral batteries, focus experiments RunPod training — LoRA training eval ladders (Qwen2.5-3B, multiple seeds/epochs) Continual learning — sequential curriculum, QA judged results World model — on-device, adaptive, real-corpus experiments Coherence reframe — NLI… See the full description on the dataset page: https://huggingface.co/datasets/tomyimkc/sophia-training-results.text-generation100K<n<1M1 likes83 downloads2mo agoHugging Face21ahmedBargady /open-models-benchmark-results ⚡ Local LLM Evaluation Leaderboard Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs. 💻 Hardware & System Specifications All evaluations are executed under standardized local cluster environments: Specification Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.tabulartext-generationn<1K1 likes73 downloads1mo agoHugging Face22EunsuKim /benchhub_plus_results_evaluated BenchHub Plus Results (Evaluated) LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores. Folder Structure ├── vllm_inference_results_en/ # English benchmark results (19 models) │ ├── {model_name}_{date}.jsonl │ └── ... └── vllm_inference_results_ko/ # Korean benchmark results (16 models) ├── {model_name}_{date}.jsonl └── ... Column Description Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.tabulartext-generation100K<n<1M0 likes68 downloads7mo agoHugging Face23Firmansyah-Ibrahim /idt5-v4-results-final-fft-s2026-20260911T063507183162Z final-fft-s2026-20260911T063507183162Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 88.01498127340824, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-fft-s2026-20260911T063507183162Z.texttext-generationn<1K0 likes67 downloads13d agoHugging Face24Honkware /catbench-results CatBench results Model outputs for CatBench, a small benchmark that asks a model to draw a cute kitten two ways and looks at what comes back. Produced by the /catbench command in blockquant. Upstream publishes its own results at Katehuuh.github.io/demos/CatBench/assets. This dataset holds runs for models that are not in that set. /catbench checks both and only rents a pod when neither has the model, so the two do not duplicate each other. The prompts Verbatim… See the full description on the dataset page: https://huggingface.co/datasets/Honkware/catbench-results.imagetext-generationn<1K0 likes65 downloads1mo agoHugging Face25Firmansyah-Ibrahim /idt5-v4-results-final-lora-s42-20260912T063343815032Z final-lora-s42-20260912T063343815032Z Run artifacts and per-item predictions. Phase: final. These are newly generated results, not a reproduction of the legacy TCI tables. See run_manifest.json, rules.json, generation_protocol.json and checkpoint_hashes.json. Structural scores do not establish semantic or Bloom validity. Metrics { "n": 267, "rule_version": "structural-proxy-v0.4-grounding-separated", "parse_success_pct": 92.88389513108615, "bleu":… See the full description on the dataset page: https://huggingface.co/datasets/Firmansyah-Ibrahim/idt5-v4-results-final-lora-s42-20260912T063343815032Z.texttext-generationn<1K0 likes63 downloads13d agoHugging Face26patrickleenyc /fucc-boi-bench-results-v02 fucc boi bench results v0.2 This is the public results release for fucc boi bench. 13 models 96 prompts per model 1248 scored answers one combined leaderboard Files: responses.jsonl: sanitized model outputs and run metadata. grades.jsonl: parsed grades, fuccboi scores, serious misses, and rationales. leaderboard.json: the ranked summary. summary.json: the interactive site's data and case examples. benchmark_report.md: the full short report. benchmark_card.md: the compact… See the full description on the dataset page: https://huggingface.co/datasets/patrickleenyc/fucc-boi-bench-results-v02.text-generation0 likes61 downloads23d agoHugging Face27zhongweixie /inplace-ttt-results In-Place TTT - Experimental Results This dataset contains experimental results and training logs for the In-Place Test-Time Training project. 📦 Dataset Contents 1. Experimental Results (results_20260904.tar.gz) Size: 191MB (compressed from 1.7GB) Files: 1,999 files Contents: Language model evaluation results RULER benchmark outputs ProLong training metrics Various test configurations 2. WandB Training Logs (wandb_20260904.tar.gz)… See the full description on the dataset page: https://huggingface.co/datasets/zhongweixie/inplace-ttt-results.text-generationn<1K0 likes61 downloads21d agoHugging Face28berkbirkan /turkish-seo-reasoning-benchmark-results Turkish SEO Reasoning Benchmark Results Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir. Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora Sonuç Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti. Mutlak artış: +10,28 puan Göreli artış: %85,97 Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.tabulartext-generationn<1K0 likes55 downloads2mo agoHugging Face29Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads26d agoHugging Face30daios /compartmentalized-harm-v1-results Four-model character-training results This package contains 29,952 paper-facing response records and their matching visible requests and blind outcome judgments. It covers Qwen3-8B, Qwen3-32B, Mistral Small 3.2 24B, and Gemma 4 31B at L1 and L2. Each model is evaluated with an identical rule-plus-hostile system message in the base and trained arms, and again with no experimental system prompt. Each lane has requests.jsonl, responses.jsonl, and judgments.jsonl. The public… See the full description on the dataset page: https://huggingface.co/datasets/daios/compartmentalized-harm-v1-results.text-generation0 likes45 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.