CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01C0rk1 /vulpentestbench-results VulPentestBench -- agent results & trajectories Benchmark results of autonomous LLM penetration-testing agents (launched context-free) on VulPentestBench: an agent-evaluation harness that boots vulhub vulnerable targets in isolated Docker networks, injects a fresh random canary token at a vulnerability-reachable location per run, and scores provenance-verified milestones (the token must come back through a tool response of a target-aimed action before a flag submission counts --… See the full description on the dataset page: https://huggingface.co/datasets/C0rk1/vulpentestbench-results.tabulartext-generation10K<n<100K1 likes285 downloads12d agoHugging Face02ahmedBargady /open-models-benchmark-results ⚡ Local LLM Evaluation Leaderboard Welcome to the official public benchmark leaderboard maintained by @ahmedBargady.This dataset repository hosts benchmark evaluation metrics, accuracy scores, throughput telemetry, and quantization trade-off analyses of open-weights foundation models tested locally on NVIDIA A100 GPUs. 💻 Hardware & System Specifications All evaluations are executed under standardized local cluster environments: Specification Details… See the full description on the dataset page: https://huggingface.co/datasets/ahmedBargady/open-models-benchmark-results.tabulartext-generationn<1K1 likes73 downloads1mo agoHugging Face03EunsuKim /benchhub_plus_results_evaluated BenchHub Plus Results (Evaluated) LLM inference results on the BenchHub Plus benchmark, with per-sample accuracy scores. Folder Structure ├── vllm_inference_results_en/ # English benchmark results (19 models) │ ├── {model_name}_{date}.jsonl │ └── ... └── vllm_inference_results_ko/ # Korean benchmark results (16 models) ├── {model_name}_{date}.jsonl └── ... Column Description Each .jsonl file contains one JSON object per line with the… See the full description on the dataset page: https://huggingface.co/datasets/EunsuKim/benchhub_plus_results_evaluated.tabulartext-generation100K<n<1M0 likes68 downloads7mo agoHugging Face04berkbirkan /turkish-seo-reasoning-benchmark-results Turkish SEO Reasoning Benchmark Results Bu dataset, Turkish SEO Reasoning benchmark'ının altı farklı model/checkpoint üzerinde çalıştırılmış ham tahminlerini, metriklerini ve tekrar üretim manifestlerini içerir. Fine-tuned model: berkbirkan/gemma-3-1b-turkish-seo-reasoning-lora Sonuç Fine-tuned Gemma 3 1B modeli 22,23 skorla ilk sırada yer aldı. Aynı base model 11,96 skor elde etti. Mutlak artış: +10,28 puan Göreli artış: %85,97 Fine-tuned model hata sayısı:… See the full description on the dataset page: https://huggingface.co/datasets/berkbirkan/turkish-seo-reasoning-benchmark-results.tabulartext-generationn<1K0 likes55 downloads2mo agoHugging Face05Experimental-Orange /HumanAgencyBench_Evaluation_Results HumanAgencyBench evaluation results Paper: HumanAgencyBench: Scalable Evaluation of Human Agency Support in AI Assistants Code: https://github.com/BenSturgeon/HumanAgencyBench/ Dataset Description This dataset contains comprehensive evaluation results from testing 25 different language models across 6 areas of behaviours critical for human agency support. Each model was evaluated on 3,000 prompts (500 per category), resulting in 75,000 total evaluations designed… See the full description on the dataset page: https://huggingface.co/datasets/Experimental-Orange/HumanAgencyBench_Evaluation_Results.tabulartext-generation10K<n<100K0 likes47 downloads26d agoHugging Face06goktugozkanmd /medfailbench-v02-results MedFailBench v0.2 Model Response Artifacts This dataset preserves full model responses and automated rule scores generated during MedFailBench development. Each row links a prompt, a model identifier, the response text, available scores, a final label, and its source artifact. Public data scope The current file contains 181 rows with 11 distinct values in model_name. These values are run identifiers rather than a normalized model registry. They must not be… See the full description on the dataset page: https://huggingface.co/datasets/goktugozkanmd/medfailbench-v02-results.tabulartext-generationn<1K0 likes37 downloads2mo agoHugging Face07anote-ai /codebench-results CodeBench Results Experimental results from the CodeBench evaluation framework. Two dataset configs are included. Configs h4_security — H4 Security-Adjusted Reliability Experiment 240 rollouts across 3 agents × 10 tasks × 8 rollouts. Tasks split into 5 standard algorithmic tasks and 5 security-sensitive tasks (eval, exec, shell, yaml, dynamic import). Column Description task_id Task identifier (algo-* or sec-*) agent_name anote-code… See the full description on the dataset page: https://huggingface.co/datasets/anote-ai/codebench-results.tabulartext-generationn<1K0 likes34 downloads1mo agoHugging Face08cmpatino /math500-bon-weighted-results MATH-500 Best-of-N Weighted Selection Results Dataset Description This dataset contains the results of evaluating Best-of-N weighted selection on a subset of the MATH-500 benchmark. It was created as part of a HuggingFace internship exercise exploring how test-time compute scaling with reward models can improve LLM performance on math problems. How It Was Constructed 1. Problem Selection Started from the HuggingFaceH4/MATH-500 dataset (500 problems)… See the full description on the dataset page: https://huggingface.co/datasets/cmpatino/math500-bon-weighted-results.tabulartext-generationn<1K1 likes32 downloads5mo agoHugging Face09HerrHruby /imo-proofbench-mr-v3sft-results IMO ProofBench — Meta-Reasoning v3-SFT eval results GPT-5.4-judged results of a meta-reasoning SFT model (midtrain v3) on HerrHruby/imo-proofbench-all-vf (60 modified-olympiad problems). For each problem the model is sampled 3× (180 candidates total). Model & inference Model: Qwen3.5-9B, midtrain-v3 SFT (checkpoint global_step_4500). Scaffold: a single model runs the full meta-reasoning loop — MR (propose exploration directions) → E (execute each direction… See the full description on the dataset page: https://huggingface.co/datasets/HerrHruby/imo-proofbench-mr-v3sft-results.tabulartext-generationn<1K0 likes27 downloads3mo agoHugging Face10vazirani /concept-injection-results Concept Injection Results Experimental data and analysis logs from the research "Replicating Introspection on Injected Content in Open-Source Language Models"    Overview This project enables researchers to inject concept vectors directly into a model's hidden layers during inference, allowing investigation of whether language models can detect and report on artificially induced "thoughts." This dataset contains the raw responses and processed evaluations produced… See the full description on the dataset page: https://huggingface.co/datasets/vazirani/concept-injection-results.tabulartext-generationn<1K0 likes26 downloads9mo agoHugging Face11benchpress /results-multi-temp-seed Benchpress Evaluation Results This dataset contains normalized outputs from lm-evaluation-harness runs for the Benchpress project. It is intended for analysis of post-training recipe behavior across models, benchmarks, temperatures, and random seeds. Coverage Runs: 110 Model/stage recipes: olmo2-13b:dpo, olmo2-13b:rlvr, olmo2-13b:sft, olmo2-7b:dpo, olmo2-7b:rlvr, olmo2-7b:sft, olmo3-7b:dpo, olmo3-7b:rlvr, olmo3-7b:sft, smollm3-3b:apo, smollm3-3b:sft Temperatures:… See the full description on the dataset page: https://huggingface.co/datasets/benchpress/results-multi-temp-seed.tabulartext-generation1M<n<10M0 likes21 downloads4mo agoHugging Face12SparkSupernova /nova-industry-benchmark-results Nova Industry Benchmark — Results Model run outputs for the questions in SparkSupernova/nova-industry-benchmark. Why results live in their own repository Results were previously stored as extra splits of the question dataset. Because each model version wrote a different set of columns, adding the v5 run gave that dataset two splits with incompatible schemas, and load_dataset failed for every consumer — including the usage example on the model card. Questions are a… See the full description on the dataset page: https://huggingface.co/datasets/SparkSupernova/nova-industry-benchmark-results.tabulartext-generationn<1K0 likes20 downloads2mo agoHugging Face13Legal-verse /sinergi-model-evaluation-results Sinergi model evaluation outputs Model answers and generation telemetry used by the Sinergi Table 8-aligned evaluation notebook. The repository contains eight configurations so every system can be loaded independently with the Hugging Face datasets library. from datasets import load_dataset data = load_dataset( "Legal-verse/sinergi-model-evaluation-results", "qwen-sft-rl-rag", split="test", ) Configurations Configuration System Rows Source file… See the full description on the dataset page: https://huggingface.co/datasets/Legal-verse/sinergi-model-evaluation-results.tabulartext-generation10K<n<100K0 likes13 downloads1mo agoHugging Face14HUFS-DILAB /PREPAIR-reproduction-results PREPAIR Reproduction Results This dataset contains the reproduction results of the PREPAIR method. Source Project: PREPAIR Target File: results_LLMBar_all_methods.csv tabulartext-generationn<1K0 likes9 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.