CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jinulee-v /expert-rag-benchmarks Expert RAG Benchmarks A unified collection of four expert-level legal RAG benchmarks, exposed as six named splits and three relational configurations: questions, documents, and qrels. The KCL split is named kcl_essay because Hugging Face split identifiers do not permit hyphens; its source name remains kcl-essay. Loading from datasets import load_dataset repo_id = "jinulee-v/expert-rag-benchmarks" questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.textquestion-answering1M<n<10M0 likes627 downloads4d agoHugging Face02yesilhealth /Health_Benchmarks LLM Health Benchmarks Dataset by Yesil Science The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge. Primary Purpose This dataset is built to: Benchmark LLMs in medical specialties and subfields. Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.textquestion-answering1K<n<10K10 likes450 downloads1y agoHugging Face03lthn /LEM-benchmarks LEM-benchmarks Canonical 8-PAC benchmark results for the Lemma model family. This dataset is an aggregated store of per-round evaluation data produced by lthn/LEM-Eval. Every row represents one model's answer to one question in one round of a paired A/B run against its unmodified base, and the dataset grows monotonically as more workers contribute — different machines, different sampling states, different hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.tabularquestion-answering10K<n<100K3 likes378 downloads5mo agoHugging Face04noel7Y /data-agent-benchmarks LongHorizon Full Data-Agent Benchmarks Companion data artifacts for five complete evaluation tracks: DataSciBench full55 / 167 metric entries DABStep full450 DABStep-Research full100 DSBench Modeling full74 LongDS full68 / 2,225 turns The companion GitHub repository contains processed manifests, evaluation code, historical API ReAct baseline code, download/preparation tools, and the frozen source lock. artifact_manifest.json records every uploaded object's size, SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.imagequestion-answering1 likes378 downloads14d agoHugging Face05yshenaw /SkillOpt_Lite_Benchmarks SkillOpt_Lite Benchmarks Train / val / test splits used by the SkillOpt_Lite project. One multi-config repo containing all six benchmarks: Config Rows (train / val / test) Content shipped searchqa 400 / 200 / 1400 Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA. docvqa 107 / 53 / 374 Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.imagequestion-answering1K<n<10K0 likes304 downloads3mo agoHugging Face06RISys-Lab /Benchmarks_CyberSec_RedSageMCQ Dataset Card for RedSage-MCQ Dataset Summary RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM". The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.tabularquestion-answering10K<n<100K0 likes263 downloads8mo agoHugging Face07RISys-Lab /Benchmarks_CyberSec_SecBench Dataset Card for SecBench (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.textquestion-answering1K<n<10K0 likes218 downloads8mo agoHugging Face08local-deep-research /ldr-benchmarks LDR Community Benchmarks (Leaderboards) Aggregated leaderboards for Local Deep Research (LDR) community benchmark runs against SimpleQA, BrowseComp, and xbench-DeepSearch. 👉 Submit results, read raw YAMLs, open PRs: github.com/LearningCircuit/ldr-benchmarks This Hugging Face dataset hosts only the aggregated CSV leaderboards. It is regenerated automatically on every merge to main in the GitHub repo above. Each CSV row represents one benchmark run (one strategy… See the full description on the dataset page: https://huggingface.co/datasets/local-deep-research/ldr-benchmarks.tabularquestion-answeringn<1K13 likes167 downloads4mo agoHugging Face09RaguTeam /ragu_benchmarks RAGU_Benchmarks MultiQ Dataset Dataset Description Dataset Summary MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.textquestion-answering1K<n<10K1 likes162 downloads9mo agoHugging Face10RISys-Lab /Benchmarks_CyberSec_CTI-Bench Dataset Card for CTIBench (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original CTIBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original CTIBench. It has been converted to… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CTI-Bench.texttext-classification1K<n<10K0 likes153 downloads8mo agoHugging Face11RLAIF-V /RLPR-Benchmarks Dataset Card for RLPR-Test GitHub | Paper News: [2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at arXiv! Dataset Summary We include the following seven benchmarks for evaluation of RLPR: Mathematical Reasoning Benchmarks: MATH-500 (Cobbe et al., 2021) Minerva (Lewkowycz et al., 2022) AIME24 General Domain Reasoning Benchmarks: MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF-V/RLPR-Benchmarks.textquestion-answeringn<1K0 likes118 downloads1y agoHugging Face12Abdulrahmankalil /enterprise-llm-inference-benchmarks-2026 🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide) A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments. 🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation) Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.tabulartext-generationn<1K1 likes96 downloads7d agoHugging Face13alex73 /benchmarks-stress-bel Пазначэнне націскаў у беларускіх амографах (Датасэты і Бэнчмаркі) Summary: This repository provides datasets and benchmark results for evaluating stress prediction in Belarusian homographs. It features three datasets (CommonVoice, a balanced 10x10 synthetic dataset, and a fully manually annotated literary text) to test context-aware stress assignment. The repository also includes benchmark results comparing statistical methods with state-of-the-art LLM-based approaches… See the full description on the dataset page: https://huggingface.co/datasets/alex73/benchmarks-stress-bel.texttoken-classification1K<n<10K0 likes94 downloads1mo agoHugging Face14RISys-Lab /Benchmarks_CyberSec_CyberMetrics Dataset Card for CyberMetric (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original CyberMetric dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original CyberMetric benchmark. It has… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CyberMetrics.textquestion-answering10K<n<100K0 likes82 downloads8mo agoHugging Face15patrickechohelloworld /well_formatted_benchmarks_pro Dataset Card for well_formatted_benchmarks_pro This is a collection of formatted benchmarks. Dataset Details Dataset Description This repo is home to formatted versions of some famous benchmarks I created this repo because current benchmark datasets on the hub generally don't have a fixed format, which is annoying when you try to use them. Language(s) (NLP): English Dataset Sources ARC Repository: Original ARC repo Demo: <user>An… See the full description on the dataset page: https://huggingface.co/datasets/patrickechohelloworld/well_formatted_benchmarks_pro.textzero-shot-classification1M<n<10M0 likes81 downloads1y agoHugging Face16RISys-Lab /Benchmarks_CyberSec_SECURE Dataset Card for SECURE (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.texttext-classification1K<n<10K0 likes79 downloads8mo agoHugging Face17pgmenon /soul-benchmarks-locomo soul.py LoCoMo Benchmark Results Benchmark results for soul.py on the LoCoMo long-conversation memory benchmark. Benchmarks repo: github.com/menonpg/soul-benchmarksInteractive results: menonpg.github.io/soul-benchmarks What is soul.py? soul.py is an open-source conversational memory layer for LLM agents. It provides multiple retrieval backends (BM25, Qdrant vector search, Relational Learning Model) and an auto-router that selects the best strategy per query.… See the full description on the dataset page: https://huggingface.co/datasets/pgmenon/soul-benchmarks-locomo.tabularquestion-answering1K<n<10K0 likes73 downloads4mo agoHugging Face18lllouo /BD-benchmarks BD-benchmarks: Denoised Benchmark Datasets Dataset Description BD-benchmarks is a comprehensive collection of denoised versions of popular NLP benchmark datasets. This repository contains original and cleaned versions of 11 widely-used benchmarks processed using two state-of-the-art denoising methods: DeepSeek-R1 and WAC-GEC (Whitespace Anomaly Correction - Grammar Error Correction). Dataset Summary This dataset addresses the critical issue of noise in… See the full description on the dataset page: https://huggingface.co/datasets/lllouo/BD-benchmarks.textquestion-answering10K<n<100K0 likes55 downloads8mo agoHugging Face19ChuGyouk /Qwen3.5-4B-nothink-benchmarks Qwen3.5-4B (non-thinking) — 13 benchmarks, multi-sample outputs with pass@k All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False) on 13 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was sent to the model, all sampled responses, their scores, and the benchmark-level metrics. Generation setup (identical for every benchmark) Model… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.texttext-generation10K<n<100K0 likes48 downloads3d agoHugging Face20Echo9k /mexican-legal-benchmarks Mexican Legal Benchmarks: Interpretation, Reasoning, and Cross-Jurisdiction Evaluation Dataset Summary The first specialized benchmark suite for evaluating language models on Mexican legal tasks. Three configs test distinct legal capabilities: practical interpretation of federal statutes, IRAC-structured reasoning chains with citation verification, and cross-jurisdiction comparison between Mexican states. 380 total samples across 3 benchmarks, drawn from Mexican federal… See the full description on the dataset page: https://huggingface.co/datasets/Echo9k/mexican-legal-benchmarks.texttext-classificationn<1K1 likes38 downloads5mo agoHugging Face21RISys-Lab /Benchmarks_CyberSec_SecEval Dataset Card for SecEval (RISys-Lab Mirror) ⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecEval benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below. Repository Intent This Hugging Face dataset is a re-host of the original SecEval benchmark. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecEval.textquestion-answering1K<n<10K0 likes32 downloads8mo agoHugging Face22AYI-NEDJIMI /security-tool-benchmarks-en Cybersecurity Tool Benchmarks - EN Bilingual dataset comparing the best cybersecurity tools by category. Created by AYI-NEDJIMI Consultants - Cybersecurity consulting firm. Dataset Contents Tool Comparisons (56 tools) Category # Tools Source Active Directory Audit 12 Top 10 AD Audit Tools 2025 EDR/XDR Solutions 12 Top 10 EDR/XDR Solutions 2025 Kubernetes Security 10 Top 10 Kubernetes Security Tools DFIR (Forensics & IR) 12 DFIR Tools… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/security-tool-benchmarks-en.textquestion-answeringn<1K0 likes20 downloads7mo agoHugging Face23AYI-NEDJIMI /security-tool-benchmarks-fr Benchmarks Outils de Cybersecurite - FR Dataset bilingue de comparaison des meilleurs outils de cybersecurite par categorie. Cree par AYI-NEDJIMI Consultants - Cabinet de conseil en cybersecurite. Contenu du Dataset Comparatifs d'Outils (56 outils) Categorie Nb Outils Source Audit Active Directory 12 Top 10 Outils Audit AD 2025 Solutions EDR/XDR 12 Top 10 Solutions EDR/XDR 2025 Securite Kubernetes 10 Top 10 Outils Securite Kubernetes DFIR… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/security-tool-benchmarks-fr.textquestion-answeringn<1K0 likes17 downloads7mo agoHugging Face24Mercury-Zz /benchmark_synthentictextquestion-answeringn<1K0 likes4 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.