datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
expert-rag-benchmarks
Expert RAG Benchmarks
A unified collection of four expert-level legal RAG benchmarks, exposed as six
named splits and three relational configurations: questions, documents, and
qrels.
The KCL split is named kcl_essay because Hugging Face split identifiers do not
permit hyphens; its source name remains kcl-essay.
Loading
from datasets import load_dataset
repo_id = "jinulee-v/expert-rag-benchmarks"
questions = load_dataset(repo_id, "questions", split="housing")… See the full description on the dataset page: https://huggingface.co/datasets/jinulee-v/expert-rag-benchmarks.Health_Benchmarks
LLM Health Benchmarks Dataset by Yesil Science
The LLM Health Benchmarks Dataset is a specialized resource for evaluating large language models (LLMs) in different medical specialties. It provides structured question-answer pairs designed to test the performance of AI models in understanding and generating domain-specific knowledge.
Primary Purpose
This dataset is built to:
Benchmark LLMs in medical specialties and subfields.
Assess the accuracy and contextual… See the full description on the dataset page: https://huggingface.co/datasets/yesilhealth/Health_Benchmarks.LEM-benchmarks
LEM-benchmarks
Canonical 8-PAC benchmark results for the Lemma model family.
This dataset is an aggregated store of per-round evaluation data produced by
lthn/LEM-Eval. Every row
represents one model's answer to one question in one round of a paired A/B
run against its unmodified base, and the dataset grows monotonically as more
workers contribute — different machines, different sampling states, different
hardware paths — which is the whole point of 8-PAC: multiple independent… See the full description on the dataset page: https://huggingface.co/datasets/lthn/LEM-benchmarks.data-agent-benchmarks
LongHorizon Full Data-Agent Benchmarks
Companion data artifacts for five complete evaluation tracks:
DataSciBench full55 / 167 metric entries
DABStep full450
DABStep-Research full100
DSBench Modeling full74
LongDS full68 / 2,225 turns
The companion GitHub repository contains processed manifests, evaluation code,
historical API ReAct baseline code, download/preparation tools, and the frozen
source lock. artifact_manifest.json records every uploaded object's size,
SHA-256… See the full description on the dataset page: https://huggingface.co/datasets/noel7Y/data-agent-benchmarks.SkillOpt_Lite_Benchmarks
SkillOpt_Lite Benchmarks
Train / val / test splits used by the SkillOpt_Lite project.
One multi-config repo containing all six benchmarks:
Config
Rows (train / val / test)
Content shipped
searchqa
400 / 200 / 1400
Full QA — id, question, list of DOC contexts, answers. Sampled from dl4ir-searchQA.
docvqa
107 / 53 / 374
Full QA + images bundled — parquet has id/question/answers/topic/image_path; PNGs live under docvqa_images/ at the repo root. Subset of… See the full description on the dataset page: https://huggingface.co/datasets/yshenaw/SkillOpt_Lite_Benchmarks.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.Benchmarks_CyberSec_SecBench
Dataset Card for SecBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.ldr-benchmarks
LDR Community Benchmarks (Leaderboards)
Aggregated leaderboards for Local Deep Research (LDR) community benchmark
runs against SimpleQA, BrowseComp, and xbench-DeepSearch.
👉 Submit results, read raw YAMLs, open PRs:
github.com/LearningCircuit/ldr-benchmarks
This Hugging Face dataset hosts only the aggregated CSV leaderboards.
It is regenerated automatically on every merge to main in the GitHub
repo above. Each CSV row represents one benchmark run (one strategy… See the full description on the dataset page: https://huggingface.co/datasets/local-deep-research/ldr-benchmarks.ragu_benchmarks
RAGU_Benchmarks
MultiQ Dataset
Dataset Description
Dataset Summary
MultiQ is a small but rich dataset designed for question answering (QA) and multi-document information retrieval tasks. It contains 169 Russian-language questions, each accompanied by a correct answer and a set of relevant Wikipedia articles serving as context for locating the answer. This dataset is suitable for evaluating models’ ability to identify precise answers based on multiple… See the full description on the dataset page: https://huggingface.co/datasets/RaguTeam/ragu_benchmarks.Benchmarks_CyberSec_CTI-Bench
Dataset Card for CTIBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original CTIBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original CTIBench. It has been converted to… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CTI-Bench.RLPR-Benchmarks
Dataset Card for RLPR-Test
GitHub | Paper
News:
[2025.06.23] 📃 Our paper detailing the RLPR framework and its comprehensive evaluation using this suite is accessible at arXiv!
Dataset Summary
We include the following seven benchmarks for evaluation of RLPR:
Mathematical Reasoning Benchmarks:
MATH-500 (Cobbe et al., 2021)
Minerva (Lewkowycz et al., 2022)
AIME24
General Domain Reasoning Benchmarks:
MMLU-Pro (Wang et al., 2024): A multitask language… See the full description on the dataset page: https://huggingface.co/datasets/RLAIF-V/RLPR-Benchmarks.enterprise-llm-inference-benchmarks-2026
🚀 Enterprise LLM Inference & Fine-Tuning Benchmarks (2026 Guide)
A curated benchmark index and architectural guide evaluating open-source foundation models, real-time inference engines (vLLM vs. TensorRT-LLM), and cloud GPU economics for enterprise deployments.
🧠 Open-Source Foundation Model Benchmarks (RAG & Code Generation)
Flagship Evaluation: Top Open-Source LLMs for Enterprise RAG & Code Generation (2026 In-Depth Guide) — Comparing Qwen 2.5 Coder, Llama… See the full description on the dataset page: https://huggingface.co/datasets/Abdulrahmankalil/enterprise-llm-inference-benchmarks-2026.benchmarks-stress-bel
Пазначэнне націскаў у беларускіх амографах (Датасэты і Бэнчмаркі)
Summary: This repository provides datasets and benchmark results for evaluating stress prediction in Belarusian homographs.
It features three datasets (CommonVoice, a balanced 10x10 synthetic dataset, and a fully manually annotated literary text)
to test context-aware stress assignment. The repository also includes benchmark results comparing statistical methods
with state-of-the-art LLM-based approaches… See the full description on the dataset page: https://huggingface.co/datasets/alex73/benchmarks-stress-bel.Benchmarks_CyberSec_CyberMetrics
Dataset Card for CyberMetric (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original CyberMetric dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original CyberMetric benchmark. It has… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_CyberMetrics.well_formatted_benchmarks_pro
Dataset Card for well_formatted_benchmarks_pro
This is a collection of formatted benchmarks.
Dataset Details
Dataset Description
This repo is home to formatted versions of some famous benchmarks
I created this repo because current benchmark datasets on the hub generally don't have a fixed format, which is annoying when you try to use them.
Language(s) (NLP): English
Dataset Sources
ARC
Repository: Original ARC repo
Demo:
<user>An… See the full description on the dataset page: https://huggingface.co/datasets/patrickechohelloworld/well_formatted_benchmarks_pro.Benchmarks_CyberSec_SECURE
Dataset Card for SECURE (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SECURE benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SECURE benchmark. It has been converted… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SECURE.soul-benchmarks-locomo
soul.py LoCoMo Benchmark Results
Benchmark results for soul.py on the LoCoMo long-conversation memory benchmark.
Benchmarks repo: github.com/menonpg/soul-benchmarksInteractive results: menonpg.github.io/soul-benchmarks
What is soul.py?
soul.py is an open-source conversational memory layer for LLM agents. It provides multiple retrieval backends (BM25, Qdrant vector search, Relational Learning Model) and an auto-router that selects the best strategy per query.… See the full description on the dataset page: https://huggingface.co/datasets/pgmenon/soul-benchmarks-locomo.BD-benchmarks
BD-benchmarks: Denoised Benchmark Datasets
Dataset Description
BD-benchmarks is a comprehensive collection of denoised versions of popular NLP benchmark datasets. This repository contains original and cleaned versions of 11 widely-used benchmarks processed using two state-of-the-art denoising methods: DeepSeek-R1 and WAC-GEC (Whitespace Anomaly Correction - Grammar Error Correction).
Dataset Summary
This dataset addresses the critical issue of noise in… See the full description on the dataset page: https://huggingface.co/datasets/lllouo/BD-benchmarks.Qwen3.5-4B-nothink-benchmarks
Qwen3.5-4B (non-thinking) — 13 benchmarks, multi-sample outputs with pass@k
All sampled outputs of Qwen/Qwen3.5-4B in non-thinking mode (enable_thinking=False)
on 13 benchmarks, with per-response correctness and pass@k / avg@n metrics. One row per problem; every row carries the exact prompt that was
sent to the model, all sampled responses, their scores, and the benchmark-level metrics.
Generation setup (identical for every benchmark)
Model… See the full description on the dataset page: https://huggingface.co/datasets/ChuGyouk/Qwen3.5-4B-nothink-benchmarks.mexican-legal-benchmarks
Mexican Legal Benchmarks: Interpretation, Reasoning, and Cross-Jurisdiction Evaluation
Dataset Summary
The first specialized benchmark suite for evaluating language models on Mexican legal tasks. Three configs test distinct legal capabilities: practical interpretation of federal statutes, IRAC-structured reasoning chains with citation verification, and cross-jurisdiction comparison between Mexican states.
380 total samples across 3 benchmarks, drawn from Mexican federal… See the full description on the dataset page: https://huggingface.co/datasets/Echo9k/mexican-legal-benchmarks.Benchmarks_CyberSec_SecEval
Dataset Card for SecEval (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecEval benchmark.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit belongs to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecEval benchmark. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecEval.security-tool-benchmarks-en
Cybersecurity Tool Benchmarks - EN
Bilingual dataset comparing the best cybersecurity tools by category.
Created by AYI-NEDJIMI Consultants - Cybersecurity consulting firm.
Dataset Contents
Tool Comparisons (56 tools)
Category
# Tools
Source
Active Directory Audit
12
Top 10 AD Audit Tools 2025
EDR/XDR Solutions
12
Top 10 EDR/XDR Solutions 2025
Kubernetes Security
10
Top 10 Kubernetes Security Tools
DFIR (Forensics & IR)
12
DFIR Tools… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/security-tool-benchmarks-en.security-tool-benchmarks-fr
Benchmarks Outils de Cybersecurite - FR
Dataset bilingue de comparaison des meilleurs outils de cybersecurite par categorie.
Cree par AYI-NEDJIMI Consultants - Cabinet de conseil en cybersecurite.
Contenu du Dataset
Comparatifs d'Outils (56 outils)
Categorie
Nb Outils
Source
Audit Active Directory
12
Top 10 Outils Audit AD 2025
Solutions EDR/XDR
12
Top 10 Solutions EDR/XDR 2025
Securite Kubernetes
10
Top 10 Outils Securite Kubernetes
DFIR… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/security-tool-benchmarks-fr.benchmark_synthentic
