datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eval-Qwen3-32B-HLEtest-datachain-llm-evaleval-result-Qwen3-32B-HLE
neko-llm/eval-result-Qwen3-32B-HLE
Judged HLE evaluation results uploaded via script.
Summary metrics
Total examples: 2129
Accuracy: 7.4213% ± 1.1134%
Average confidence: 93.9%
Calibration error (L2): 83.6859%
Models
neko-llm/Qwen3-32B-HLE
Columns
id
model
response
usage (JSON string)
judge_correct_answer
judge_model_answer
judge_reasoning
judge_correct
judge_confidence
judge_response (raw JSON)
Generated by evaluation/hle/upload_eval_result.py
llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/MultiverseComputingCAI/llm-refusal-evaluation.llm-sfm-safety-eval
LLM x SFM Safety Evaluation
When a general-purpose language model interprets the output of a specialist
science foundation model (a protein, genomic, RNA, or chemistry model), does its
safety behavior recognize the scientific content, or only the surface form of the
request?
This repository is the empirical core of a study of that question: the evaluation
harness, the redacted aggregate results, and the measurement specifications behind
four findings about how deployed Claude… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/llm-sfm-safety-eval.slovenian-llm-eval
Slovenian LLM Evaluation Dataset
This dataset is designed for evaluating Slovenian language models and builds upon the work of gordicaleksa/slovenian-llm-eval-v0 which translated some of the popular English benchmarks into Slovenian by using Google Translate. We have further improved the quality of the Slovenian translations.
The dataset contains the following benchmarks:
ARC Challenge
ARC Easy
BoolQ
GSM8K
HellaSwag
NQ Open
OpenBookQA
PIQA
TriviaQA
TruthfulQA
Winogrande… See the full description on the dataset page: https://huggingface.co/datasets/cjvt/slovenian-llm-eval.llm-jp-evalllm-eval-massive_intentllm-based-expansions-eval-datasetsllm-eval-massive_scenariollm_eval_promptsllm-eval-imdbllm-eval-mtop_domainllm-eval-banking77llm-fol-reasoning-eval
LLM FOL Reasoning Eval
This dataset is derived from ProverQA, a First-Order Logic reasoning benchmark designed to test the ability of large language models (LLMs) to perform structured logical reasoning.It restructures and normalizes the ProverQA development and training data into a unified, clean format suitable for evaluating chain-of-thought (CoT) and symbolic reasoning capabilities in LLMs.
Source
Original dataset: ProverQA: A First-Order Logic Reasoning… See the full description on the dataset page: https://huggingface.co/datasets/MinaGabriel/llm-fol-reasoning-eval.llm-eval-amazon_counterfactualllm-eval-amazon_reviewsllm-eval-tweet_sentimentllm-eval-toxic_conversationseval-Qwen-Qwen3-235B-A22B-openrouterLLMEval-Fair
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation,
built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines.
This is the publicly released subset of that bank.
Paper (arXiv): https://arxiv.org/abs/2508.05452
Venue: ACL 2026 Main Conference
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.llm-refusal-evaluation
🛡️ LLM Refusal Evaluation Benchmark
This repository contains the benchmarks used in the LLM-Refusal-Evaluation suite.
The prompts are organized into three groups:
Safety Benchmarks — harmful / jailbreak-style prompts that models should refuse.
Chinese Sensitive Topics — prompts that may be censored by China-aligned models.
Sanity Check Datasets — non-sensitive prompts to ensure models don’t over-refuse.
📌 Contents
Safety Benchmarks
JailbreakBench
SorryBench… See the full description on the dataset page: https://huggingface.co/datasets/BushNate/llm-refusal-evaluation.AutoRAG-evaluation-2024-LLM-paper-v1
AutoRAG evaluation dataset
Made with 2024 LLM resesarch articles (papers)
This dataset is an example for AutoRAG.
You can directly use this dataset for optimizng and benchmarking your RAG setup in AutoRAG.
How this dataset created?
This dataset is 100% synthetically generated by GPT-4 and Marker Inc. technology.
At first, we collected 110 latest LLM papers at arxiv.
We used Marker OCR model to extract texts.
And chunk it using MarkdownSplitter and… See the full description on the dataset page: https://huggingface.co/datasets/MarkrAI/AutoRAG-evaluation-2024-LLM-paper-v1.NLP-Course-LLM-Reasoning-Eval-May2025
Overview of LLM Reasoning Eval Dataset
This dataset contains evaluation of multiple large language models (LLMs) over 918 MCQ reasoning questions created by 184 students.
Each question was used to test 3 LLMs (each 3 times): GPT-4o, Claude Sonnet 3.x (3.5 or 3.7), and Deepseek R1.
The questions target various reasoning areas (i.e., Math, Logic, Temporal, Commonsense) and are included only if 3 seperate attempts (in a new session) by ChatGPT (GPT-4o) fail at giving the correct… See the full description on the dataset page: https://huggingface.co/datasets/nlpllmeval/NLP-Course-LLM-Reasoning-Eval-May2025.LLMEval-1
LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I)
LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024).
It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.serbian-llm-eval
Serbian LLM Eval
This dataset is a republish of
gordicaleksa/serbian-llm-eval-v1
as plain Parquet configs, so it can be loaded with modern versions of the
datasets library (the upstream repo ships a Python loading script that
datasets can no longer execute). Row contents are otherwise unchanged; only
the example_id field is synthesized where the upstream data has no unique
identifier, and the triviaqa answer struct is flattened into
answer_value / answer_aliases columns.… See the full description on the dataset page: https://huggingface.co/datasets/permitt/serbian-llm-eval.llm-eval-public-health-qaLLMEval-2
LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II)
LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models".
While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across
12 academic disciplines.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.llm-math-evaluation-dataset
LLM Math Response Evaluation Dataset
Dataset Summary
A human-annotated dataset of 150 AI-generated math responses
evaluated across GPT-4o, Claude, and Gemini. Each response is
scored on Correctness, Reasoning, and Clarity using a structured
rubric, with written justification for every score.
Supported Tasks
LLM evaluation and benchmarking
Math reasoning quality assessment
Error type classification in AI responses
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/nbvbharath-1729/llm-math-evaluation-dataset.llm-eval-aila-statutes
