datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_eval_promptsLLMEval-Fair
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation,
built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines.
This is the publicly released subset of that bank.
Paper (arXiv): https://arxiv.org/abs/2508.05452
Venue: ACL 2026 Main Conference
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.LLMEval-2
LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II)
LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models".
While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across
12 academic disciplines.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.LLMEval-1
LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I)
LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024).
It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.LLMEval-Med
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
LLMEval-Med is a physician-validated benchmark for evaluating Large Language Models on real-world
medical tasks. The questions are drawn from real electronic health records and expert-designed
clinical scenarios, and the LLM-as-Judge evaluation pipeline is calibrated against medical experts.
Paper (arXiv): https://arxiv.org/abs/2506.04078
ACL Anthology:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Med.LLMEval-Logic
LLMEval-Logic — Public 80% Release
A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening.
📄 Paper (arXiv): https://arxiv.org/abs/2605.19597
🌐 Project: https://llmeval.com/
🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic
🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic
⚠️ This is the 80% public release
LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/vigneshwar234/llm-eval-benchmark.LLMEval-Reality_CheckThis repository contains model evaluation results with the following setup:
Models evaluated:
GPT-5
Gemini-2.5-Pro
Claude-4.5-Sonnet
Datasets included:
MMLU-Pro (test split)
GPQA (three subsets)
MATH-500
MMMU-Pro (standard 10 options and vision versions)
Each split in the dataset corresponds to one benchmark.
Schema
All datasets have been standardized to a unified schema with the following features:
dataset_info:
features:
- id: int64
- prompt: string… See the full description on the dataset page: https://huggingface.co/datasets/HappyEval/LLMEval-Reality_Check.llm-evaluation-sft-100k
LLM Evaluation SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering LLM evaluation methodology, benchmarking, safety assessment, and prompt optimization. Designed to train AI assistants that can help ML engineers and researchers rigorously evaluate and improve language models.
Dataset Description
This dataset covers the full spectrum of LLM evaluation practice across 7 specialized categories. Each record follows the… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/llm-evaluation-sft-100k.
