datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-eval-massive_intentllm-eval-massive_scenariollm-eval-imdbllm-eval-mtop_domainllm_eval_promptsllm-eval-banking77llm-eval-amazon_reviewsllm-eval-amazon_counterfactualllm-eval-toxic_conversationsllm-eval-tweet_sentimentLLMEval-Fair
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation,
built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines.
This is the publicly released subset of that bank.
Paper (arXiv): https://arxiv.org/abs/2508.05452
Venue: ACL 2026 Main Conference
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.llm-eval-public-health-qaLLMEval-2
LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II)
LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models".
While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across
12 academic disciplines.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.LLMEval-1
LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I)
LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024).
It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.llm-eval-aila-statutesllm-eval-hc3-financeLLMEval-Logic
LLMEval-Logic — Public 80% Release
A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening.
📄 Paper (arXiv): https://arxiv.org/abs/2605.19597
🌐 Project: https://llmeval.com/
🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic
🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic
⚠️ This is the 80% public release
LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.llm-eval-sts22_v2llm-eval-fquadllm-eval-biossesllm-eval-legalbench-consumer-contractsllm-eval-stsbenchmarkllm-eval-twitter-hjernellm-eval-requestsllm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/vigneshwar234/llm-eval-benchmark.llm-eval-sprint_duplicate_questionsllm-eval-sickrllm-eval-sts17llm-eval-medrxiv_clustering_p2p_v2
