datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm-eval-flakiness-trajectories
Llm Eval Flakiness Trajectories
Rights & intended use: legacy public research corpus / portfolio
artifact. Hosted frontier-model outputs are research-only inputs under
project policy (synthetic-factory#161):
intended_use: research_only, project_training_policy: blocked. Not
training data for any model-weight update. Machine-readable record:
rights.json.
Release status: The raw, uncurated payload is now published under
data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/llm-eval-flakiness-trajectories.llm-eval-massive_intentllm-eval-massive_scenariollm_eval_promptsllm-eval-imdbllm-eval-mtop_domainllm-eval-banking77llm-eval-amazon_counterfactualllm-eval-amazon_reviewsllm-eval-tweet_sentimentllm-eval-toxic_conversationsLLMEval-Fair
LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models
LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation,
built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines.
This is the publicly released subset of that bank.
Paper (arXiv): https://arxiv.org/abs/2508.05452
Venue: ACL 2026 Main Conference
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.LLMEval-1
LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I)
LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024).
It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.llm-eval-public-health-qaLLMEval-2
LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II)
LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab),
released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models".
While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across
12 academic disciplines.
Paper: https://arxiv.org/abs/2312.07398
Project website: https://llmeval.com/
GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.llm-eval-aila-statutesptb-llmevalmedLLMEval-Med
LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation
LLMEval-Med is a physician-validated benchmark for evaluating Large Language Models on real-world
medical tasks. The questions are drawn from real electronic health records and expert-designed
clinical scenarios, and the LLM-as-Judge evaluation pipeline is calibrated against medical experts.
Paper (arXiv): https://arxiv.org/abs/2506.04078
ACL Anthology:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Med.llm-eval-hc3-financeLLMEval-Logic
LLMEval-Logic — Public 80% Release
A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening.
📄 Paper (arXiv): https://arxiv.org/abs/2605.19597
🌐 Project: https://llmeval.com/
🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic
🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic
⚠️ This is the 80% public release
LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.llm-eval-benchmark
LLM Evaluation Benchmark
A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness.
Sourced from MMLU and TruthfulQA, cleaned and formatted for the
LLM Evaluation Framework.
Dataset Summary
Split
Samples
Use
train
500
Fine-tuning reference / training baselines
validation
200
Hyperparameter tuning
test
500
Final benchmark — use this for fair comparisons
Total
1,200
Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.llm-eval-sts22_v2llm-eval-fquadllm-eval-biossesllm-eval-legalbench-consumer-contractsllm-eval-requestsllm-eval-sickrllm-eval-sprint_duplicate_questionsllm-eval-twitter-hjernellm-eval-sts17
