CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01k-mktr /llm_eval_promptstextquestion-answering1K<n<10K1 likes124 downloads2y agoHugging Face02llmeval-fdu /LLMEval-Fair LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation, built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines. This is the publicly released subset of that bank. Paper (arXiv): https://arxiv.org/abs/2508.05452 Venue: ACL 2026 Main Conference Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.textquestion-answering100K<n<1M0 likes99 downloads5mo agoHugging Face03llmeval-fdu /LLMEval-2 LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II) LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models". While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across 12 academic disciplines. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.textquestion-answering1K<n<10K0 likes79 downloads5mo agoHugging Face04llmeval-fdu /LLMEval-1 LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I) LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024). It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.textquestion-answering1K<n<10K0 likes79 downloads5mo agoHugging Face05llmeval-fdu /LLMEval-Med LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation LLMEval-Med is a physician-validated benchmark for evaluating Large Language Models on real-world medical tasks. The questions are drawn from real electronic health records and expert-designed clinical scenarios, and the LLM-as-Judge evaluation pipeline is calibrated against medical experts. Paper (arXiv): https://arxiv.org/abs/2506.04078 ACL Anthology:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Med.question-answeringn<1K0 likes70 downloads5mo agoHugging Face06llmeval-fdu /LLMEval-Logic LLMEval-Logic — Public 80% Release A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening. 📄 Paper (arXiv): https://arxiv.org/abs/2605.19597 🌐 Project: https://llmeval.com/ 🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic 🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic ⚠️ This is the 80% public release LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.textquestion-answeringn<1K5 likes64 downloads4mo agoHugging Face07sohaibdevv /llm-eval-benchmark LLM Evaluation Benchmark A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness. Sourced from MMLU and TruthfulQA, cleaned and formatted for the LLM Evaluation Framework. Dataset Summary Split Samples Use train 500 Fine-tuning reference / training baselines validation 200 Hyperparameter tuning test 500 Final benchmark — use this for fair comparisons Total 1,200 Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.textquestion-answering1K<n<10K0 likes60 downloads4mo agoHugging Face08vigneshwar234 /llm-eval-benchmark LLM Evaluation Benchmark A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness. Sourced from MMLU and TruthfulQA, cleaned and formatted for the LLM Evaluation Framework. Dataset Summary Split Samples Use train 500 Fine-tuning reference / training baselines validation 200 Hyperparameter tuning test 500 Final benchmark — use this for fair comparisons Total 1,200 Quick… See the full description on the dataset page: https://huggingface.co/datasets/vigneshwar234/llm-eval-benchmark.textquestion-answering1K<n<10K1 likes53 downloads4mo agoHugging Face09HappyEval /LLMEval-Reality_CheckThis repository contains model evaluation results with the following setup: Models evaluated: GPT-5 Gemini-2.5-Pro Claude-4.5-Sonnet Datasets included: MMLU-Pro (test split) GPQA (three subsets) MATH-500 MMMU-Pro (standard 10 options and vision versions) Each split in the dataset corresponds to one benchmark. Schema All datasets have been standardized to a unified schema with the following features: dataset_info: features: - id: int64 - prompt: string… See the full description on the dataset page: https://huggingface.co/datasets/HappyEval/LLMEval-Reality_Check.textquestion-answering10K<n<100K0 likes26 downloads11mo agoHugging Face10stindardlogic /llm-evaluation-sft-100k LLM Evaluation SFT 100K A synthetic supervised fine-tuning dataset of 100,000 high-quality conversations covering LLM evaluation methodology, benchmarking, safety assessment, and prompt optimization. Designed to train AI assistants that can help ML engineers and researchers rigorously evaluate and improve language models. Dataset Description This dataset covers the full spectrum of LLM evaluation practice across 7 specialized categories. Each record follows the… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/llm-evaluation-sft-100k.text-generation100K<n<1M0 likes21 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.