CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rmems /llm-eval-flakiness-trajectories Llm Eval Flakiness Trajectories Rights & intended use: legacy public research corpus / portfolio artifact. Hosted frontier-model outputs are research-only inputs under project policy (synthetic-factory#161): intended_use: research_only, project_training_policy: blocked. Not training data for any model-weight update. Machine-readable record: rights.json. Release status: The raw, uncurated payload is now published under data/raw/. It is available for inspection and… See the full description on the dataset page: https://huggingface.co/datasets/rmems/llm-eval-flakiness-trajectories.1 likes363 downloads3d agoHugging Face02mteb /llm-eval-massive_intenttext10K<n<100K0 likes154 downloads6mo agoHugging Face03mteb /llm-eval-massive_scenariotext10K<n<100K0 likes135 downloads6mo agoHugging Face04k-mktr /llm_eval_promptstextquestion-answering1K<n<10K1 likes132 downloads2y agoHugging Face05mteb /llm-eval-imdbtext10K<n<100K0 likes126 downloads7mo agoHugging Face06mteb /llm-eval-mtop_domaintext10K<n<100K0 likes125 downloads6mo agoHugging Face07mteb /llm-eval-banking77text10K<n<100K0 likes122 downloads7mo agoHugging Face08mteb /llm-eval-amazon_counterfactualtext10K<n<100K0 likes114 downloads6mo agoHugging Face09mteb /llm-eval-amazon_reviewstext1M<n<10M0 likes108 downloads7mo agoHugging Face10mteb /llm-eval-tweet_sentimenttext10K<n<100K0 likes108 downloads6mo agoHugging Face11mteb /llm-eval-toxic_conversationstext10K<n<100K0 likes107 downloads6mo agoHugging Face12llmeval-fdu /LLMEval-Fair LLMEval-Fair: A Large-Scale Longitudinal Study on Robust and Fair Evaluation of Large Language Models LLMEval-Fair is a 30-month longitudinal study on the robustness and fairness of LLM evaluation, built on a proprietary bank of 220,000+ graduate-level Chinese questions spanning 13 academic disciplines. This is the publicly released subset of that bank. Paper (arXiv): https://arxiv.org/abs/2508.05452 Venue: ACL 2026 Main Conference Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Fair.textquestion-answering100K<n<1M0 likes98 downloads5mo agoHugging Face13llmeval-fdu /LLMEval-1 LLMEval-1: Large-Scale Chinese LLM Evaluation (Phase I) LLMEval-1 is the Phase I evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models" (AAAI 2024). It is a Chinese, generative-QA benchmark designed to study how large language models should be evaluated. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub: https://github.com/llmeval/LLMEval-1… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-1.textquestion-answering1K<n<10K0 likes80 downloads5mo agoHugging Face14mteb /llm-eval-public-health-qatextn<1K0 likes79 downloads4mo agoHugging Face15llmeval-fdu /LLMEval-2 LLMEval-2: Professional Domain Evaluation of Chinese LLMs (Phase II) LLMEval-2 is the Phase II evaluation dataset of the LLMEval project (Fudan NLP Lab), released alongside the AAAI 2024 paper "LLMEval: A Preliminary Study on How to Evaluate Large Language Models". While LLMEval-1 focuses on general capabilities, LLMEval-2 targets professional domain evaluation across 12 academic disciplines. Paper: https://arxiv.org/abs/2312.07398 Project website: https://llmeval.com/ GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-2.textquestion-answering1K<n<10K0 likes77 downloads5mo agoHugging Face16mteb /llm-eval-aila-statutestextn<1K0 likes70 downloads6mo agoHugging Face17HaimingW /ptb-llmevalmedtabularn<1K0 likes70 downloads3mo agoHugging Face18llmeval-fdu /LLMEval-Med LLMEval-Med: A Real-world Clinical Benchmark for Medical LLMs with Physician Validation LLMEval-Med is a physician-validated benchmark for evaluating Large Language Models on real-world medical tasks. The questions are drawn from real electronic health records and expert-designed clinical scenarios, and the LLM-as-Judge evaluation pipeline is calibrated against medical experts. Paper (arXiv): https://arxiv.org/abs/2506.04078 ACL Anthology:… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Med.question-answeringn<1K0 likes67 downloads5mo agoHugging Face19mteb /llm-eval-hc3-financetextn<1K0 likes65 downloads4mo agoHugging Face20llmeval-fdu /LLMEval-Logic LLMEval-Logic — Public 80% Release A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening. 📄 Paper (arXiv): https://arxiv.org/abs/2605.19597 🌐 Project: https://llmeval.com/ 🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic 🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic ⚠️ This is the 80% public release LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.textquestion-answeringn<1K5 likes63 downloads4mo agoHugging Face21sohaibdevv /llm-eval-benchmark LLM Evaluation Benchmark A 1,200-sample curated benchmark dataset for evaluating LLMs on factual accuracy and truthfulness. Sourced from MMLU and TruthfulQA, cleaned and formatted for the LLM Evaluation Framework. Dataset Summary Split Samples Use train 500 Fine-tuning reference / training baselines validation 200 Hyperparameter tuning test 500 Final benchmark — use this for fair comparisons Total 1,200 Quick… See the full description on the dataset page: https://huggingface.co/datasets/sohaibdevv/llm-eval-benchmark.textquestion-answering1K<n<10K0 likes62 downloads4mo agoHugging Face22mteb /llm-eval-sts22_v2text1K<n<10K0 likes57 downloads6mo agoHugging Face23mteb /llm-eval-fquadtextn<1K0 likes57 downloads4mo agoHugging Face24mteb /llm-eval-biossestextn<1K0 likes56 downloads7mo agoHugging Face25mteb /llm-eval-legalbench-consumer-contractstextn<1K0 likes53 downloads4mo agoHugging Face26vector-institute /llm-eval-requeststabularn<1K0 likes52 downloads2y agoHugging Face27mteb /llm-eval-sickrtextn<1K0 likes52 downloads6mo agoHugging Face28mteb /llm-eval-sprint_duplicate_questionstextn<1K0 likes51 downloads7mo agoHugging Face29mteb /llm-eval-twitter-hjernetextn<1K0 likes51 downloads6mo agoHugging Face30mteb /llm-eval-sts17text1K<n<10K0 likes50 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.