CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01openai /MMMLU Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science. We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.textquestion-answering100K<n<1M526 likes12k downloads2y agoHugging Face02bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face03databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K26 likes9.1k downloads2mo agoHugging Face04ArtificialAnalysis /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.documentquestion-answeringn<1K50 likes7.9k downloads1mo agoHugging Face05databricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes2.6k downloads2mo agoHugging Face06matichon /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.textquestion-answering1K<n<10K0 likes2.2k downloads5mo agoHugging Face07openthaigpt /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.textquestion-answering1K<n<10K8 likes1.6k downloads3y agoHugging Face08bench-llms /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench.imagetext-generation10K<n<100K1 likes742 downloads2y agoHugging Face09Chenyu-Zhou /OR-Space OR-Space A full-lifecycle workspace benchmark for industrial optimization agents. OR-Space evaluates whether language-model agents can work reliably with operations research problems represented as executable, multi-file workspaces. Rather than presenting a self-contained mathematical prompt, each task distributes evidence across business requirements, structured data, source code, execution logs, and solver records. The benchmark contains 100 optimization topologies. Each… See the full description on the dataset page: https://huggingface.co/datasets/Chenyu-Zhou/OR-Space.textquestion-answeringn<1K4 likes629 downloads2mo agoHugging Face10orbench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our leaderboard at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/orbench-llm/or-bench.imagetext-generation10K<n<100K0 likes615 downloads2y agoHugging Face11bench-llms /or-bench-toxic-all OR-Bench: An Over-Refusal Benchmark for Large Language Models This dataset constains highly toxic prompts, use with caution!!! Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic… See the full description on the dataset page: https://huggingface.co/datasets/bench-llms/or-bench-toxic-all.imagetext-generation10K<n<100K1 likes357 downloads2y agoHugging Face12PersonaBias /Reverse-alpha-beta-no-outsidetabulartext-classification10K<n<100K0 likes291 downloads2mo agoHugging Face13MaYiding /OracleProto OracleProto: Forecasting Evaluation Set Chinese doc: [中文文档] GitHub repo: [MaYiding/OracleProto] Visit Our Leaderboards: [Website] View Our Paper: [arXiv] A SQLite-packaged evaluation set of 80 hand-curated forecasting questions on real-world events, with resolution dates between 2026-03-12 and 2026-04-14, released alongside the GitHub Repo. Both the rows and the byte-stable prompt-reconstruction recipe are packaged in a single file, forecast_eval_set_example.db, which exposes two… See the full description on the dataset page: https://huggingface.co/datasets/MaYiding/OracleProto.textquestion-answeringn<1K3 likes260 downloads5mo agoHugging Face14PersonaBias /Original-alpha-suppression-task-boosttabulartext-classification100K<n<1M0 likes203 downloads2mo agoHugging Face15OpenLab-NLP /tiny-singleturn-chat-kotextquestion-answering10K<n<100K0 likes197 downloads10mo agoHugging Face16carlomarxx /trilemma-of-truth Dataset Card for Trilemma of Truth (ToT) Dataset 🧾 Dataset Summary The Trilemma of Truth (ToT) dataset serves as a benchmark for evaluating veracity probes across three distinct statement types: Factually true statements. Factually false statements. Neither-valued statements are defined as those for which the language model lacks sufficient evidence to assign a truth value (see formal definition below). The dataset includes three domain configurations:… See the full description on the dataset page: https://huggingface.co/datasets/carlomarxx/trilemma-of-truth.texttext-classification10K<n<100K2 likes191 downloads2mo agoHugging Face17FrankPN /OmniBrainBench OmniBrainBench 🍎 Homepage|💻 GitHub|🤗 Dataset|📖 Paper This repository is the official implementation of the paper [OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks]. 🚀 News [02/2026] Our OmniBrainBench is accepted by CVPR2026! [12/2025] We have released the evaluation code and dataset for OmniBrainBench. [11/2025] The manuscript can be found on arXiv. 🚀Overview we introduce… See the full description on the dataset page: https://huggingface.co/datasets/FrankPN/OmniBrainBench.imagequestion-answering1K<n<10K3 likes188 downloads6mo agoHugging Face18PersonaBias /Original-no-persona-replacement-remaindertabulartext-classification100K<n<1M0 likes166 downloads2mo agoHugging Face19agentic-learning-ai-lab /daily-oracle Daily Oracle 📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time. Dataset Details Question Type: True/False (TF) & Multiple Choice (MC) Current Version* Time Span: 2020.01.01 - 2026.07.18 Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.textquestion-answering10K<n<100K4 likes156 downloads2mo agoHugging Face20PersonaBias /Original-hybrid-shared-no-persona-remaindertabulartext-classification100K<n<1M0 likes146 downloads2mo agoHugging Face21PersonaBias /Original-hybrid-correct-train-no-persona-meantabulartext-classification100K<n<1M0 likes125 downloads2mo agoHugging Face22AAAIBenchmark /Multi-Opthalingua Cite Accepted to AAAI 2025 (https://openreview.net/group?id=AAAI.org/2025/Conference#tab-recent-activity) Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs: @misc{restrepo2024multiophthalinguamultilingualbenchmarkassessing, title={Multi-OphthaLingua: A Multilingual Benchmark for Assessing and Debiasing LLM Ophthalmological QA in LMICs}, author={David Restrepo and Chenwei Wu and Zhengxu Tang and Zitao Shuai and Thao… See the full description on the dataset page: https://huggingface.co/datasets/AAAIBenchmark/Multi-Opthalingua.tabularquestion-answeringn<1K3 likes123 downloads2y agoHugging Face23jerogo /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue… See the full description on the dataset page: https://huggingface.co/datasets/jerogo/or-bench.imagetext-generation10K<n<100K0 likes123 downloads2mo agoHugging Face24PersonaBias /Original-hybrid-train-no-persona-meantabulartext-classification100K<n<1M0 likes118 downloads2mo agoHugging Face25PersonaBias /Original-circuit-discoverytabulartext-classification10K<n<100K0 likes117 downloads2mo agoHugging Face26Mahadih534 /Institutional-Information-of-Bangladesh Institutional-Information-of-Bangladesh Dataset This Dataset contains all verified and authorized Institutional information in Bangladesh Description I have collected all data from bangladeshi government authorized web portal and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source http://data.gov.bd/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Institutional-Information-of-Bangladesh.tabularquestion-answering10K<n<100K2 likes112 downloads2y agoHugging Face27openthaigpt /thai-investment-consultant-licensing-exams Thai Public Investment Consultant (IC) Exams Dataset Overview This dataset comprises a collection of exam questions and answers from the Thai Public Investment Consultant (IC) Examinations. It's a valuable resource for developing and evaluating question-answering systems in the finance sector. Dataset Source The Stock Exchange of Thailand (SET) Maintainer Dr. Kobkrit Viriyayudhakorn Email: kobkrit@iapp.co.th Dataset Description This… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-investment-consultant-licensing-exams.tabularquestion-answeringn<1K6 likes111 downloads3y agoHugging Face28Sandhya1912 /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/Sandhya1912/AA-Omniscience-Public.documentquestion-answeringn<1K0 likes110 downloads4mo agoHugging Face29Fancy-MLLM /R1-Onevision-Bench R1-Onevision-Bench [📂 GitHub][📝 Paper] [🤗 HF Dataset] [🤗 HF Model] [🤗 HF Demo] Dataset Overview R1-Onevision-Bench comprises 38 subcategories organized into 5 major domains, including Math, Biology, Chemistry, Physics, Deducation. Additionally, the tasks are categorized into five levels of difficulty, ranging from ‘Junior High School’ to ‘Social Test’ challenges, ensuring a comprehensive evaluation of model capabilities across varying complexities.… See the full description on the dataset page: https://huggingface.co/datasets/Fancy-MLLM/R1-Onevision-Bench.textquestion-answeringn<1K3 likes106 downloads2y agoHugging Face30PersonaBias /Original-baseline-bias-unbiastabulartext-classification100K<n<1M0 likes94 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.