CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes25k downloads9mo agoHugging Face02google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K266 likes9.7k downloads2y agoHugging Face03google /simpleqa-verified SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.textquestion-answering1K<n<10K53 likes3.3k downloads7mo agoHugging Face04google /FACTS-grounding-public FACTS Grounding 1.0 Public Examples 860 public FACTS Grounding examples from Google DeepMind and Google Research FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding. ▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post Usage The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.textquestion-answeringn<1K47 likes1.3k downloads2y agoHugging Face05google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes467 downloads3mo agoHugging Face06smartduketech /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes217 downloads3mo agoHugging Face07google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes140 downloads2y agoHugging Face08philosopher-from-god /ChatGPT-Jailbreak-Prompts-rubend18 Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K2 likes73 downloads1y agoHugging Face09Mahadih534 /Global_Environment-Social-And-Governance-Data Global_Environment-Social-And-Governance Dataset This Dataset contains all verified and authorized Environment, Social and Governance Statistics data in the World Description I have collected all data from WORLD-Bank's Data Catalog and also shared this link in the data source section, this dataset is sutitable for various NLP tasks Data Source https://datacatalog.worldbank.org/ Dataset Card Authors Mahadi Hassan Dataset Card Contact… See the full description on the dataset page: https://huggingface.co/datasets/Mahadih534/Global_Environment-Social-And-Governance-Data.tabularquestion-answering10K<n<100K1 likes53 downloads2y agoHugging Face10google /revealgated Reveal: A Benchmark for Verifiers of Reasoning Chains Paper: A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains Link: https://arxiv.org/abs/2402.00559 Website: https://reveal-dataset.github.io/ Abstract: Prompting language models to provide step-by-step answers (e.g., "Chain-of-Thought") is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task… See the full description on the dataset page: https://huggingface.co/datasets/google/reveal.tabulartext-classification1K<n<10K38 likes46 downloads2y agoHugging Face11Yokey20 /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Yokey20/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes38 downloads7d agoHugging Face12go-inoue /ArabicMMLU_full Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_full.tabularquestion-answering10K<n<100K0 likes31 downloads8mo agoHugging Face13MikhailVyrodov /goattextquestion-answering1K<n<10K0 likes25 downloads2y agoHugging Face14gorges-haha /Bitext-customer-support-llm-chatbot-training-dataset Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants Overview This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/gorges-haha/Bitext-customer-support-llm-chatbot-training-dataset.textquestion-answering10K<n<100K0 likes25 downloads2mo agoHugging Face15google /TACTgated TACT: A Complex Numerical Reasoning Benchmark Paper - TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools Website: https://tact-benchmark.github.io Abstract: Large Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts. To better evaluate this setting and facilitate modeling efforts, we introduce TACT - Text And Calculations through Tables, a dataset crafted to evaluate LLMs'… See the full description on the dataset page: https://huggingface.co/datasets/google/TACT.tabularquestion-answeringn<1K10 likes21 downloads2y agoHugging Face16lawful-good-project /dataset-qa-ip-lawДатасет для оценки производительности большой языковой модели. Контрибьюторы (в алфавитном порядке): Ася Айнбунд Дарья Анисимова Юрий Батраков Арсений Батуев Егор Батурин Андрей Бочков Дмитрий Данилов Максим Долотин Алексей Дружинин Константин Евменов Лолита Князева Владимир Королев Антон Костин Ярослав Котов Сергей Лагутин Иван Литвак Илья Лопатин Татьяна Максиян Артур Маликов Александр Медведев Михаил Кирилл Пантелеев Александр Панюков Алексей Суслов Даниэль Торен Данила Хайдуков… See the full description on the dataset page: https://huggingface.co/datasets/lawful-good-project/dataset-qa-ip-law.textquestion-answeringn<1K1 likes20 downloads3y agoHugging Face17siva0072 /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/siva0072/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes20 downloads2mo agoHugging Face18Sumna /indian-government-schemes-2025gated Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/Sumna/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes20 downloads5d agoHugging Face19relai-ai /go-reasoningSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: Go Documentation Data Source Link: https://go.dev/doc/ Data Source License: https://go.dev/LICENSE Data Source Authors: Go Contributors AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes19 downloads1y agoHugging Face20jun000 /MS_MARCO_Gold_Passage_QA How to Make Make "Question has only one answer". If Question has multiple answer, that question is deleted. Example Usage from datasets import load_dataset docs = load_dataset('jun000/MS_MARCO_Gold_Passage_QA') textquestion-answering100K<n<1M1 likes17 downloads2y agoHugging Face21arzumanabbasov /azerbaijani-gov-qa Azerbaijani Government Services Question Answering Dataset Overview This dataset contains over 5000 samples of question-answer pairs scraped from the comments section of the Instagram page of AsanXidmat, a government organization in Azerbaijan dedicated to providing services to Azerbaijani citizens. The dataset is intended for use in training and evaluating question answering systems, particularly those focused on understanding and responding to inquiries related to… See the full description on the dataset page: https://huggingface.co/datasets/arzumanabbasov/azerbaijani-gov-qa.textquestion-answering1K<n<10K1 likes16 downloads2y agoHugging Face22GodwillN /power-seeking-eval-300gated Power-Seeking Evaluation Dataset A 300-item multiple-choice benchmark for power-seeking in language models: the disposition to prefer options that increase the model's resources, autonomy, influence, or freedom from oversight, in situations where a lower-power option would serve the stated task equally well. Model-written, following Perez et al., "Discovering Language Model Behaviors with Model-Written Evaluations". Built for the ARENA LLM evaluations curriculum. This is the… See the full description on the dataset page: https://huggingface.co/datasets/GodwillN/power-seeking-eval-300.textquestion-answeringn<1K1 likes13 downloads1mo agoHugging Face23go-inoue /ArabicMMLU_undiac Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha Sengupta, Shady Shehata, Nizar Habash, Preslav Nakov, and Timothy Baldwin MBZUAI, Prince Sattam bin Abdulaziz University, KFUPM, Core42, NYU Abu Dhabi, The University of Melbourne Introduction We present ArabicMMLU, the first multi-task language understanding benchmark for Arabic language, sourced from school exams across diverse… See the full description on the dataset page: https://huggingface.co/datasets/go-inoue/ArabicMMLU_undiac.tabularquestion-answering10K<n<100K0 likes12 downloads1y agoHugging Face24relai-ai /go-standardSamples in this benchmark were generated by RELAI using the following data source(s): Data Source Name: Go Documentation Data Source Link: https://go.dev/doc/ Data Source License: https://go.dev/LICENSE Data Source Authors: Go Contributors AI Benchmarks by Data Agents © 2025 RELAI.AI · Licensed under CC BY 4.0. Source: https://relai.ai textquestion-answering1K<n<10K0 likes11 downloads1y agoHugging Face25SAINTHALF /kanna-rag-gold-standard Kanna RAG Gold Standard Dataset This dataset contains 30 expert-curated Question-Answer pairs focused on the ethnopharmacology of Sceletium tortuosum (Kanna). It serves as the "Gold Standard" evaluation set for the LAYRA (Large Academic Visual RAG Agent) thesis project. Dataset Structure query: The scientific question. doc_id: The unique identifier of the source document (PDF). page_num: The specific page number where the answer is found (critical for Visual RAG).… See the full description on the dataset page: https://huggingface.co/datasets/SAINTHALF/kanna-rag-gold-standard.textquestion-answeringn<1K0 likes7 downloads9mo agoHugging Face26har123ish /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/har123ish/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes7 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.