datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WikiProfile
WikiProfile
WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances.
Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.trending-words-google
Google Trending Words Dataset (2001-2024)
Dataset Description
This dataset contains Google trending words and search terms from 2001 to 2024, capturing 24 years of internet culture, major events, and global trends. The dataset includes 2,784 entries across 93 standardized categories, providing a comprehensive view of what captured the world's attention over more than two decades.
Dataset Summary
Total Entries: 2,784
Years Covered: 2001-2024 (24 years)… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/trending-words-google.reveal
Reveal: A Benchmark for Verifiers of Reasoning Chains
Paper: A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
Link: https://arxiv.org/abs/2402.00559
Website: https://reveal-dataset.github.io/
Abstract:
Prompting language models to provide step-by-step answers (e.g., "Chain-of-Thought") is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task… See the full description on the dataset page: https://huggingface.co/datasets/google/reveal.TACT
TACT: A Complex Numerical Reasoning Benchmark
Paper - TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
Website: https://tact-benchmark.github.io
Abstract: Large Language Models (LLMs) often do not perform well on queries that require the aggregation of information across texts. To better evaluate this setting and facilitate modeling efforts, we introduce TACT - Text And Calculations through Tables, a dataset crafted to evaluate LLMs'… See the full description on the dataset page: https://huggingface.co/datasets/google/TACT.
