CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes23k downloads9mo agoHugging Face02google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K266 likes9.7k downloads2y agoHugging Face03google /Synthetic-Persona-Chat Dataset Card for SPC: Synthetic-Persona-Chat Dataset Abstract from the paper introducing this dataset: High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.text10K<n<100K139 likes4k downloads3y agoHugging Face04google /simpleqa-verified SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.textquestion-answering1K<n<10K52 likes3.4k downloads7mo agoHugging Face05google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K153 likes1.5k downloads4y agoHugging Face06google /FACTS-grounding-public FACTS Grounding 1.0 Public Examples 860 public FACTS Grounding examples from Google DeepMind and Google Research FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding. ▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post Usage The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.textquestion-answeringn<1K47 likes1.4k downloads2y agoHugging Face07google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes455 downloads3mo agoHugging Face08aurman /GoogleTrendArchive Google Trend Archive: Global Real-Time Search Trends (2024-2026) Dataset Details Dataset Description This dataset contains over 10.2 million trending search instances from Google's Trending Now feature, collected continuously from November 28, 2024 to May 17, 2026 across all available geographic locations (200+ countries/regions). Unlike aggregated retrospective tools like Google Trends, Trending Now captures search queries experiencing real-time… See the full description on the dataset page: https://huggingface.co/datasets/aurman/GoogleTrendArchive.tabulartext-classification10M<n<100M5 likes454 downloads4mo agoHugging Face09UniqueData /messengers-reviews-google-play Reviews on Messengers Dataset - Review dataset The Reviews on Messengers Dataset is a comprehensive collection of 200 the most recent customer reviews on 6 messengers obtained from the popular app store, Google Play. See the list of the apps below. This dataset encompasses reviews written in 5 different languages: English, French, German, Italian, Japanese. 💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/messengers-reviews-google-play.tabulartext-classification1K<n<10K3 likes342 downloads1y agoHugging Face10lara-popovic /google_play_store_reviewstabularn<1K0 likes179 downloads2mo agoHugging Face11google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes143 downloads2y agoHugging Face12najeh-halawani /google-ads-transparencytabular10M<n<100M0 likes134 downloads2d agoHugging Face13megloughney /googleAnalyticsCustomerRevenuePredictiontabular10K<n<100K0 likes128 downloads9mo agoHugging Face14opdullah /turkish-google-maps-15M Turkish Google Maps Reviews Bu veri seti, Türkiye’deki işletmelere ait Türkçe Google Maps yorumlarını içerir. Her kayıt: yorum metni yorum puanı işletme adı işletme kategorisi gibi bilgileri içerir. Veri seti, özellikle büyük ölçekli Türkçe NLP çalışmaları için uygundur. Contents Veri setinde aşağıdaki türde alanlar bulunmaktadır: yorum metni (review_text) yorum puanı (rating) işletme adı (place_name) işletme kategorisi (category) kategori listesi (category_list)… See the full description on the dataset page: https://huggingface.co/datasets/opdullah/turkish-google-maps-15M.tabulartext-classification10M<n<100M12 likes93 downloads6mo agoHugging Face15Roy229 /fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001 fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001 A curated registry of points of interest in downtown Portland, Oregon. License This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY). You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source. Contents data.csv - sample points of interest with coordinates… See the full description on the dataset page: https://huggingface.co/datasets/Roy229/fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testterm001.tabularn<1K0 likes81 downloads28d agoHugging Face16almogtavor /google-analogy-dataset Google Analogy Dataset (Columnized) This is a columnized version of the Google analogy dataset by Mikolov et al. (2013): https://github.com/nicholas-leonard/word2vec/blob/master/questions-words.txt The dataset contains word analogy questions grouped by subjects such as: capital-common-countries (e.g., Athens Greece Tokyo Japan) currency (e.g., USA dollar Japan yen) gram3-comparative (e.g., big bigger cold colder) The original dataset is widely used in word embedding evaluation… See the full description on the dataset page: https://huggingface.co/datasets/almogtavor/google-analogy-dataset.text10K<n<100K2 likes80 downloads1y agoHugging Face17Agisight /google-smol-en-ru Карточка Google Smol to Russian Человеческий перевод датасета Smol от Google Translate на Русский язык от Андрея Анисимова. Вычитка от Фархада Фаткуллина и David Dalé´. Детали датасета Описание датасета Если вы хотите добавить переводы на другой язык, пожалуйста, создайте копию этого документа и выполняйте переводы в ней. Инструкции: Переведите русский (или английский) текст на ваш язык – лучше начать с файла smoldoc.csv. С Русского языка… See the full description on the dataset page: https://huggingface.co/datasets/Agisight/google-smol-en-ru.texttranslation1K<n<10K3 likes68 downloads10mo agoHugging Face18google /rfm-rm-as-user-dataset RFM Reward Model As User Dataset This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section. Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.tabulartext-generation10K<n<100K10 likes67 downloads11mo agoHugging Face19google /mittens MiTTenS: A Dataset for Evaluating Misgendering in Translation Misgendering is the act of referring to someone in a way that does not reflect their gender identity. Translation systems, including foundation models capable of translation, can produce errors that result in misgendering harms. To measure the extent of such potential harms when translating into and out of English, we introduce a dataset, MiTTenS, covering 26 languages from a variety of language families and scripts… See the full description on the dataset page: https://huggingface.co/datasets/google/mittens.tabulartranslation1K<n<10K8 likes58 downloads3y agoHugging Face20google /gemma3n-slicing-configsThis repository contains configurations to slice Gemma 3n E4B, which is enabled thanks to it being a MatFormer. The E4B model can be sliced into small models, trading off quality and latency/compute requirements. We recommend exploring the [MatFormer Lab](TODO: add link) to getting started with slicing Gemma 3n E4B yourself. For each configuration, we calculate the MMLU accuracy. Although these are not the only configurations possible, they are optimal configurations identified by calculating… See the full description on the dataset page: https://huggingface.co/datasets/google/gemma3n-slicing-configs.tabularn<1K9 likes56 downloads1y agoHugging Face21jason1966 /alhamdulliah123_google-play-store-apps-ratings-reviews Google Play Store Apps – Ratings, Reviews A comprehensive dataset to explore app performance, user ratings, installs Dataset Info Source: Kaggle Original Size: 0.02 MB Kaggle Downloads: 145 Files: 1 Files google_play_store_apps_famous.csv Mirrored from Kaggle tabular1K<n<10K0 likes56 downloads6mo agoHugging Face22justinqbui /covid_fact_checked_google_apiThis dataset was gathered from the Google Fact Checker API, using an automatic web scraper. 10,000 facts were pulled, but for the sake of simplicity, only ones were the ratings were singular words "false" or "true", were kept, which filtered it down to ~3000 fact checks, with about 90% of the facts being false. annotations_creators: expert-generated language_creators: crowdsourced languages: en-US licenses: unknown multilinguality: monolingual pretty_name: polifact-covid-fact-checker… See the full description on the dataset page: https://huggingface.co/datasets/justinqbui/covid_fact_checked_google_api.text1K<n<10K1 likes51 downloads5y agoHugging Face23zhuq41 /filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_ugvtc2 BlueWave Logistics Delivery-Point Registry This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch). Files registry.csv — the master delivery-point registry. review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst). registry.csv schema Columns: id,branch,address,city,state,status,review_month id: delivery-point identifier (e.g. DP-101). branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_ugvtc2.textn<1K0 likes49 downloads1mo agoHugging Face24google /revealgated Reveal: A Benchmark for Verifiers of Reasoning Chains Paper: A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains Link: https://arxiv.org/abs/2402.00559 Website: https://reveal-dataset.github.io/ Abstract: Prompting language models to provide step-by-step answers (e.g., "Chain-of-Thought") is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task… See the full description on the dataset page: https://huggingface.co/datasets/google/reveal.tabulartext-classification1K<n<10K38 likes48 downloads2y agoHugging Face25zhuq41 /filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4 BlueWave Logistics Delivery-Point Registry This dataset maintains the delivery-point registry for BlueWave Logistics (regional freight & dispatch). Files registry.csv — the master delivery-point registry. review_decisions.csv — the latest Q3 2026 review report (published by the operations analyst). registry.csv schema Columns: id,branch,address,city,state,status,review_month id: delivery-point identifier (e.g. DP-101). branch: operations branch… See the full description on the dataset page: https://huggingface.co/datasets/zhuq41/filesystem_fetch_hf_playwright_googlemap_terminal_github_scholarly_8016_bwdelreg_1zytb4.textn<1K0 likes44 downloads28d agoHugging Face26Roy229 /fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001 fetch_huggingface_google_map_terminal_github_7958-poi-downtown-portland-testrun001 A curated registry of points of interest in downtown Portland, Oregon. License This dataset is licensed under the Open Data Commons Attribution License 1.0 (ODC-BY). You are free to share, create, and adapt the data for any purpose, including commercial use, provided you give attribution to the source. Contents data.csv - sample data. tabularn<1K0 likes42 downloads29d agoHugging Face27Roy229 /google-cloud_github_fetch_huggingface_terminal_6737_bv8nvl1etabularn<1K0 likes41 downloads1mo agoHugging Face28Roy229 /google-cloud_github_fetch_huggingface_terminal_6737_308aqj4ztabularn<1K0 likes41 downloads28d agoHugging Face29suhani-sarvam /google-dakshinatext100K<n<1M2 likes40 downloads2y agoHugging Face30Roy229 /fetch_huggingface_google_map_terminal_github_7958-street-mobility-testrun001 fetch_huggingface_google_map_terminal_github_7958-street-mobility-testrun001 Street network mobility and accessibility attributes for downtown Portland. License This dataset is licensed under the Apache License 2.0. Commercial use, modification, and distribution are permitted. Contents data.csv - sample data. textn<1K0 likes40 downloads29d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.