CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01google /deepsearchqa DeepSearchQA A 900-prompt factuality benchmark from Google DeepMind, designed to evaluate agents on difficult multi-step information-seeking tasks across 17 different fields. ▶ Google DeepMind Release Blog Post▶ DeepSearchQA Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark DeepSearchQA is a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional… See the full description on the dataset page: https://huggingface.co/datasets/google/deepsearchqa.textquestion-answeringn<1K132 likes24k downloads9mo agoHugging Face02google /frames-benchmark FRAMES: Factuality, Retrieval, And reasoning MEasurement Set FRAMES is a comprehensive evaluation dataset designed to test the capabilities of Retrieval-Augmented Generation (RAG) systems across factuality, retrieval accuracy, and reasoning. Our paper with details and experiments is available on arXiv: https://arxiv.org/abs/2409.12941. Dataset Overview 824 challenging multi-hop questions requiring information from 2-15 Wikipedia articles Questions span diverse topics… See the full description on the dataset page: https://huggingface.co/datasets/google/frames-benchmark.texttext-classificationn<1K266 likes9.7k downloads2y agoHugging Face03google /Synthetic-Persona-Chat Dataset Card for SPC: Synthetic-Persona-Chat Dataset Abstract from the paper introducing this dataset: High-quality conversational datasets are essential for developing AI models that can communicate with users. One way to foster deeper interactions between a chatbot and its user is through personas, aspects of the user's character that provide insights into their personality, motivations, and behaviors. Training Natural Language Processing (NLP) models on a diverse and… See the full description on the dataset page: https://huggingface.co/datasets/google/Synthetic-Persona-Chat.text10K<n<100K139 likes4.3k downloads3y agoHugging Face04google /simpleqa-verified SimpleQA Verified A 1,000-prompt factuality benchmark from Google DeepMind and Google Research, designed to reliably evaluate LLM parametric knowledge. ▶ SimpleQA Verified Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code Benchmark SimpleQA Verified is a 1,000-prompt benchmark for reliably evaluating Large Language Models (LLMs) on short-form factuality and parametric knowledge. The authors from Google DeepMind and Google Research… See the full description on the dataset page: https://huggingface.co/datasets/google/simpleqa-verified.textquestion-answering1K<n<10K52 likes3.5k downloads7mo agoHugging Face05google /MusicCaps Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g., "A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.tabulartext-to-speech1K<n<10K153 likes1.5k downloads4y agoHugging Face06google /FACTS-grounding-public FACTS Grounding 1.0 Public Examples 860 public FACTS Grounding examples from Google DeepMind and Google Research FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding. ▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post Usage The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.textquestion-answeringn<1K47 likes1.4k downloads2y agoHugging Face07GotThatData /kraken-trading-data 📈 Kraken Trading Data Collection Overview High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis. This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling. 📊 Included Trading Pairs Pair Asset Base Currency Typical Daily Volume XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.tabular10K<n<100K6 likes881 downloads8mo agoHugging Face08Salesteq /arabic-dialects-gold20 arabic-dialects-gold20 660 sentences: 33 Arabic lects × 20 sentences, each with fully diacritized dialectal orthography, the undiacritized surface form, gold IPA, an engine draft, an English gloss, machine-verified phonetic feature tags, per-row verification metadata, and notes citing the dialectological literature that grounds the row. Columns (TSV, UTF-8, one file per lect): id, sentence, raw, ipa, ipa_o2i, gloss_en, features, notes, fable_corrections, verification… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20.texttext-to-speechn<1K0 likes834 downloads2mo agoHugging Face09google /WikiProfile WikiProfile WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances. Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.tabularquestion-answering1K<n<10K20 likes466 downloads3mo agoHugging Face10aurman /GoogleTrendArchive Google Trend Archive: Global Real-Time Search Trends (2024-2026) Dataset Details Dataset Description This dataset contains over 10.2 million trending search instances from Google's Trending Now feature, collected continuously from November 28, 2024 to May 17, 2026 across all available geographic locations (200+ countries/regions). Unlike aggregated retrospective tools like Google Trends, Trending Now captures search queries experiencing real-time… See the full description on the dataset page: https://huggingface.co/datasets/aurman/GoogleTrendArchive.tabulartext-classification10M<n<100M5 likes446 downloads4mo agoHugging Face11RIPS-Goog-23 /IIT-CDIP-CSVtext10K<n<100K3 likes411 downloads3y agoHugging Face12BrightData /Goodreads-Books Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly. Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices. For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.tabulartext-classification1M<n<10M20 likes377 downloads2y agoHugging Face13namanvats /harbor-goose-openhands-benchmark Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro Trial-level results from a small controlled study comparing two agent harnesses — Goose and OpenHands-SDK — on a frozen 40-task Harbor Terminal-Bench-Pro slice. All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend. Key Findings Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.tabularn<1K3 likes348 downloads5mo agoHugging Face14toolathon123 /data-govt-nz-mirror data.govt.nz — Mirror Catalogue (hourly snapshot) Mirror publication of the New Zealand open government data catalogue (data.govt.nz). Each row of catalog.csv is a dataset record as harvested from the data.govt.nz CKAN instance (national agencies and local councils). License declaration This mirror catalogue is published under the Creative Commons Attribution 4.0 International (CC BY 4.0) licence. Individual dataset records reference their own source licence in… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/data-govt-nz-mirror.text1K<n<10K0 likes344 downloads1mo agoHugging Face15UniqueData /messengers-reviews-google-play Reviews on Messengers Dataset - Review dataset The Reviews on Messengers Dataset is a comprehensive collection of 200 the most recent customer reviews on 6 messengers obtained from the popular app store, Google Play. See the list of the apps below. This dataset encompasses reviews written in 5 different languages: English, French, German, Italian, Japanese. 💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/messengers-reviews-google-play.tabulartext-classification1K<n<10K3 likes343 downloads1y agoHugging Face16mrm8488 /goemotions GoEmotions GoEmotions is a corpus of 58k carefully curated comments extracted from Reddit, with human annotations to 27 emotion categories or Neutral. Number of examples: 58,009. Number of labels: 27 + Neutral. Maximum sequence length in training and evaluation datasets: 30. On top of the raw data, we also include a version filtered based on reter-agreement, which contains a train/test/validation split: Size of training dataset: 43,410. Size of test dataset: 5,427. Size of… See the full description on the dataset page: https://huggingface.co/datasets/mrm8488/goemotions.tabular100K<n<1M11 likes317 downloads5y agoHugging Face17gopalkalpande /bbc-news-summary About Dataset Context Text summarization is a way to condense the large amount of information into a concise form by the process of selection of important information and discarding unimportant and redundant information. With the amount of textual information present in the world wide web the area of text summarization is becoming very important. The extractive summarization is the one where the exact sentences present in the document are used as summaries. The extractive… See the full description on the dataset page: https://huggingface.co/datasets/gopalkalpande/bbc-news-summary.text1K<n<10K23 likes300 downloads4y agoHugging Face18netop /gotsf-ds 📶 Beam-Level (5G) Time-Series Dataset 📚 Citation This dataset is released alongside the following paper: Fechete, L., et al. “Goal-Oriented Time-Series Forecasting: Foundation Framework Design.” Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Singapore. If you use this dataset, please cite the above work. This dataset introduces a novel multivariate time series specifically curated to support research in enabling accurate prediction of KPIs… See the full description on the dataset page: https://huggingface.co/datasets/netop/gotsf-ds.tabulartime-series-forecasting1K<n<10K13 likes297 downloads9mo agoHugging Face19NoeFlandre /landuse-sentence-relevance-golden-human-set Land-use sentence relevance golden human set This release contains the final 300-row V3 benchmark in English plus one parallel CSV for each of the 84 non-English project-provided sat-3l-sm language codes. There are 85 language files in total. Files Every file is at data/translations/<iso>/v3-final-<iso>.csv. The nine columns are: sentence, label, polygon_name, h3_cell, latitude, longitude, source, region, source_url. The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.tabulartext-classification10K<n<100K0 likes253 downloads1d agoHugging Face20dmariaa70 /GO-MO GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis This is the official dataset repository for the GO-MO traffic dataset. The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain). GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.tabulartime-series-forecasting1B<n<10B0 likes250 downloads3mo agoHugging Face21smartduketech /indian-government-schemes-2025 Indian Government Schemes Dataset 2026 Dataset Description The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields. Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India This dataset powers SchemeFit — India's government scheme finder for citizens and businesses. What Makes This Different Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.tabulartext-classification1K<n<10K0 likes245 downloads3mo agoHugging Face22toolathon123 /oceania-gov-open-data-catalog Oceania Government Open Data — Combined Catalogue (hourly snapshot) Combined regional catalogue of Oceania (Australia + New Zealand) public-service open data harvested from both data.gov.au and data.govt.nz portals, including state, territory and local-council publishers. License declaration License: other (see below). Records in this catalogue inherit the licence of their source dataset. Where the source declares a standard open licence the record is tagged with… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/oceania-gov-open-data-catalog.text1K<n<10K0 likes227 downloads1mo agoHugging Face23dotfantasy /godot-codetext10K<n<100K8 likes223 downloads3y agoHugging Face24Eitanli /goodreadsDataset Card for "goodreads" Must-read books summary Features: Book - Name of the book. Soemtimes this includes the details of the Series it belongs to inside a parenthesis. This information can be further extracted to analyse only series. Author - Name of the book's Author Description - The book's description as mentioned on Goodreads Genres - Multiple Genres as classified on Goodreads. Could be useful for Multi-label classification or Content based recommendation and Clustering. Average… See the full description on the dataset page: https://huggingface.co/datasets/Eitanli/goodreads.tabular10K<n<100K9 likes201 downloads3y agoHugging Face25Salesteq /arabic-dialects-gold20-code-switch gold20-code-switch Code-switched Arabic sentences with IPA: 20 rows per lect across 33 Arabic lects (the same roster as the sibling TigreGotico/arabic-dialects-gold20). Each row embeds foreign material in a dialectal Arabic frame: inline Latin-script English (and French, for the lects whose live contact language is French), Arabic-script loanwords (سيرفس، كاش، موبايل-class), and Arabizi (Latin-written Arabic with digit gutturals). Columns (TSV, UTF-8, one file per lect): id… See the full description on the dataset page: https://huggingface.co/datasets/Salesteq/arabic-dialects-gold20-code-switch.texttext-to-speechn<1K0 likes184 downloads2mo agoHugging Face26lara-popovic /google_play_store_reviewstabularn<1K0 likes180 downloads2mo agoHugging Face27double-blind-anonymous /go-mo-dataset GO-MO, a massive Graph agumented Open urban MObility dataset This is the official dataset repository for the GO-MO traffic dataset. The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain). GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024). Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.tabulartime-series-forecasting1B<n<10B0 likes160 downloads8mo agoHugging Face28pszemraj /govreport-summarization-8192 GovReport Summarization - 8192 tokens ccdv/govreport-summarization with the changes of: data cleaned with the clean-text python package total tokens for each column computed and added in new columns according to the long-t5 tokenizer (done after cleaning) train info RangeIndex: 8200 entries, 0 to 8199 Data columns (total 4 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 report 8200 non-null… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/govreport-summarization-8192.tabularsummarization10K<n<100K3 likes155 downloads9mo agoHugging Face29google /granola-entity-questions GRANOLA Entity Questions Dataset Card Dataset details Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions) Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.tabularquestion-answering10K<n<100K12 likes144 downloads2y agoHugging Face30theonegareth /antam_historical_gold_prices Unofficial Antam gold price history (IDR) Antam gold selling prices in Indonesian rupiah per gram, compiled from the public price chart on the official Antam Logam Mulia site. 5,156 records covering 2010-01-04 through 2026-08-14. This is an unofficial compilation for research, analysis and teaching. See the disclaimer at the end before you rely on it for anything else. Files Antam_historical_gold_prices.csv is the one to use: Column Type Meaning Time… See the full description on the dataset page: https://huggingface.co/datasets/theonegareth/antam_historical_gold_prices.tabular10K<n<100K2 likes136 downloads10d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.