datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MusicCaps
Dataset Card for MusicCaps
Dataset Summary
The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g.,
"A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.kraken-trading-data
📈 Kraken Trading Data Collection
Overview
High-frequency cryptocurrency market data from Kraken exchange - perfect for algorithmic trading, time-series forecasting, and market microstructure analysis.
This dataset includes real-time price, volume, and order book data for 9 major cryptocurrency trading pairs, collected via WebSocket streaming and REST API polling.
📊 Included Trading Pairs
Pair
Asset
Base Currency
Typical Daily Volume
XXBTZUSD… See the full description on the dataset page: https://huggingface.co/datasets/GotThatData/kraken-trading-data.WikiProfile
WikiProfile
WikiProfile is a factual knowledge benchmark for evaluating how well language models encode and recall factual knowledge. It comprises 2,150 facts, each paired with 10 questions, for a total of 21,500 question instances.
Each fact is grounded in the first paragraph (summary) of an English Wikipedia page and is defined as a proposition between two entities, a subject and an object (e.g., "Oasis played their first gig at the Boardwalk club" → subject: Oasis, object:… See the full description on the dataset page: https://huggingface.co/datasets/google/WikiProfile.GoogleTrendArchive
Google Trend Archive: Global Real-Time Search Trends (2024-2026)
Dataset Details
Dataset Description
This dataset contains over 10.2 million trending search instances from Google's Trending Now feature, collected continuously from November 28, 2024 to May 17, 2026 across all available geographic locations (200+ countries/regions). Unlike aggregated retrospective tools like Google Trends, Trending Now captures search queries experiencing real-time… See the full description on the dataset page: https://huggingface.co/datasets/aurman/GoogleTrendArchive.Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books"
Dataset Summary
Explore a collection of millions of books with the Goodreads dataset, comprising over 6.3M structured records and 14 data fields updated and refreshed regularly.
Each entry includes all major data points such as URLs, book IDs, titles, authors, ratings, number of ratings, reviews, summaries, genres, publication dates, author details and prices.
For a complete list of data points, please refer to the full "Data… See the full description on the dataset page: https://huggingface.co/datasets/BrightData/Goodreads-Books.harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.messengers-reviews-google-play
Reviews on Messengers Dataset - Review dataset
The Reviews on Messengers Dataset is a comprehensive collection of 200 the most recent customer reviews on 6 messengers obtained from the popular app store, Google Play. See the list of the apps below.
This dataset encompasses reviews written in 5 different languages: English, French, German, Italian, Japanese.
💴 For Commercial Usage: To discuss your requirements, learn about the price and buy the dataset, leave a request… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/messengers-reviews-google-play.goemotions
GoEmotions
GoEmotions is a corpus of 58k carefully curated comments extracted from Reddit,
with human annotations to 27 emotion categories or Neutral.
Number of examples: 58,009.
Number of labels: 27 + Neutral.
Maximum sequence length in training and evaluation datasets: 30.
On top of the raw data, we also include a version filtered based on reter-agreement, which contains a train/test/validation split:
Size of training dataset: 43,410.
Size of test dataset: 5,427.
Size of… See the full description on the dataset page: https://huggingface.co/datasets/mrm8488/goemotions.gotsf-ds
📶 Beam-Level (5G) Time-Series Dataset
📚 Citation
This dataset is released alongside the following paper:
Fechete, L., et al. “Goal-Oriented Time-Series Forecasting: Foundation Framework Design.” Proceedings of the AAAI Conference on Artificial Intelligence, 2026, Singapore.
If you use this dataset, please cite the above work.
This dataset introduces a novel multivariate time series specifically curated to support research in enabling accurate prediction of KPIs… See the full description on the dataset page: https://huggingface.co/datasets/netop/gotsf-ds.landuse-sentence-relevance-golden-human-set
Land-use sentence relevance golden human set
This release contains the final 300-row V3 benchmark in English plus one
parallel CSV for each of the 84 non-English project-provided sat-3l-sm
language codes. There are 85 language files in total.
Files
Every file is at
data/translations/<iso>/v3-final-<iso>.csv. The nine columns are:
sentence, label, polygon_name, h3_cell, latitude, longitude,
source, region, source_url.
The Dataset Viewer exposes these files as 85… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/landuse-sentence-relevance-golden-human-set.GO-MO
GO-MO: A large-scale graph-augmented traffic dataset for data-driven spatio-temporal traffic analysis
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).… See the full description on the dataset page: https://huggingface.co/datasets/dmariaa70/GO-MO.indian-government-schemes-2025
Indian Government Schemes Dataset 2026
Dataset Description
The most comprehensive structured dataset of Indian central and state government schemes — 4,693 schemes across all ministries and states, with machine-readable eligibility fields.
Maintained by SmartDuke Technologies · Coimbatore, Tamil Nadu, India
This dataset powers SchemeFit — India's government scheme finder for citizens and businesses.
What Makes This Different
Most existing Indian… See the full description on the dataset page: https://huggingface.co/datasets/smartduketech/indian-government-schemes-2025.goodreadsDataset Card for "goodreads"
Must-read books summary
Features:
Book - Name of the book. Soemtimes this includes the details of the Series it belongs to inside a parenthesis. This information can be further extracted to analyse only series.
Author - Name of the book's Author
Description - The book's description as mentioned on Goodreads
Genres - Multiple Genres as classified on Goodreads. Could be useful for Multi-label classification or Content based recommendation and Clustering.
Average… See the full description on the dataset page: https://huggingface.co/datasets/Eitanli/goodreads.google_play_store_reviewsgo-mo-dataset
GO-MO, a massive Graph agumented Open urban MObility dataset
This is the official dataset repository for the GO-MO traffic dataset.
The GO-MO dataset is a traffic dataset extracted from the publicly available Open Data Portal of the City Council of Madrid (Spain).
GO-MO comprises more than 1.5 billion records of three traffic-related metrics together with spatio-temporal data and metadata, spanning a ten-year period (2015-2024).
Additionally, the GO-MO dataset introduces two graph… See the full description on the dataset page: https://huggingface.co/datasets/double-blind-anonymous/go-mo-dataset.govreport-summarization-8192
GovReport Summarization - 8192 tokens
ccdv/govreport-summarization with the changes of:
data cleaned with the clean-text python package
total tokens for each column computed and added in new columns according to the long-t5 tokenizer (done after cleaning)
train info
RangeIndex: 8200 entries, 0 to 8199
Data columns (total 4 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 report 8200 non-null… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/govreport-summarization-8192.granola-entity-questions
GRANOLA Entity Questions Dataset Card
Dataset details
Dataset Name: GRANOLA-EQ (Granularity of Labels Entity Questions)
Paper: Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
Abstract: Factual questions typically can be answered correctly at different levels of granularity. For example, both "August 4, 1961" and "1961" are correct answers to the question "When was Barack Obama born?"". Standard question answering (QA)… See the full description on the dataset page: https://huggingface.co/datasets/google/granola-entity-questions.antam_historical_gold_prices
Unofficial Antam gold price history (IDR)
Antam gold selling prices in Indonesian rupiah per gram, compiled from the public price chart on
the official Antam Logam Mulia site. 5,156 records covering 2010-01-04 through 2026-08-14.
This is an unofficial compilation for research, analysis and teaching. See the disclaimer at the
end before you rely on it for anything else.
Files
Antam_historical_gold_prices.csv is the one to use:
Column
Type
Meaning
Time… See the full description on the dataset page: https://huggingface.co/datasets/theonegareth/antam_historical_gold_prices.call-playbook
☎️ The Call Playbook Dataset
Real-world B2B sales conversations for text classification
A dataset by Gong.io Research
Annotated samples drawn from anonymized enterprise sales conversations across 5 binary classification tasks.
📄 Read the paper (ACL Anthology)
·
arXiv
🗂️ Dataset Summary
The Call Playbook Dataset contains annotated samples from real enterprise sales conversations across 5 binary classification tasks… See the full description on the dataset page: https://huggingface.co/datasets/gong-io-research/call-playbook.medicines_from_zakupki_gov_ruДанные для исследования существования focal points (https://www.jstor.org/stable/3132148) в гос. закупках лекарств в России.
googleAnalyticsCustomerRevenuePredictiongoogle-ads-transparencyfootball-charts-results-goal-timing
Football Charts — Match Results and Goal Timing
Football Charts publishes results, fixtures, league tables and the minute of
every goal for 93 leagues in 42 countries, including the lower divisions and
women's competitions most sources skip. Free JSON API and MCP server; the
results and goal-timing dataset is CC BY 4.0 with a DOI
(10.5281/zenodo.22295583).
Most public football datasets cover the big five European leagues and stop at
the final score. This one reaches Serie C, 3.… See the full description on the dataset page: https://huggingface.co/datasets/Damir81/football-charts-results-goal-timing.governance_2kM3D-RefSeg
Dataset Description
3D Medical Image Referring Segmentation Dataset (M3D-RefSeg),
consisting of 210 3D images, 2,778 masks, and text annotations.
Dataset Introduction
3D medical segmentation is one of the main challenges in medical image analysis. In practical applications,
a more meaningful task is referring segmentation,
where the model can segment the corresponding region based on given text descriptions.
However, referring segmentation requires image-mask-text… See the full description on the dataset page: https://huggingface.co/datasets/GoodBaiBai88/M3D-RefSeg.turkish-plu-goal-inferenceHomepage: https://github.com/GGLAB-KU/turkish-plu
turkish-google-maps-15M
Turkish Google Maps Reviews
Bu veri seti, Türkiye’deki işletmelere ait Türkçe Google Maps yorumlarını içerir.
Her kayıt:
yorum metni
yorum puanı
işletme adı
işletme kategorisi
gibi bilgileri içerir.
Veri seti, özellikle büyük ölçekli Türkçe NLP çalışmaları için uygundur.
Contents
Veri setinde aşağıdaki türde alanlar bulunmaktadır:
yorum metni (review_text)
yorum puanı (rating)
işletme adı (place_name)
işletme kategorisi (category)
kategori listesi (category_list)… See the full description on the dataset page: https://huggingface.co/datasets/opdullah/turkish-google-maps-15M.SC-train-valid-test_SDG-Descriptionsnli-label:
(0) entailment
(2) contradiction
premier-league-first-goal-impact-2025-26
What Is the First Goal Worth? — 2025/26 Premier League
Match-level data behind a 5DollarFootballAPI study of how the first confirmed goal changed
Bet365's normalized in-play win probabilities during the 2025/26 Premier League season.
Across 347 usable matches, the median within-match increase in the scoring team's normalized
win probability was 23.1 percentage points (bootstrap 95% CI: 22.4–24.2). The median
first goal after minute 75 moved the probability by 62.2 points… See the full description on the dataset page: https://huggingface.co/datasets/5dollarfootballapi/premier-league-first-goal-impact-2025-26.scifig-bench
SciFig-Bench: Scientific Figure & Multi-Scale Caption Dataset Card
Dataset Description
This dataset contains high-quality scientific figures, architecture diagrams, charts, and visualizations extracted from scientific papers (arXiv & local PDFs). It includes multi-level captions (Small, Medium, Large), extracted embedded OCR text, paper metadata, and alignment quality scores.
Total Figure Records: 589
Average CLIP Alignment Score: 0.2721
Average Composite Score:… See the full description on the dataset page: https://huggingface.co/datasets/Goutam112/scifig-bench.
