datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ai-model-popularity
Datamata AI Model Popularity Index
Weekly popularity of the most-downloaded and trending Hugging Face models: trailing downloads, likes, the model's task and its trending rank. One row per model from the most recent weekly snapshot.
Latest snapshot: 2026-09-20
Models in this release: 50
Updated: weekly
Licence: CC BY 4.0 — free to use and adapt, including commercially, with attribution.
Source & methodology: https://www.datamatastudios.com/datasets
Quickstart… See the full description on the dataset page: https://huggingface.co/datasets/datamatastudios/ai-model-popularity.StackV1-popular
Stack V1 with popular programming languages
Javascript
Python
C
C++
SQL
Cuda
passages_gutenberg_popularPopularNovelsspotify_popular_trackspopular-tokenizerschub_popular_charactersdaily-papers-popularity
From the Frontier Research Team at Takara.ai, we present Daily Papers Popularity — a dataset tracking the popularity of Hugging Face Papers with arXiv metadata. It aggregates daily paper entries with votes, IDs, titles, abstracts (backfilled via the HF API), and URLs, enabling analysis of patterns in paper reception and engagement.
Daily Papers Popularity
Columns: date, arxiv_id, votes, title, abstract, url
Format: Parquet
Load
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/takara-ai/daily-papers-popularity.12_popular_russia_mushrooms_edible_poisonouspope-coco-popularpopularity-enriched-qa-datasets
Popularity-Enriched QA Datasets
This dataset repo hosts popularity-enriched versions of PopQA, Natural Questions, and TriviaQA.
Each subset retains the enrichment schema produced by this notebook (question + pron and popularity metrics).
Subsets
pop_qa: Popularity-enriched PopQA test split
natural_questions: Wikipedia-provenance Natural Questions validation set
trivia_qa: TriviaQA validation subset matched to KILT and original TriviaQA IDs
hotpot_qa: HotPotQA validation… See the full description on the dataset page: https://huggingface.co/datasets/Cyro1/popularity-enriched-qa-datasets.AudioHallucination_AudioCaps-Popularpopular-street-ac3873
popular-street-ac3873
Synthetic products test data: 40 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/naoki72/popular-street-ac3873.entity_popularity
Entity Popularity Dataset
This dataset contains information for about 26,000 entities, including the Wikipedia article title, QID, and the annual article view count for the year 2021.
The annual article view count can be considered as an indicator of the popularity of a entity.
Languages
This dataset is composed in English.
Dataset Structure
from datasets import load_dataset
dataset = load_dataset("masaki-sakata/entity_popularity")["en"]
print(dataset)
#… See the full description on the dataset page: https://huggingface.co/datasets/masaki-sakata/entity_popularity.youtube_top_popular_videos_commentssinhala-comment-popularity-predictionpopular_names_spell_correctionSouth_African_Popular_Surnames
SA Popular Surnames Dataset
Authors: Minah Mojela (@minahmojela) and Tadiwa Tine, Umkho-AI
Dataset Summary
This dataset contains 651 structured records documenting popular South African
surnames — their linguistic/cultural group, clan or lineage context, meaning and
etymology, and notable historical bearers — across 10 South African population
groups. It is a companion release to the South African History Dataset
and the SA Tribal & Cultural Practices Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Umkho-AI/South_African_Popular_Surnames.popular_english_wordspopular_tracks
Spotify-Africa Popular Tracks | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/popular_tracks.Manga-Popularis-EN
Dataset Card: Manga-Popularis-EN
Dataset Details
Dataset Description
This dataset contains information about art books, including titles, authors, descriptions, image paths, prices, publication dates, publishers, ISBN numbers, and sizes.
License: GPL-3.0
Dataset Sources
The dataset was compiled from various sources, including online bookstores, publisher websites, and catalogs. On date 4/21/24.
Dataset Structure
The dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/Tsunnami/Manga-Popularis-EN.popular_artists
Spotify-Africa Popular Artists | Africa (Electric Sheep Africa metadata inventory)
Size category: n<1K - Formats: parquet - Sector: culture_language - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/popular_artists.laurashin_Why_CoinFund_Believes_Worldcoin_Could_Become_More_Popular_Than_Bitcoin_-_Ep__525models-text-generation-popular-PRIVATEpdf_exp_train_data__pdf_sci_qa_unverified_filtered_popular_urlspopular-cardspopular-posts-3d-scatter-plotPopularMovieDataset
Movie Data Scraper
⚠️ SOME DATA IS MESSED UP ⚠️
some of the colums aren't properly registered with huggingface, so be careful with using this data.
Overview
This Project Scraped movies from 1990-2003 (due to api limitations) and has popular movies.
Data Structure
The resulting Data file will show the following columns:
Title: The title of the movie.
Year: The release year of the movie.
Genre: The genre(s) of the movie.
Director: The director of the… See the full description on the dataset page: https://huggingface.co/datasets/ZelonPrograms/PopularMovieDataset.popular_anime_corpus
popular_anime_corpus
有名なアニメのあらすじや主人公の情報をまとめたコーパスです。
1960年代以降の人気アニメに関するテキストをWikipediaから抽出し、テキストを要約して、Alpaca形式でコーパスにまとめています。下記のようなデータとなっています。
{
"train": [
{
"instruction": "主人公は誰ですか?",
"input": "名探偵コナン 絶海の探偵",
"output": "主人公は「江戸川 コナン(えどがわ コナン)」です。\n\n江戸川 コナンは、本来の姿は「東の高校生探偵」として名を馳せている工藤新一だが、黒ずくめの組織に飲まされた毒薬・APTX4869の副作用で小学生の姿になっている。本作の主人公であり、物語の中心人物として活躍します。"},
{
"instruction": "主人公やあらすじを教えてください。",
"input": "ヴァイオレット・エヴァーガーデン",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/kujirahand/popular_anime_corpus.dataset_popularity
