datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
media-metadata-gutenberg-books
TigreGotico/media-metadata-gutenberg-books
Rich entity dataset scraped by metadatarr
scraper gutenberg_books.
Rows: 78,738
Fields
gutenberg_id
title
authors
translators
subjects
bookshelves
languages
copyright
media_type
download_count
has_text
has_epub
entity_type
Source
Generated by scrapers/gutenberg_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-openlibrary-books
TigreGotico/media-metadata-openlibrary-books
Rich entity dataset scraped by metadatarr
scraper openlibrary_books.
Rows: 4,098,190
Fields
olid
title
subtitle
authors
author_key
first_publish_year
subjects
isbn_10
isbn_13
publisher
language
number_of_pages_median
ebook_access
has_fulltext
edition_count
cover_i
Source
Generated by scrapers/openlibrary_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-classical-composers
TigreGotico/media-metadata-classical-composers
Rich entity dataset scraped by metadatarr
scraper classical_composers.
Rows: 15,587
Fields
composer_id
name
country
life
birth
death
period
image_url
url
bio
radio_id
notable
must_know
n_recordings
n_performers
n_albums
n_works_listed
n_albums_listed
Source
Generated by scrapers/classical_composers.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-musicbrainz-artists
TigreGotico/media-metadata-musicbrainz-artists
Rich entity dataset scraped by metadatarr
scraper musicbrainz_artists.
Rows: 1,520,644
Fields
mb_id
name
sort_name
type
gender
country
area
begin_date
end_date
ended
disambiguation
aliases
tags
ipi_codes
isni_codes
Source
Generated by scrapers/musicbrainz_artists.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-librivox-audiobooks
TigreGotico/media-metadata-librivox-audiobooks
Rich entity dataset scraped by metadatarr
scraper librivox_audiobooks.
Rows: 22,151
Fields
librivox_id
title
description
url_text_source
language
copyright_year
num_sections
url_rss
url_zip_file
url_project
url_librivox
project_type
totaltimesecs
genres
authors
readers
Source
Generated by scrapers/librivox_audiobooks.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-artists
Unified Music Artists
Cross-database artist dataset unifying MusicBrainz, TheAudioDB, Metal Archives, ProgArchives, Jazz, Classical Composers, Bandcamp, SoundCloud, and YouTube Music into one row per artist with flat canonical ID columns.
1,662,320 total rows — full outer union across all sources.
Canonical ID columns
All nullable — present only when the artist was found in that database:
Column
Source
Type
mb_id
MusicBrainz
UUID string
adb_id… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-artists.media-metadata-tvmaze-shows
TigreGotico/media-metadata-tvmaze-shows
Rich entity dataset scraped by metadatarr
scraper tvmaze_shows.
Rows: 88,297
Fields
tvmaze_id
name
type
language
genres
status
runtime
average_runtime
premiered
ended
network_name
network_country
rating_average
schedule_time
schedule_days
summary
official_site
imdb_id
thetvdb_id
tvrage_id
image_medium
Source
Generated by scrapers/tvmaze_shows.py. See the metadatarr repo for the full
pipeline and scraper… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-tvmaze-shows.media-metadata-metal-archives
TigreGotico/media-metadata-metal-archives
Rich entity dataset scraped by metadatarr
scraper metal_archives.
Rows: 245,976
Fields
band_id
name
url
country
location
status
formed_in
years_active
genres
themes
current_label_id
current_label_name
comment
Source
Generated by scrapers/metal_archives.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-jikan-manga
TigreGotico/media-metadata-jikan-manga
Rich entity dataset scraped by metadatarr
scraper jikan_manga.
Rows: 83,790
Fields
mal_id
title
title_english
title_japanese
aliases
type
status
chapters
volumes
published_from
published_to
authors
serializations
genres
themes
demographics
score
scored_by
rank
popularity
members
synopsis
background
approved
Source
Generated by scrapers/jikan_manga.py. See the metadatarr repo for the full
pipeline and scraper… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-jikan-manga.media-metadata-podcastindex-podcasts
TigreGotico/media-metadata-podcastindex-podcasts
Rich entity dataset scraped by metadatarr
scraper podcastindex_podcasts.
Rows: 112,717
Fields
itunes_id
title
author
image
genres
url
description
language
episode_count
explicit
feed_url
country_charts
source
entity_type
Source
Generated by scrapers/podcastindex_podcasts.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-deezer-playlists
TigreGotico/media-metadata-deezer-playlists
Rich entity dataset scraped by metadatarr
scraper deezer_playlists.
Rows: 16,724
Fields
deezer_id
title
description
nb_tracks
duration_seconds
fans
creation_date
genres
creator_name
creator_id
is_editorial
source_genre
tracks
url
Source
Generated by scrapers/deezer_playlists.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-progarchives-artists
TigreGotico/media-metadata-progarchives-artists
Rich entity dataset scraped by metadatarr
scraper progarchives_artists.
Rows: 9,879
Fields
artist_id
name
genre
country
bio
url
n_albums
Source
Generated by scrapers/progarchives_artists.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-listennotes-podcasts
TigreGotico/media-metadata-listennotes-podcasts
Rich entity dataset scraped by metadatarr
scraper listennotes_podcasts.
Rows: 492
Fields
ln_id
ln_url
title
author
description
image
language
genres
episode_count
listen_score
global_rank
website
entity_type
Source
Generated by scrapers/listennotes_podcasts.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-jazz-artists
TigreGotico/media-metadata-jazz-artists
Rich entity dataset scraped by metadatarr
scraper jazz_artists.
Rows: 7,941
Fields
artist_slug
name
genres
genre
country
bio
url
n_albums
Source
Generated by scrapers/jazz_artists.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-radiobrowser-stations
TigreGotico/media-metadata-radiobrowser-stations
Rich entity dataset scraped by metadatarr
scraper radiobrowser_stations.
Rows: 58,923
Fields
stationuuid
name
url
url_resolved
homepage
favicon
country
countrycode
state
language
language_codes
tags
codec
bitrate
hls
votes
clickcount
clicktrend
last_check_ok
entity_type
Source
Generated by scrapers/radiobrowser_stations.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-anilist-anime
TigreGotico/media-metadata-anilist-anime
Rich entity dataset scraped by metadatarr
scraper anilist_anime.
Rows: 10,000
Fields
anilist_id
mal_id
title_romaji
title_english
title_native
type
format
status
episodes
duration
chapters
volumes
country_of_origin
source_material
start_date
end_date
season
season_year
genres
tags
studios
studio_ids
average_score
popularity
favourites
is_adult
Source
Generated by scrapers/anilist_anime.py. See the… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-anilist-anime.media-metadata-fanedits
TigreGotico/media-metadata-fanedits
Rich entity dataset scraped by metadatarr
scraper fanedits.
Rows: 2,881
Fields
fanedit_id
slug
title
url
cover_url
faneditor
original_title
genre
franchise
fanedit_type
original_release_date
original_running_time
imdb_id
fanedit_release_date
fanedit_running_time
time_cut
time_added
subtitles
available_in
release_information
synopsis
additional_notes
special_thanks
cuts_and_additions
intention
awards
editor_rating
user_rating… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-fanedits.media-metadata-wikidata-entities
TigreGotico/media-metadata-wikidata-entities
Rich entity dataset scraped by metadatarr
scraper wikidata_entities.
Rows: 325,563
Fields
wikidata_id
label_en
description_en
country
country_qid
inception_year
dissolved_year
website
entity_type
Source
Generated by scrapers/wikidata_entities.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-unified-artists
TigreGotico/media-metadata-unified-artists
Rich entity dataset scraped by metadatarr
scraper unified_artists.
Rows: 2,208,545
Fields
mb_id
adb_id
ma_id
progarchives_id
jazz_id
classical_id
wikidata_id
bandcamp_id
bandcamp_url
soundcloud_id
soundcloud_username
youtube_channel_id
sources
name
sort_name
type
gender
disambiguation
ended
aliases
country
area
begin_date
end_date
tags
style
mood
classical_period
biography_en
members
ipi_codes
isni_codes
website
social… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-unified-artists.media-metadata-steam-games
TigreGotico/media-metadata-steam-games
Rich entity dataset scraped by metadatarr
scraper steam_games.
Rows: 82,236
Fields
steam_appid
name
developer
publisher
score_rank
positive_reviews
negative_reviews
owners
average_playtime_forever
average_playtime_2weeks
median_playtime_forever
price_usd
discount_pct
ccu
type
genres
categories
release_date
is_free
platforms_windows
platforms_mac
platforms_linux
metacritic_score
short_description
Source… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-steam-games.media-metadata-ytmusic-playlists
TigreGotico/media-metadata-ytmusic-playlists
Rich entity dataset scraped by metadatarr
scraper ytmusic_playlists.
Rows: 3,679
Fields
ytm_id
title
description
track_count
source_category
source_section
tracks
Source
Generated by scrapers/ytmusic_playlists.py. See the metadatarr repo for the full
pipeline and scraper source code.
