media-metadata
media-metadata-gutenberg-books
TigreGotico/media-metadata-gutenberg-books
Rich entity dataset scraped by metadatarr
scraper gutenberg_books.
Rows: 78,738
Fields
gutenberg_id
title
authors
translators
subjects
bookshelves
languages
copyright
media_type
download_count
has_text
has_epub
entity_type
Source
Generated by scrapers/gutenberg_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-openlibrary-books
TigreGotico/media-metadata-openlibrary-books
Rich entity dataset scraped by metadatarr
scraper openlibrary_books.
Rows: 4,098,190
Fields
olid
title
subtitle
authors
author_key
first_publish_year
subjects
isbn_10
isbn_13
publisher
language
number_of_pages_median
ebook_access
has_fulltext
edition_count
cover_i
Source
Generated by scrapers/openlibrary_books.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-classical-composers
TigreGotico/media-metadata-classical-composers
Rich entity dataset scraped by metadatarr
scraper classical_composers.
Rows: 15,587
Fields
composer_id
name
country
life
birth
death
period
image_url
url
bio
radio_id
notable
must_know
n_recordings
n_performers
n_albums
n_works_listed
n_albums_listed
Source
Generated by scrapers/classical_composers.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-musicbrainz-artists
TigreGotico/media-metadata-musicbrainz-artists
Rich entity dataset scraped by metadatarr
scraper musicbrainz_artists.
Rows: 1,520,644
Fields
mb_id
name
sort_name
type
gender
country
area
begin_date
end_date
ended
disambiguation
aliases
tags
ipi_codes
isni_codes
Source
Generated by scrapers/musicbrainz_artists.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-librivox-audiobooks
TigreGotico/media-metadata-librivox-audiobooks
Rich entity dataset scraped by metadatarr
scraper librivox_audiobooks.
Rows: 22,151
Fields
librivox_id
title
description
url_text_source
language
copyright_year
num_sections
url_rss
url_zip_file
url_project
url_librivox
project_type
totaltimesecs
genres
authors
readers
Source
Generated by scrapers/librivox_audiobooks.py. See the metadatarr repo for the full
pipeline and scraper source code.
media-metadata-artists
Unified Music Artists
Cross-database artist dataset unifying MusicBrainz, TheAudioDB, Metal Archives, ProgArchives, Jazz, Classical Composers, Bandcamp, SoundCloud, and YouTube Music into one row per artist with flat canonical ID columns.
1,662,320 total rows — full outer union across all sources.
Canonical ID columns
All nullable — present only when the artist was found in that database:
Column
Source
Type
mb_id
MusicBrainz
UUID string
adb_id… See the full description on the dataset page: https://huggingface.co/datasets/LeData/media-metadata-artists.
