CoolFace
Datasetpublic

twangodev/librivox-mirror

LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,724 Published sections 493,186 Audio hours 132,549.7 Audio languages 86 Quarantined books 610 Last updated (UTC) 2026-09-21T15:06:16.273192Z Audio by language Language Hours English 131,600.3 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.

sourceHugging Facecc-by-4.0updated 8h agoView on Hugging Face
0likes23kdownloads
Dataset Card

LibriVox Mirror

![License: CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) ![Hugging Face dataset](https://huggingface.co/datasets/twangodev/librivox-mirror) Last updated

Fast, structured, continuously updated LibriVox audio mirror.

Current snapshot

MetricValue
Published books21,724
Published sections493,186
Audio hours132,549.7
Audio languages86
Quarantined books610
Last updated (UTC)2026-09-21T15:06:16.273192Z

Audio by language

LanguageHours
English131,600.3
German417.0
Spanish160.9
French103.8
Portuguese37.4
Polish34.1
Dutch25.8
Italian22.5
Chinese14.7
Japanese13.9
Ukrainian12.1
Danish11.6
Russian11.1
Hindi9.2
Arabic8.9
Latin7.3
Catalan5.1
Esperanto5.0
Romanian4.9
Hebrew4.5
Greek3.3
Swedish2.1
Finnish2.0
Multilingual1.8
Ancient Greek1.8
Javanese1.7
Hungarian1.6
Tagalog1.6
Old English1.5
Yiddish1.4
Persian/Farsi1.3
Urdu1.3
Galician1.2
Norwegian1.1
Luxembourgish1.0
Bulgarian1.0
Afrikaans0.8
Indonesian0.8
Czech0.7
Marathi0.7
Slovenian0.7
Sanskrit0.6
Tamil0.6
Kapampangan0.5
Latvian0.5
Middle English0.5
Braj0.5
Balinese0.4
Malay0.4
Minangkabau0.4
Cantonese Chinese0.4
Buginese0.4
Acehnese0.4
Sundanese0.3
Serbian0.3
Bengali0.3
Oriya0.3
Nynorsk0.3
Volapük0.3
Occitan0.2
Walloon0.2
Korean0.2
Western Frisian0.2
Croatian0.2
Slovak0.2
Maori0.2
Turkish0.2
Faroese0.2
Bisaya/Cebuano0.2
Old Javanese0.1
Rajasthani0.1
Nepali0.1
Irish0.1
Garo0.1
Old Norse0.1
Scottish Gaelic0.1
Lithuanian0.1
Welsh0.0
Macedonian0.0
Maltese0.0
Old Sundanese0.0
Assamese0.0
Low German0.0
Belarusian0.0
Palatine German0.0
Old Tupi0.0

Dataset structure

The default preview config contains a bounded set of browser-playable original MP3s. Full sections and books indexes remain typed Parquet; canonical audio is streamed from WebDataset shards under data/.

Use language and hash_partition to construct stable downstream subsets or evaluation splits. Source metadata features preserve unmodeled LibriVox and Internet Archive fields for provenance.

License and attribution

The copyrightable mirror-specific compilation, curation, normalized metadata, indexes, and documentation are licensed under CC BY 4.0. Please cite this dataset when using those parts. Original LibriVox audio remains public domain in the United States and is not relicensed by this mirror. Rules may differ by jurisdiction.

Citation

bibtex
@misc{ding2026librivoxmirror,
  author       = {James Ding},
  title        = {LibriVox Mirror},
  year         = {2026},
  publisher    = {Hugging Face},
  howpublished = {\url{https://huggingface.co/datasets/twangodev/librivox-mirror}},
  note         = {Continuously updated dataset}
}

Provenance and integrity

Each sample links to its LibriVox project and exact Internet Archive source file. Upstream checksums and a mirror SHA-256 are preserved, and audio is not transcoded.