twangodev/librivox-mirror
LibriVox Mirror Fast, structured, continuously updated LibriVox audio mirror. Current snapshot Metric Value Published books 21,724 Published sections 493,186 Audio hours 132,549.7 Audio languages 86 Quarantined books 610 Last updated (UTC) 2026-09-21T15:06:16.273192Z Audio by language Language Hours English 131,600.3 German 417.0 Spanish 160.9 French 103.8 Portuguese 37.4 Polish 34.1 Dutch 25.8… See the full description on the dataset page: https://huggingface.co/datasets/twangodev/librivox-mirror.
LibriVox Mirror
 
Fast, structured, continuously updated LibriVox audio mirror.
Current snapshot
Audio by language
Dataset structure
The default preview config contains a bounded set of browser-playable original MP3s. Full sections and books indexes remain typed Parquet; canonical audio is streamed from WebDataset shards under data/.
Use language and hash_partition to construct stable downstream subsets or evaluation splits. Source metadata features preserve unmodeled LibriVox and Internet Archive fields for provenance.
License and attribution
The copyrightable mirror-specific compilation, curation, normalized metadata, indexes, and documentation are licensed under CC BY 4.0. Please cite this dataset when using those parts. Original LibriVox audio remains public domain in the United States and is not relicensed by this mirror. Rules may differ by jurisdiction.
Citation
@misc{ding2026librivoxmirror,
author = {James Ding},
title = {LibriVox Mirror},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/datasets/twangodev/librivox-mirror}},
note = {Continuously updated dataset}
}Provenance and integrity
Each sample links to its LibriVox project and exact Internet Archive source file. Upstream checksums and a mirror SHA-256 are preserved, and audio is not transcoded.
