datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emilia-ja-plus-metadata
Emilia Dataset JA Plus — normalized metadata
This is an audit-backed, metadata-only derivative of
ayousanz/Emilia-Dataset-JA-Plus. It does not bundle audio payloads.
Verified snapshot statistics
Metric
Value
Metadata rows
78,748
Unique IDs
78,748
Duplicate IDs
0
Unique speakers
10,046
Duration represented by metadata
145.06 hours
Language labels
ja: 78,748
Unique transcripts
64,034
Duplicate transcript rows
14,714
Japanese-labelled rows… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/emilia-ja-plus-metadata.voicevox-voice-corpus-metadata
VOICEVOX synthetic Japanese speech corpus
This metadata derivative documents ayousanz/voicevox-voice-corpus at
39dff6b254bf3118accad04446238ae3147156a6.
Audio payloads were not downloaded during this metadata audit.
Verified repository inventory
Corpus
Voice/style directories
WAV files
WAV bytes
ROHAN-corpus
87
400,200
90.16 GB
ita-corpus
87
36,888
6.64 GB
tsukuyomi-chan-corpus
87
8,700
3.07 GB
Total WAV paths: 445,788
Total WAV bytes: 99.87… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/voicevox-voice-corpus-metadata.
