datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
f5_tts_ru_accent
Original datasets:
https://huggingface.co/datasets/mozilla-foundation/common_voice_17_0
https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd
https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield
https://huggingface.co/datasets/bond005/sova_rudevices
https://huggingface.co/datasets/Aniemore/resd_annotated
SageLM-F5-TTSne-tts-f5-grt
NE-TTS F5 Garo (grt)
F5-TTS training dataset for Garo (grt). Contains 24,772 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
24,772
Hours
29.6h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-grt
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-grt.f5tts-vietnamese-datasetne-tts-f5-ccp
NE-TTS F5 Chakma (ccp)
F5-TTS training dataset for Chakma (ccp). Contains 10,689 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
10,689
Hours
14.3h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-ccp
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-ccp.F5-TTS-Hindif5_tts_ru_accent
Original datasets:
https://huggingface.co/datasets/mozilla-foundation/common_voice_17_0
https://huggingface.co/datasets/bond005/sberdevices_golos_10h_crowd
https://huggingface.co/datasets/bond005/sberdevices_golos_100h_farfield
https://huggingface.co/datasets/bond005/sova_rudevices
https://huggingface.co/datasets/Aniemore/resd_annotated
f5tts-kabyle-dataset
F5-TTS Kabyle Dataset
Clean, deduplicated audio-text dataset for Kabyle (Taqbaylit / Tamaziɣt) TTS fine-tuning with F5-TTS.
Statistics
Metric
Value
Total clips
59,462
Total duration
41.30 hours
Sample rate
24 kHz mono
Avg clip length
2.50s
Min clip length
1.00s
Max clip length
12.65s
Unique phrases
59,462 (0% duplicates)
Unique characters
112
Sources
Tatoeba (67.8%) + Common Voice 26 tiny (32.2%)
Source Datasets… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/f5tts-kabyle-dataset.ne-tts-f5-nag
NE-TTS F5 Nagamese (nag)
F5-TTS training dataset for Nagamese (nag). Contains 9,688 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
9,688
Hours
14.5h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-nag
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-nag.4tts-f5ttshindi_F5-TTSf5tts-voice-bygptF5-TTS_benchne-tts-f5-lus
NE-TTS F5 Mizo (lus)
F5-TTS training dataset for Mizo (lus). Contains 8,554 clips at 24kHz (SNR >= 20dB) from the cleaned NE-TTS dataset, formatted for F5-TTS training.
Stats
Metric
Value
Clips
8,554
Hours
14.5h
Sample rate
24kHz
SNR filter
>= 20dB
Source
ne-tts-lus
Schema
Column
Type
Description
audio
Audio
24kHz WAV audio
text
string
Cleaned transcript
Note
Audio is upsampled from 16kHz… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/ne-tts-f5-lus.F5-TTS-Small_audio_testing_datasettongyi-f5ttsf5-tts-datasetf5_ttsF5-TTSF5-TTS-LOUCUTORESF5-TTSf5-ttsF5TTS-0868683336f5tts-persian-dataset
