datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
onevoice-asr
OneVoice ASR
Generated Vietnamese TTS WAV dataset for ASR training.
Layout follows Hugging Face AudioFolder:
train/metadata.jsonl + train/*.wav
validation/metadata.jsonl + validation/*.wav
test/metadata.jsonl + test/*.wav
onevoice
OneVoice
Vietnamese-English garment-factory lexicon, MT, and generated ASR audio.
Data and QC provenance
Vietnamese audio contains the restored VieNeu/OmniVoice assets. Rows without a recorded ASR QC result are marked qc_status=unverified rather than presented as QC-passed.
English audio is generated from text_en with one Qwen3-TTS VoiceDesign identity. INT8 describes inference weight quantization; exported WAV audio remains ordinary PCM.
Only English rows that… See the full description on the dataset page: https://huggingface.co/datasets/hatakekksheeshh/onevoice.common-voice-corpus-20hatrang-voice-4h
VieNeu-TTS Vietnamese Speech Dataset
Dataset Description
Vietnamese text-to-speech dataset for fine-tuning VieNeu-TTS models. Contains paired audio recordings and Vietnamese text transcriptions.
Dataset Structure
Data Fields
audio: Audio recording (WAV, 24kHz, mono, 16-bit PCM)
transcription: Vietnamese text transcription of the audio
Data Splits
Split
Samples
train
1,805
Dataset Statistics
Total audio files:… See the full description on the dataset page: https://huggingface.co/datasets/quocs/hatrang-voice-4h.hat_asr_sixian_broadcast_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
231.77
88,262
3,633,737
9.45
4.36
Total
-
231.77
88,262
3,633,737
9.45
4.36
hatespeech_synthesized_datasethat_asr_sixian_reading_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
203.03
106,493
2,137,244
6.86
2.92
Total
-
203.03
106,493
2,137,244
6.86
2.92
hat_asr_sixian_reading_cm
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_sx
Hakka_Sixian
204.60
105,513
2,152,732
6.98
2.92
0
0
Total
-
204.60
105,513
2,152,732
6.98
2.92
0
0
hat_asr_hailu_reading_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Hailu
客語_海陸
197.58
104,555
2,018,861
6.80
2.84
Total
-
197.58
104,555
2,018,861
6.80
2.84
hat_asr_sixian_reading_clean_r
hat_asr_sixian_reading_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_sixian_reading_clean.
Summary
Subset: Hakka_Sixian
Dialect: 客語四縣
Train samples: 106493
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
203.03
106,493
5,597,645
6.86
7.66
Total
-
203.03
106,493
5,597,645
6.86
7.66
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_sixian_reading_clean_r.HatawASR-p10hat_asr_sixian_broadcast_clean_r
hat_asr_sixian_broadcast_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_sixian_broadcast_clean.
Summary
Subset: Hakka_Sixian
Dialect: 客語四縣
Train samples: 88263
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
231.77
88,263
8,172,269
9.45
9.79
Total
-
231.77
88,263
8,172,269
9.45
9.79
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_sixian_broadcast_clean_r.hat_asr_hailu_reading_clean_r
hat_asr_hailu_reading_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_hailu_reading_clean.
Summary
Subset: Hakka_Hailu
Dialect: 客語海陸
Train samples: 104555
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Hailu
客語_海陸
197.58
104,555
5,287,046
6.80
7.43
Total
-
197.58
104,555
5,287,046
6.80
7.43
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_hailu_reading_clean_r.hat_asr_nansixian_reading_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_NanSixian
客語_南四縣
71.49
36,070
770,713
7.13
2.99
Total
-
71.49
36,070
770,713
7.13
2.99
hat_asr_nansixian_reading_cm
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_nsx
Hakka_NanSixian
72.40
34,992
775,610
7.45
2.98
0
0
Total
-
72.40
34,992
775,610
7.45
2.98
0
0
hat_asr_hailu_reading_cm
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_hl
Hakka_Hailu
199.61
103,519
2,035,542
6.94
2.83
0
0
Total
-
199.61
103,519
2,035,542
6.94
2.83
0
0
ha-tts-csv
Hausa TTS Dataset
Dataset Description
This dataset contains Hausa language text-to-speech (TTS) recordings from multiple speakers. It includes audio files paired with their corresponding Hausa text transcriptions.
Dataset Structure
The dataset is organized as follows:
├── data/
│ ├── metadata.csv # Metadata (source, audio paths, text)
│ └── audio_files/
│ ├── 97f373e8-f6e6-.../ # Speaker 1 audio files
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/Aybee5/ha-tts-csv.hat_asr_nansixian_reading_clean_r
hat_asr_nansixian_reading_clean_r
This dataset is an enhanced -R variant of formospeech/hat_asr_nansixian_reading_clean.
Summary
Subset: Hakka_NanSixian
Dialect: 客語南四縣
Train samples: 36070
Audio: enhanced 24 kHz WAV
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_NanSixian
客語_南四縣
71.49
36,070
2,046,782
7.14
7.95
Total
-
71.49
36,070
2,046,782
7.14
7.95
Processing
Start from the original… See the full description on the dataset page: https://huggingface.co/datasets/formospeech/hat_asr_nansixian_reading_clean_r.HateSpeechDetection_Detoxy_VCTK_LJSpeech_CV_MELDha-tts-mixedDataset conversion helper
Install dependencies:
pip install -r requirements.txt
Run the script from the repo root to create the parquet and copy audio files into data/audio_files:
python3 scripts/create_parquet.py
Resulting files:
data/dataset.parquet (contains columns: source, text, audio)
data/audio_files//.wav (copied audio files)
Notes:
The audio column contains relative paths starting with data/audio_files/... so the whole data/ folder can be uploaded to Hugging Face or copied… See the full description on the dataset page: https://huggingface.co/datasets/Aybee5/ha-tts-mixed.hat_tts_hailu_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Hailu
客語_海陸
53.62
24,573
641,201
7.85
3.32
Total
-
53.62
24,573
641,201
7.85
3.32
hat_tts_sixian_clean
TRAIN
Subset
lang_group
hours
n_utts
n_chars
secs/utt
chars/sec
Hakka_Sixian
客語_四縣
56.47
26,519
681,363
7.67
3.35
Total
-
56.47
26,519
681,363
7.67
3.35
hat_tts_sixian
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_sx
Hakka_Sixian
56.61
26,316
695,847
7.74
3.41
0
0
Total
-
56.61
26,316
695,847
7.74
3.41
0
0
hat_tts_hailu
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_hl
Hakka_Hailu
54.08
24,134
655,843
8.07
3.37
0
0
Total
-
54.08
24,134
655,843
8.07
3.37
0
0
cmu_hat_ltihat_asr_sixian_broadcast
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_sx
Hakka_Sixian
240.64
80,073
3,757,980
10.82
4.34
0
0
Total
-
240.64
80,073
3,757,980
10.82
4.34
0
0
hat_asr_nansixian_reading_mc
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_nsx
Hakka_NanSixian
72.04
34,992
775,610
7.41
2.99
0
0
Total
-
72.04
34,992
775,610
7.41
2.99
0
0
Hate-SpeechHatawASR-p1hat_asr_hailu_reading_mc
TRAIN
Subset
lang_name
hours
n_utts
n_chars_in_utts
secs/utt
chars/sec
n_sents
n_chars_in_sents
hak_hl
Hakka_Hailu
198.88
103,519
2,035,542
6.92
2.84
0
0
Total
-
198.88
103,519
2,035,542
6.92
2.84
0
0
