CoolFace
Datasetpublic

malaysia-ai/Low-Language-TTS

Low-Language-TTS A held-out TTS/ASR test set for the long tail of malaysia-ai/Multilingual-TTS: 50 of the lowest-resource languages in that corpus, 25 utterances each (1250 rows, 2.56 hours). Every row carries the three things needed to score a NeuCodec speech-token model without touching the parent corpus: column what audio the original clip, exactly as stored upstream (mostly mp3) tokens NeuCodec speech tokens at 50 tokens/s — the <|s_N|> ids the Multilingual-TTS… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Low-Language-TTS.

sourceHugging Faceupdated 7d agoView on Hugging Face
0likes57downloads
Dataset Card

Low-Language-TTS

A held-out TTS/ASR test set for the long tail of malaysia-ai/Multilingual-TTS: 50 of the lowest-resource languages in that corpus, 25 utterances each (1250 rows, 2.56 hours).

Every row carries the three things needed to score a NeuCodec speech-token model without touching the parent corpus:

columnwhat
audiothe original clip, exactly as stored upstream (mostly mp3)
tokensNeuCodec speech tokens at 50 tokens/s — the `<\s_N\>` ids the Multilingual-TTS models emit
post-normalizednormalized transcription (postnormalizer.py output, the reference text for CER)
textraw upstream transcript
languageGlotLID v3 label, e.g. mos_Latn
speaker, subset, audio_filenameprovenance back into the parent corpus
duration, sampling_ratedecoded clip length in seconds, and its sample rate
lid_prob, lid_marginlanguage-ID confidence of the row (see below)

How the languages were picked

Ranking the corpus' GlotLID histogram and taking the rarest labels does not work, and this was measured: of the 500 rarest labels, exactly 3 survived row-level verification. Short transcripts are misclassified constantly (a three-word Hausa line lands on wnc_Latn) and und_<script> labels are punctuation false positives, so that tail is detector noise, not languages.

A language is instead taken seriously when some subset of the corpus was collected for it — at least 50 rows and at least 40% of that subset. 177 of the 2,100 labels clear that bar, and they rank into a genuine low-resource tail. The 50 with the fewest rows in the parent corpus that can still fill 25 verified utterances are what you see here.

Individual rows are then verified too, because a subset's minority rows are not its language. A row is kept only when

  • the normalized transcript is >= 40 characters and >= 5 words (characters instead of words for scripts written without spaces), long enough for GlotLID to be reliable,
  • re-running GlotLID v3 reproduces the stored label with probability
= 0.9 and top1-top2 margin >= 0.5,
  • NeuCodec tokens exist for the clip and decode to 50-1500 tokens (1-30s),
  • the transcript has no more words than speech tokens,
  • the decoded audio length and len(tokens) / 50 agree to within 20% — audio and tokens live in different zips upstream, and this is what catches a mispairing.

Rows are deduplicated by text and spread over speakers (4 per speaker, relaxed only when a language would otherwise not reach 25 — much of the tail is single-speaker corpora). On long, confident rows the stored and re-predicted labels agree 97-99% of the time, so the language tag here is considerably more trustworthy than the parent corpus' raw language column.

Languages

corpus rows is how many rows that language has in the whole 121.8M-row parent corpus.

languagerowscorpus rowsminutessubset(s)
guc_Latn251,1282.2wayuuCOtest
njm_Latn251,2112.3ne-asr-njm, ne-asr-njm-aug
nre_Latn251,8152.3ne-asr-nre-aug, ne-tts-nre
nri_Latn252,1822.1ne-asr-nri-aug
njo_Latn253,4222.7ne-asr-njo-aug
qup_Latn254,1152.6killkan
gos_Latn254,6001.6gos-demo, gronings
mjw_Latn254,7932.1ne-asr-mjw, ne-tts-mjw
bci_Latn254,8473.9baoule-common-voice
kdj_Latn255,9591.8karamojong-speech-dataset, turkana-speech-dataset
teo_Latn257,0792.1turkana-speech-dataset
iba_Latn257,3484.7iban-speech
mni_Mtei258,8463.1meiteimayek-audio-parallel-corpus
zgh_Tfng259,4451.3TOSD
crh_Latn2513,7992.2qirimtatar-tts, tts-crh-sevil-fixed
suk_Latn2514,8234.1Sukuma-Voices, Sukuma-Voices-ACL
fon_Latn2515,2402.1fongbe-speech-zenodo
smj_Latn2515,3392.7salmon-asr-smj
bre_Latn2515,6262.3mlsuperbbr
hat_Latn2516,0293.5cmuhaitiancreole_speech
oci_Latn2516,5606.4occitan-speech-dataset
lad_Latn2517,4452.6ladino-karen-TTS
tgk_Cyrl2518,1765.1tajik-asr-augmented-test
kha_Latn2518,1833.9khasi-tts
hac_Arab2518,2681.9hawrami_speech
trp_Latn2519,3661.9ne-asr-trp-aug
gom_Latn2520,0244.1konkani-bible-audio
kbd_Cyrl2520,2142.4sixuxaryijirimak7
ekk_Latn2522,4994.9estonian-speech-dataset
crh_Cyrl2522,8292.1audiobooks
sdh_Arab2524,3032.8southern-kurdish-asr
mya_Mymr2524,4913.5burmese-synthetic-speech-corpus, myanmar-speech-dataset-for-asr
lat_Latn2525,5583.3Latin-Audio
slv_Latn2529,1123.9slovenian-speech-dataset
fin_Latn2530,7403.9TTS-Finnish
lin_Latn2532,4293.0LRSC
sna_Latn2535,4866.0mixedshonadataset, sna-dataset-annotated, sna-manasseh-150-raw
lao_Laoo2539,3903.8lao-speech-dataset
vag_Latn2540,3332.6vagla-speech-text-parallel
lvs_Latn2546,6274.2latvian-speech-dataset
bod_Tibt2550,3922.2tibetanwztts
san_Deva2550,5432.8shrutilipi_sanskrit
fry_Latn2551,9622.4frisian-asr-cv22
chv_Cyrl2552,0552.3chuvash_voice
fuv_Latn2555,7252.8ASR_pulaar
plt_Latn2558,9173.8malagasy-nwt-bible
tir_Ethi2560,8935.3phonetico-speech
guj_Gujr2561,0823.0gujarati-f-openslr
gsw_Latn2561,1052.2swiss-german-city-sentences_v2
xsm_Latn2561,2772.4kasem-speech-text-parallel

Usage

python
from datasets import load_dataset

ds = load_dataset('malaysia-ai/Low-Language-TTS', split='test')
row = ds[0]
row['audio']['array'], row['tokens'], row['post-normalized'], row['language']

Built by low-language-testset/build_testset.py in the Multilingual-Speech-Model repo, from malaysia-ai/Multilingual-TTS (audio + NeuCodec tokens) and malaysia-ai/Multilingual-TTS-language (language + normalized transcription).