datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Malaysian-STT-Whisper
Malaysian STT Whisper format
Heavy postprocessing and post-translation to improve pseudolabeled Whisper Large V3. Also include word level timestamp.
Postprocessing
Check repetitive trigrams.
Verify Voice Activity using Silero-VAD.
Verify scores using Force Alignment.
Post-translation
We use mesolitica/nanot5-base-malaysian-translation-v2.1.
Dataset involved
Malaysian context v2
Singaporean context
Indonesian context
Mandarin audio
Tamil audio… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-STT-Whisper.fleurs-r-neucodec-all-languages
FLEURS-R NeuCodec All Languages
FLEURS-R metadata, source audio and precomputed NeuCodec speech tokens for 102
locales, plus a speaker label FLEURS itself does not ship.
Layout
data/{locale}-{split}.parquet — metadata, one row per utterance (this is what the
viewer shows).
audio/{locale}-{split}.zip — source FLEURS-R audio, 24kHz mono PCM16 WAV, members
named audio/{locale}/{split}/{id}.wav (the path column).
neucodec/{locale}-{split}-rank{N}.zip — NeuCodec… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/fleurs-r-neucodec-all-languages.Malaysia-Whisper-Berita
Malaysia Whisper Berita Dataset
Dataset Description
This is a custom audio dataset designed specifically for fine-tuning Automatic Speech Recognition (ASR) models, such as OpenAI's Whisper, on the Malay language (Bahasa Melayu).
The audio consists of high-quality Malaysian news broadcasts (Berita), which provide excellent examples of formal, standard Malay (Bahasa Baku) and domain-specific vocabulary (e.g., politics, economy, current events).
Language: Malay (ms-MY)… See the full description on the dataset page: https://huggingface.co/datasets/PishangShedappp/Malaysia-Whisper-Berita.
