CoolFace
Datasetpublicgated

yyhenggg/EN-MALAY-CS-FILTERED

EN-MALAY-CS-FILTERED English–Malay code-switching conversational speech from the IMDA National Speech Corpus (2021), segmented to utterance level and cleaned with a 3-model agreement filter: each utterance was transcribed by three ASR models — openai/whisper-large-v3, MERaLiON/MERaLiON-2-10B-ASR, and Qwen/Qwen3-ASR-1.7B — and an utterance is removed when all three models score WER > 60% against the reference transcript (all models agreeing the reference is unreliable).… See the full description on the dataset page: https://huggingface.co/datasets/yyhenggg/EN-MALAY-CS-FILTERED.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes7downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.

yyhenggg/EN-MALAY-CS-FILTERED · CoolFace