CoolFace
Datasetpublicgated

yyhenggg/EN-MALAY-CS-FILTERED

EN-MALAY-CS-FILTERED English–Malay code-switching conversational speech from the IMDA National Speech Corpus (2021), segmented to utterance level and cleaned with a 3-model agreement filter: each utterance was transcribed by three ASR models — openai/whisper-large-v3, MERaLiON/MERaLiON-2-10B-ASR, and Qwen/Qwen3-ASR-1.7B — and an utterance is removed when all three models score WER > 60% against the reference transcript (all models agreeing the reference is unreliable).… See the full description on the dataset page: https://huggingface.co/datasets/yyhenggg/EN-MALAY-CS-FILTERED.

sourceHugging Faceotherupdated 2mo agoView on Hugging Face
0likes8downloads

No commit history came back for main. The revision may not exist, or the source declined the request.

yyhenggg/EN-MALAY-CS-FILTERED · CoolFace