CoolFace
Datasetpublic

malaysia-ai/Malaysian-STT

Malaysian-STT Prepare streaming and whole mode Speech-to-Text Malaysian context dataset, suitable to train streaming LLM base or Encoder-Decoder such as Whisper. Merged 30 seconds chunk into one audio file, can up to 10 minutes. Segmentize based on silent at least 0.3 seconds. Reject low score based on force alignment. Reject timestamp anomaly based on force alignment. Dataset involved Dialects IMDA Malaysian context Malaysia Parliament Science context… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/Malaysian-STT.

sourceHugging Faceupdated 1y agoView on Hugging Face
2likes1.1kdownloads

No commit history came back for main. The revision may not exist, or the source declined the request.