CoolFace
Datasetpublic

Digital-Divide-Data/khmer-speech-dataset

Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
26likes3.8kdownloads
settings

This repository belongs to Digital-Divide-Data on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namekhmer-speech-dataset
visibilitypublic
licencecc-by-sa-4.0
gatedno
ownerDigital-Divide-Data
Account settings
Digital-Divide-Data/khmer-speech-dataset · CoolFace