CoolFace
Datasetpublic

KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate

Disclaimer: The original dataset can be found here. It is published by Digital Divide Data Cambodia (DDD-Cambodia). License: Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0). Please attribute Digital Divide Data if you use this dataset in any way. Objective of this dataset Add English translation: a new column en_translate is added to the original dataset (only from parquet 000 to 159 of the original… See the full description on the dataset page: https://huggingface.co/datasets/KrorngAI/DDD-Cambodia-khmer-speech-dataset-parquet-000-159-en-translate.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes436downloads
Dataset Card

_Disclaimer:_ The original dataset can be found here. It is published by Digital Divide Data Cambodia (DDD-Cambodia).

License:

Khmer ASR Cultural Dataset's license is Creative Commons Attribution Share Alike 4.0 International (CC-BY-SA-4.0). Please attribute Digital Divide Data if you use this dataset in any way.

Objective of this dataset

_Add English translation: a new column `entranslate is added to the original dataset (only from parquet 000 to 159 of the original dataset). It is the English translation of items in column transcript`.

You can use this dataset to train your AI model to translate Khmer audio into English text.

Acknowledge

The transcripts are all translated to English by `TranslateKH`. Deeply thanks to TranslateKH team for granting me a temporary access to their API, so that I can create this dataset. Their kindness helps pushing the research in AI for Khmer language.

Dataset Card Author

  • —ឈ្មោះ: បណ្ឌិត ឃុន គីមអាង
  • —Name: KHUN Kimang (Ph.D.)