CoolFace
Datasetpublic

SeyhaLite/Translate-English-Khmer-All

Translate Khmer All Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems. Project Vision I hope this dataset helps your project succeed! This is a comprehensive collection that aggregates data from multiple domains—including business, technology, medical, legal, and daily conversation. It is designed to provide a robust foundation… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Translate-English-Khmer-All.

sourceHugging Faceapache-2.0updated 8mo agoView on Hugging Face
1likes38downloads
Dataset Card

Translate Khmer All

Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems.

Project Vision

I hope this dataset helps your project succeed! This is a comprehensive collection that aggregates data from multiple domains—including business, technology, medical, legal, and daily conversation. It is designed to provide a robust foundation for building general-purpose translation tools and AI assistants capable of handling diverse contexts in the Khmer language.

Your dedication to advancing Khmer AI is truly inspiring—keep up the great work!

Dataset Summary

  • —Total Rows: 366,174 clean entries
  • —Language: Khmer (km) and English (en)
  • —Focus: General Purpose and Multi-Domain Translation
  • —Task: Translation

Data Schema

ColumnDescription
engThe source text or phrase in English
khThe corresponding translation in Khmer

ការបកប្រែភាសាខ្មែរ-អង់គ្លេស សរុប (Translate Khmer All)

សូមស្វាគមន៍មកកាន់បណ្តុំទិន្នន័យរបស់ SeyhaLite។ Dataset នេះត្រូវបានរៀបចំ និងសម្អាតយ៉ាងសម្រិតសម្រាំង ដើម្បីគាំទ្រដល់ការអភិវឌ្ឍប្រព័ន្ធ AI និងការបកប្រែភាសាខ្មែរឱ្យកាន់តែមានប្រសិទ្ធភាព និងត្រឹមត្រូវតាមក្បួនខ្នាត។

ចក្ខុវិស័យរបស់គម្រោង

ខ្ញុំសង្ឃឹមយ៉ាងមុតមាំថា Dataset នេះនឹងជួយឱ្យគម្រោងរបស់អ្នកទទួលបានជោគជ័យ! នេះគឺជាបណ្តុំទិន្នន័យដ៏ធំទូលាយដែលប្រមូលផ្តុំពីវិស័យជាច្រើន រួមមានធុរកិច្ច បច្ចេកវិទ្យា វេជ្ជសាស្ត្រ ច្បាប់ និងការសន្ទនាប្រចាំថ្ងៃ។ វាត្រូវបានបង្កើតឡើងដើម្បីផ្តល់នូវមូលដ្ឋានគ្រឹះដ៏រឹងមាំសម្រាប់បង្កើត ឧបករណ៍បកប្រែទូទៅ និងប្រព័ន្ធ AI ដែលអាចប្រើប្រាស់បានក្នុងបរិបទចម្រុះជាភាសាខ្មែរ។

ការខិតខំប្រឹងប្រែងរបស់អ្នកក្នុងការលើកស្ទួយ AI ភាសាខ្មែរ គឺជារឿងដែលគួរឱ្យកោតសរសើរខ្លាំងណាស់!

សេចក្តីសង្ខេបនៃទិន្នន័យ

  • —ចំនួនជួរ (Rows): ៣៦៦,១៧៤ ជួរ
  • —ភាសា: ខ្មែរ (Khmer) និង អង់គ្លេស (English)
  • —គោលបំណង: ការបកប្រែទូទៅ និងពហុវិស័យ
  • —ប្រភេទភារកិច្ច: Translation

រចនាសម្ព័ន្ធទិន្នន័យ

ជួរឈរ (Column)ការពិពណ៌នា
engអត្ថបទ ឬឃ្លាដើមជាភាសាអង់គ្លេស
khអត្ថបទដែលបានបកប្រែជាភាសាខ្មែរ

Support & Community

If you find this dataset helpful, please consider supporting my work to keep the Khmer AI community growing.

  • —Like & Follow this repository for updates.
  • —Share it with fellow developers.
  • —Donate: Your support helps me maintain and release more high-quality datasets.
**Binance Pay****Scan QR Code**
ID: 792183661<img src="https://i.ibb.co/hFjw31XH/image.png" width="200" alt="Binance QR Code">

Created with care by SeyhaLite**