SeyhaLite/Translate-English-Khmer-All
Translate Khmer All Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems. Project Vision I hope this dataset helps your project succeed! This is a comprehensive collection that aggregates data from multiple domains—including business, technology, medical, legal, and daily conversation. It is designed to provide a robust foundation… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Translate-English-Khmer-All.
Translate Khmer All
Welcome to the SeyhaLite collection. This dataset has been meticulously curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) and Translation systems.
Project Vision
I hope this dataset helps your project succeed! This is a comprehensive collection that aggregates data from multiple domains—including business, technology, medical, legal, and daily conversation. It is designed to provide a robust foundation for building general-purpose translation tools and AI assistants capable of handling diverse contexts in the Khmer language.
Your dedication to advancing Khmer AI is truly inspiring—keep up the great work!
Dataset Summary
- Total Rows: 366,174 clean entries
- Language: Khmer (km) and English (en)
- Focus: General Purpose and Multi-Domain Translation
- Task: Translation
Data Schema
ការបកប្រែភាសាខ្មែរ-អង់គ្លេស សរុប (Translate Khmer All)
សូមស្វាគមន៍មកកាន់បណ្តុំទិន្នន័យរបស់ SeyhaLite។ Dataset នេះត្រូវបានរៀបចំ និងសម្អាតយ៉ាងសម្រិតសម្រាំង ដើម្បីគាំទ្រដល់ការអភិវឌ្ឍប្រព័ន្ធ AI និងការបកប្រែភាសាខ្មែរឱ្យកាន់តែមានប្រសិទ្ធភាព និងត្រឹមត្រូវតាមក្បួនខ្នាត។
ចក្ខុវិស័យរបស់គម្រោង
ខ្ញុំសង្ឃឹមយ៉ាងមុតមាំថា Dataset នេះនឹងជួយឱ្យគម្រោងរបស់អ្នកទទួលបានជោគជ័យ! នេះគឺជាបណ្តុំទិន្នន័យដ៏ធំទូលាយដែលប្រមូលផ្តុំពីវិស័យជាច្រើន រួមមានធុរកិច្ច បច្ចេកវិទ្យា វេជ្ជសាស្ត្រ ច្បាប់ និងការសន្ទនាប្រចាំថ្ងៃ។ វាត្រូវបានបង្កើតឡើងដើម្បីផ្តល់នូវមូលដ្ឋានគ្រឹះដ៏រឹងមាំសម្រាប់បង្កើត ឧបករណ៍បកប្រែទូទៅ និងប្រព័ន្ធ AI ដែលអាចប្រើប្រាស់បានក្នុងបរិបទចម្រុះជាភាសាខ្មែរ។
ការខិតខំប្រឹងប្រែងរបស់អ្នកក្នុងការលើកស្ទួយ AI ភាសាខ្មែរ គឺជារឿងដែលគួរឱ្យកោតសរសើរខ្លាំងណាស់!
សេចក្តីសង្ខេបនៃទិន្នន័យ
- ចំនួនជួរ (Rows): ៣៦៦,១៧៤ ជួរ
- ភាសា: ខ្មែរ (Khmer) និង អង់គ្លេស (English)
- គោលបំណង: ការបកប្រែទូទៅ និងពហុវិស័យ
- ប្រភេទភារកិច្ច: Translation
រចនាសម្ព័ន្ធទិន្នន័យ
Support & Community
If you find this dataset helpful, please consider supporting my work to keep the Khmer AI community growing.
- Like & Follow this repository for updates.
- Share it with fellow developers.
- Donate: Your support helps me maintain and release more high-quality datasets.
Created with care by SeyhaLite**
