CoolFace
Datasetpublic

kalixlouiis/HFcourse-english-burmese-parallel-corpus

HFcourse-English-Burmese-Parallel-Corpus Dataset Description Dataset Summary The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich… See the full description on the dataset page: https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
12likes23downloads
Dataset Card

HFcourse-English-Burmese-Parallel-Corpus

Table of Contents

Dataset Description

Dataset Summary

The HFcourse-English-Burmese-Parallel-Corpus is a collection of English and Burmese parallel sentence pairs, specifically designed to support research and development in Neural Machine Translation (NMT) for the Myanmar language. It comprises 2,503 meticulously aligned sentence pairs, extracted from the subtitles of the Hugging Face Course videos. This dataset aims to enrich the available resources for Myanmar language processing in AI and Machine Learning fields, thereby fostering advancements in these areas.

Languages

The dataset contains text in two languages:

  • —English (en)
  • —Burmese (my)

Purpose

This dataset was created with the primary goals of:

  1. 1.Training Neural Machine Translation (NMT) Models: Providing high-quality parallel data to train and evaluate NMT systems for English-Burmese translation.
  2. 2.Enhancing Myanmar Language Resources: Contributing to the growing body of computational linguistic resources for the Burmese language, which is crucial for its development in AI and ML.
  3. 3.Facilitating Research and Development: Supporting researchers and developers in creating more robust and accurate language technologies for Myanmar speakers.

Dataset Structure

Data Instances

Each instance in the dataset represents an aligned English and Burmese sentence pair. Here are examples:

json
{
  "id": 1,
  "English": "Welcome to the Hugging Face Course.",
  "Burmese": "Hugging Face သင်တန်းမှ ကြိုဆိုပါတယ်။"
}
json
{
  "id": 2,
  "English": "This course has been designed to teach you all about the Hugging Face ecosystem, how to use the dataset and model hub as well as all our open-source libraries.",
  "Burmese": "ဒီသင်တန်းကို Hugging Face ရဲ့ ဂေဟစနစ် အကြောင်း၊ dataset နဲ့ model hub တွေကို ဘယ်လိုအသုံးပြုရမလဲ၊ ကျွန်ုပ်တို့ရဲ့ open-source library တွေအားလုံးကို ဘယ်လိုသုံးရမလဲဆိုတာတွေကို သင်ကြားပေးဖို့ ရေးဆွဲထားတာပါ။"
}
json
{
  "id": 3,
  "English": "Here is the Table of Contents.",
  "Burmese": "ဒီမှာတော့ သင်တန်းရဲ့ အကြောင်းအရာများ အညွှန်း ဖြစ်ပါတယ်။"
}

Data Fields

The dataset is provided in a CSV format with the following fields:

  • —id: (int) A unique identifier for each sentence pair.
  • —English: (str) The English sentence.
  • —Burmese: (str) The corresponding Burmese translation.

Data Splits

The dataset currently consists of a single split:

  • —Total Instances: 2,503

As it's primarily intended for training and research, users may create their own training, validation, and test splits as needed.

Dataset Creation

Source Data

The data originates from the subtitles of the 🤗 Hugging Face Course videos (Video 80). The original playlist can be found here. The subtitles were extracted and then meticulously converted into paired English-Burmese sentences.

Annotations

The dataset does not contain manual annotations beyond the initial translation and alignment process derived from the original subtitles. The pairing of English and Burmese sentences is considered a form of annotation, where each Burmese sentence is the translation of its corresponding English sentence.

Cleaning Steps

The raw subtitle data underwent the following cleaning procedures:

  1. 1.Duplicate Removal: Identified and eliminated any exact duplicate sentence pairs to ensure data uniqueness.
  2. 2.Structural Repair: Addressed and resolved CSV parsing errors to maintain data integrity and consistent formatting.
  3. 3.Manual Review: A thorough manual review was conducted to ensure the accuracy of the sentence pairings and overall data quality.

Considerations for Using the Data

Domain Specificity

This dataset is highly focused on technical terminology, particularly in the fields of machine learning, large language models (LLMs), and general artificial intelligence concepts. Users should be aware that while it provides excellent coverage for this domain, its applicability to other domains (e.g., general conversation, news, legal) might be limited without further domain adaptation.

License

The HFcourse-English-Burmese-Parallel-Corpus is released under the Creative Commons Attribution 4.0 International (CC BY 4.0) License.

This means you are free to:

  • —Share: copy and redistribute the material in any medium or format.
  • —Adapt: remix, transform, and build upon the material for any purpose, even commercially.

Under the following terms:

  • —Attribution: You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.

When using this dataset, please provide attribution to the original creator and the dataset itself. For example, by citing this dataset card.

Citation

If you use this dataset in your research or project, please cite it as follows:

bibtex
@dataset{hfcourse_english_burmese_parallel_corpus_2025,
  title        = {HFcourse-English-Burmese-Parallel-Corpus},
  author       = {Khant Sint Heinn},
  year         = {2025},
  publisher    = {Hugging Face Datasets},
  url          = {https://huggingface.co/datasets/kalixlouiis/HFcourse-english-burmese-parallel-corpus},
  license      = {CC BY 4.0}
}

Contact

For any questions or inquiries regarding this dataset, please feel free to contact the dataset creator: Khant Sint Heinn(Kalix Louis).

📧 - kalixlouiis@gmail.com

About the Author

Khant Sint Heinn, working under the name Kalix Louis, is a Machine Learning Engineer focused on Natural Language Processing (NLP), data foundations, and open-source AI development. His work is centered on improving support for the Burmese (Myanmar) language in modern AI systems by building high-quality datasets, practical tools, and scalable infrastructure for language technology.

He is currently the Lead Developer at DatarrX, where he develops data pipelines, manages large-scale data collection workflows, and helps create open-source resources for researchers, developers, and organizations. His experience includes data engineering, web scripting, dataset curation, and building systems that support real-world machine learning applications.

Khant Sint Heinn is especially interested in advancing low-resource languages and making AI more accessible to underrepresented communities. Through his open-source contributions, he works to strengthen the Burmese (Myanmar) tech ecosystem and provide reliable building blocks for future language models, search systems, and intelligent applications.

His goal is simple: to turn limited language resources into practical opportunities through clean data, useful tools, and community-driven innovation.

Connect with the Author: GitHub | Hugging Face | Kaggle