DatarrX/burmese-VOA
Dataset Card for Burmese VOA News Dataset This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category. Dataset Summary The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/burmese-VOA.
Dataset Card for Burmese VOA News Dataset
This dataset is a comprehensive collection of Burmese news articles crawled from Voice of America (VOA) Burmese. It is specifically curated and processed for Natural Language Processing (NLP) tasks, focusing on high-quality news content, including the "Science and Technology" category.
Dataset Summary
The Burmese VOA Dataset contains 270,546 rows of news articles. The data has been meticulously scraped and structured into a machine-readable format (JSONL) to support the development of Burmese Language Models (LLMs), text summarization, and sentiment analysis.
- Source: VOA Burmese
- Organization: DatarrX
- Curator & Developer: Khant Sint Heinn (Kalix Louis)
- Format: JSONL
- Size: 270,546 records
- License: Creative Commons Attribution 4.0 International (CC-BY-4.0)
Dataset Structure
Each entry in the dataset contains the following features:
Metadata & Curation
This dataset was created using custom web-scripting tools to filter and extract clean text from the VOA Burmese website. Special attention was given to the Science and Technology section to provide high-value technical vocabulary in the Burmese language.
The curation process involved:
- Web scraping via optimized Python scripts.
- Filtering specific categories for quality control.
- Formatting into standardized JSONL for easy integration with Hugging Face
datasetslibrary.
Usage & Licensing
This dataset is released under the CC-BY-4.0 License.
You are free to:
- Share: Copy and redistribute the material in any medium or format.
- Adapt: Remix, transform, and build upon the material for any purpose, even commercially.
Under the following terms:
- Attribution: You must give appropriate credit to the creator (Khant Sint Heinn) and the organization (DatarrX). You should provide a link to the dataset and indicate if changes were made.
Citation & Credit
If you use this dataset in your research or production environment, please credit the author as follows:
Dataset: Burmese VOA Dataset (2026) Curated by: Khant Sint Heinn (Kalix Louis) Organization: DatarrX Source: https://huggingface.co/datasets/DatarrX/burmese-VOA
About the Author
Khant Sint Heinn (Kalix Louis) is a Machine Learning Engineer specializing in NLP and Data Foundations. As the lead developer at DatarrX, he focuses on architecting robust data pipelines and curating high-quality open-source datasets to advance the Burmese Natural Language Processing ecosystem.
His work centers on bridging the gap in low-resource language data through advanced web-scripting, data engineering, and scalable ML foundations.
- Organization: DatarrX
- Expertise: Machine Learning, Natural Language Processing, Data Engineering
Disclaimer: This dataset is intended for educational and research purposes. All original content rights belong to VOA Burmese.
