CoolFace
Datasetpublic

zomi-tedim-ai/zomi-tedim-bible-corpus

Zomi Tedim–Burmese Parallel Corpus (Bible Verses) Dataset Summary A parallel text corpus of Bible verses in Tedim (Zomi) and Burmese, extracted from USX source files. This dataset is intended as a foundational resource for Tedim-language NLP research, including machine translation, language modeling, and text generation for this low-resource language. Dataset Structure Each record contains the following fields: Field Description id Unique… See the full description on the dataset page: https://huggingface.co/datasets/zomi-tedim-ai/zomi-tedim-bible-corpus.

sourceHugging Facecc-by-4.0updated 23d agoView on Hugging Face
0likes60downloads
Dataset Card

Zomi Tedim–Burmese Parallel Corpus (Bible Verses)

Dataset Summary

A parallel text corpus of Bible verses in Tedim (Zomi) and Burmese, extracted from USX source files. This dataset is intended as a foundational resource for Tedim-language NLP research, including machine translation, language modeling, and text generation for this low-resource language.

Dataset Structure

Each record contains the following fields:

FieldDescription
idUnique identifier for the verse
bookBible book code (e.g. 1CH, ACT, AMO)
chapterChapter number
verseVerse number
tedimVerse text in Tedim (Zomi)
burmeseVerse text in Burmese
englishVerse text in English (where available)
sourceSource file/reference
notesAdditional notes, if any

Data is provided in JSON and JSONL formats.

Data Source

Verses were extracted from 20 USX Bible book files:

1CH, 1CO, 1JN, 1KI, 1PE, 1SA, 1TH, 1TI, 2CH, 2CO, 2JN, 2KI, 2PE, 2SA, 2TH, 2TI, 3JN, ACT, AMO, COL

This yielded 7,289 aligned verse pairs.

Intended Use

  • —Low-resource language modeling for Tedim (Zomi)
  • —Machine translation (Tedim ↔ Burmese ↔ English)
  • —Linguistic and NLP research on Tedim/Chin languages
  • —Foundation for future Tedim datasets covering non-religious domains

Limitations

  • —Vocabulary and register are limited to Biblical text and may not reflect everyday spoken or written Tedim.
  • —Coverage is currently limited to 20 of the 66 Bible books; more books will be added over time.
  • —English alignment may be incomplete for some verses.

Roadmap

This dataset is part of an ongoing effort to build broader Tedim language resources, in collaboration with the zolai-ai project.

Creator

Thang Deih Piang — Zomi GPT AI creator

Citation

If you use this dataset, please cite it as:

Zomi Tedim–Burmese Parallel Bible Corpus

Contact

Maintained by Thang Deih Piang. For collaboration inquiries, please open a discussion on this dataset's Community tab.