CoolFace
Datasetpublic

zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412

English–Zomi Parallel Corpus (1.78M) This dataset contains 1.78 million English–Zomi sentence pairs, created to support machine translation, linguistic research, and large‑scale language model training. It is fully open and permissively licensed for commercial and non‑commercial use. 🌐 Linguistic Background: Zomi, Tedim Chin, and ISO Codes Zomi is the endonym (self‑chosen name) of the people and their language.However, Zomi does not yet have an official ISO… See the full description on the dataset page: https://huggingface.co/datasets/zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412.

sourceHugging Facecc0-1.0updated 5mo agoView on Hugging Face
1likes148downloads

zomi-language-corpora/English-Zomi-OPUS_Tatoeba_v20230412 · main · files are served by the source, never re-hosted here