datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
zo-bible
zo-bible
A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng)
Family: Zo Languages
Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.unigrams
unigrams
A frequency-weighted lexical dataset containing 98,906 Hmar unigrams and active loanwords with occurrence counts compiled directly from the Foundation's verified Hmar corpus (over 4.14 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 98,906 unigram tokens (compiled across 4,146,783 words)
Format: JSONL (data/train.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/unigrams.numeral-words
numeral-words
A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 999,999 parallel rows (1 to 999,999)
Format: Compressed JSONL (data/train-*.jsonl.gz)
License: Apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.
