hmar
Datasets
All datasets matching “hmar”corpus-archive
corpus-archive
[!WARNING]
Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets.
This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.zo-bible
zo-bible
A sentence-aligned parallel Bible corpus covering 8 closely related Zo speech varieties and English across 10 translation versions (30,974 canonical verses).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr), Mizo (lus), Paite (pck), Vaiphei (vap), Thadou (tcz), Gangte (gnb), Zou (zom), English (eng)
Family: Zo Languages
Volume: 30,974 verse anchors across 66 canonical books (10 translation editions)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/zo-bible.wordlist
wordlist
A structured digital lexicon containing 43,509 Hmar words, phrases, definitions, and translations compiled from five lexicographical sources.
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 43,509 lexical entries across 5 dictionary files
Format: JSON (data/dictionary-001.json to data/dictionary-005.json)
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/wordlist.unigrams
unigrams
A frequency-weighted lexical dataset containing 98,906 Hmar unigrams and active loanwords with occurrence counts compiled directly from the Foundation's verified Hmar corpus (over 4.14 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 98,906 unigram tokens (compiled across 4,146,783 words)
Format: JSONL (data/train.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/unigrams.sentences
sentences
A multi-register sentence corpus for the Hmar language (hmr, ISO 639-3), containing 260,492 train sentences and 5,316 evaluation sentences (~4.15 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 260,492 train sentences | 5,316 evaluation sentences (265,808 total, ~4.15M words)
Validation: 100% verified via hmaraniam… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/sentences.hma_rc26_object_images_dataset
HMA RC26 Incheon RGB Image Dataset
Dataset Overview
This dataset combines synthetic images and annotated RGB images captured with ASUS Xtion Pro Live and ORBBEC Gemini 336L. It was used to train a YOLOv8 segmentation model for object recognition at RoboCup@Home 2026 in Incheon (RC26) [https://github.com/RoboCupAtHome/Incheon2026].
Data Collection
Synthetic Data
The simulator generated 1,000 images using domain randomization. Each… See the full description on the dataset page: https://huggingface.co/datasets/HibikinoMusashiHome/hma_rc26_object_images_dataset.
