datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
corpus-archive
corpus-archive
[!WARNING]
Experimental Dataset Architecture: The repository structure, metadata tiers, category taxonomies, and catalog indexing formats are currently under active design and evaluation. All specifications, metadata keys, and JSON schemas detailed below represent representational examples and intended targets.
This repository serves as a structured digital textual archive preserving Hmar literature, historical accounts, school textbooks, dictionaries, parallel… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/corpus-archive.unigrams
unigrams
A frequency-weighted lexical dataset containing 98,906 Hmar unigrams and active loanwords with occurrence counts compiled directly from the Foundation's verified Hmar corpus (over 4.14 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 98,906 unigram tokens (compiled across 4,146,783 words)
Format: JSONL (data/train.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/unigrams.sentences
sentences
A multi-register sentence corpus for the Hmar language (hmr, ISO 639-3), containing 260,492 train sentences and 5,316 evaluation sentences (~4.15 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 260,492 train sentences | 5,316 evaluation sentences (265,808 total, ~4.15M words)
Validation: 100% verified via hmaraniam… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/sentences.numeral-words
numeral-words
A 3-way parallel digital dataset containing 999,999 spelled-out Hmar & English number words mapped in sequential numerical order (1 to 999,999).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Languages: Hmar (hmr, ISO 639-3, Glottolog: hmar1241), English (en)
Family: Zo Languages
Volume: 999,999 parallel rows (1 to 999,999)
Format: Compressed JSONL (data/train-*.jsonl.gz)
License: Apache-2.0
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/numeral-words.hmingtluon
hmingtluon
A procedural synthetic anthroponyms dataset containing 10,193,568 (10.19 million) unique Hmar full names, designed for Named Entity Recognition (NER), entity resolution, synthetic data augmentation, and anthroponymic research.
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 10,193,568 unique records across 3 splits (Train / Validation /… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/hmingtluon.llm4hls
