unigram
Datasets
All datasets matching “unigram”unigrams
unigrams
A frequency-weighted lexical dataset containing 98,906 Hmar unigrams and active loanwords with occurrence counts compiled directly from the Foundation's verified Hmar corpus (over 4.14 million words).
Maintained by the Hmar Heritage Foundation (hmarheritage.pages.dev).
Overview
Language: Hmar (hmr, ISO 639-3, Glottolog: hmar1241)
Family: Zo Languages
Volume: 98,906 unigram tokens (compiled across 4,146,783 words)
Format: JSONL (data/train.jsonl)… See the full description on the dataset page: https://huggingface.co/datasets/hmar-heritage-org/unigrams.LogicNLI-reimpl
Dataset Card for "lnli"
More Information needed
details_msu-rcc-lair__ruadapt_solar_10.7_darulm_unigram_proj_init_twostage_v1kgz_dataset_chunked_medium_unigram_64kloam_1e_minus_02_unigrams_v2fineweb2-fil-1B-unigram-tokenized
