CoolFace
Datasetpublicgated

MWirelabs/northeast-languages-test-set

Northeast Languages Test Set A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages. Languages This dataset contains test data for 9 Northeast Indian languages: Assamese (asm) - 500 sentences Garo (grt) - 500 sentences Khasi (kha) - 500 sentences Kokborok (trp) - 500 sentences Meitei (mni) - 500 sentences Mizo (lus) - 500 sentences Naga (nag) - 500 sentences Nyishi (njz) - 500 sentences Pnar… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.

sourceHugging Facecc-by-4.0updated 4mo agoView on Hugging Face
0likes7downloads
settings

This repository belongs to MWirelabs on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namenortheast-languages-test-set
visibilitypublic
licencecc-by-4.0
gatedyes
ownerMWirelabs
Account settings
MWirelabs/northeast-languages-test-set · CoolFace