CoolFace
Datasetpublicgated

MWirelabs/northeast-languages-test-set

Northeast Languages Test Set A curated test set of 500 deduplicated sentences per language for evaluating language models on Northeast Indian languages. Languages This dataset contains test data for 9 Northeast Indian languages: Assamese (asm) - 500 sentences Garo (grt) - 500 sentences Khasi (kha) - 500 sentences Kokborok (trp) - 500 sentences Meitei (mni) - 500 sentences Mizo (lus) - 500 sentences Naga (nag) - 500 sentences Nyishi (njz) - 500 sentences Pnar… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-languages-test-set.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes7downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
MWirelabs/northeast-languages-test-set · CoolFace