CoolFace
Datasetpublicgated

MWirelabs/northeast-india-voices

Northeast India Voices A multilingual transcribed speech corpus covering five indigenous and regional languages of Northeast India: Khasi, Garo, Mizo, Nagamese, and Kokborok. Dataset Summary Language Family Utterances Khasi Austroasiatic 14,974 Nagamese Indo-Aryan (Creole) 7,386 Mizo Tibeto-Burman 7,250 Kokborok Tibeto-Burman 3,278 Garo Tibeto-Burman 1,612 Total 34,500 Data Collection Recorded by native speakers across… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-india-voices.

sourceHugging Facecc-by-4.0updated 3mo agoView on Hugging Face
0likes51downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.