datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bihari-languages-upos
Bihari Languages UPOS Dataset
This dataset provides Part-of-Speech (POS) tags for Angika (anp), Magahi (mag), and Bhojpuri (bho), parallelly aligned with Hindi (hi). The annotations follow the Universal Dependencies (UD) Universal Part-of-Speech (UPOS) standard.
This work is part of research conducted at the Department of Computer Science and Engineering, IIT Bombay.
Dataset Details
Languages: Angika, Magahi, Bhojpuri, Hindi
Task: Token Classification (Part-of-Speech… See the full description on the dataset page: https://huggingface.co/datasets/snjev310/bihari-languages-upos.news
