CoolFace
Datasetpublic

l3cube-pune/marathi-pos-tagger

L3Cube-MahaPOS: Marathi Part-of-Speech Tagging Dataset Dataset Description L3Cube-MahaPOS is one of the first large-scale, manually annotated Part-of-Speech (POS) tagging datasets for Marathi — an Indo-Aryan language spoken by over 83 million people. The dataset comprises 32,354 sentences sourced from Marathi news text and annotated with a 16-tag scheme aligned with the Universal Dependencies (UD) v2 framework. This dataset is part of the L3Cube-MahaNLP family of… See the full description on the dataset page: https://huggingface.co/datasets/l3cube-pune/marathi-pos-tagger.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes51downloads
Dataset Card

L3Cube-MahaPOS: Marathi Part-of-Speech Tagging Dataset

Dataset Description

L3Cube-MahaPOS is one of the first large-scale, manually annotated Part-of-Speech (POS) tagging datasets for Marathi — an Indo-Aryan language spoken by over 83 million people. The dataset comprises 32,354 sentences sourced from Marathi news text and annotated with a 16-tag scheme aligned with the Universal Dependencies (UD) v2 framework.

This dataset is part of the L3Cube-MahaNLP family of Marathi NLP resources, which includes L3Cube-MahaSent, L3Cube-MahaNER, L3Cube-MahaSent-MD, and L3Cube-MahaSocialNER.

For more details refer our MahaPOS paper.<br> Model trained on this dataset is available here l3cube-pune/marathi-pos-tagger.


Dataset Summary

PropertyValue
LanguageMarathi (mr)
TaskToken Classification / POS Tagging
Annotation SchemeUniversal Dependencies v2 (16 tags)
DomainNews text
Total Sentences32,354
Total Tokens472,459
Avg. Sentence Length14.6 tokens
LicenseCC BY 4.0

Splits

SplitSentencesTokensAvg. Length
Train22,652332,41814.7
Validation4,84871,16314.7
Test4,85468,87814.2
Total32,354472,45914.6

Splits are stratified to maintain proportional representation of all 16 tag classes and balanced domain coverage across train, validation, and test.


Tag Set & Distribution

TagDescriptionToken Count% of CorpusEval
NOUNCommon noun144,59731.9%
VERBMain verb58,83713.0%
PUNCTPunctuation50,04411.0%
ADPAdposition39,8978.8%
ADJAdjective35,5277.8%
ADVAdverb28,1276.2%
AUXAuxiliary verb23,1555.1%
NUMNumeral15,7113.5%
DETDeterminer14,1453.1%
PRONPronoun14,0913.1%
PROPNProper noun10,1262.2%
CCONJCoordinating conjunction9,6002.1%
PARTParticle4,5911.0%
SCONJSubordinating conjunction3,2540.7%
INTJInterjection1,4400.3%
POSTPPostposition23<0.01%❌*
Total453,165
`POSTP` is present in training data but excluded from primary macro-F1 evaluation* due to critically low test-set support (3 tokens). Any single misclassification shifts its F1 by over 33 percentage points, making the metric statistically unreliable.
X (foreign/unclassifiable tokens): used during annotation but excluded entirely from all data splits. Sentences containing at least one X-tagged token were removed.

Data Collection

Source

Raw text was collected from Marathi news portals covering diverse topical domains:

  • Politics
  • Sports
  • Culture
  • Technology
  • Local affairs

News text was chosen for three reasons:

  1. 1.Formal register with consistent orthography — facilitates reproducible preprocessing
  2. 2.Broad vocabulary coverage including technical and domain-specific terminology
  3. 3.Captures contemporary Marathi as actively used in published media

Collection Process

  • HTML content extracted using domain-specific scrapers
  • Boilerplate elements (navigation, ads, metadata) removed via heuristic filtering
  • Duplicate and near-duplicate articles identified using MinHash locality-sensitive hashing and discarded
  • Sentence boundary detection using a rule-based system tuned to Marathi punctuation conventions (Devanagari danda as primary terminator)

Preprocessing Pipeline

A four-stage pipeline was applied uniformly across all splits:

Stage 1 — Unicode Normalisation

  • All text converted to Unicode NFC form
  • ZWJ/ZWNJ characters retained where they alter character appearance, removed where semantically redundant
  • Visually identical but code-point-distinct Devanagari characters (e.g., anusvara variants) canonicalised

Stage 2 — Tokenisation

  • Devanagari-aware tokeniser splitting on whitespace and explicit punctuation
  • Compound postpositions and clitics handled via a native-speaker-developed exception lexicon
  • English words embedded in Marathi text treated as single tokens

Stage 3 — Noise Filtering

  • Tokens normalised to placeholders: numerals with currency → <NUM>, URLs/emails → <URL>
  • Emoticons/emoji removed
  • Sentences with fewer than 3 tokens or more than 120 tokens excluded

Stage 4 — POS Annotation

  • Cleaned tokens passed to the annotation workflow

Edge-Case Rules

Token TypeDecision
Pure numeral (e.g., 42)Retain as NUM
Mixed numeral-unit (e.g., 42km)Split: NUM + NOUN
URL / emailReplace with <URL>; tag as NOUN
Emoticon / emojiRemove from sentence
English word in Marathi contextRetain; tag per context
Danda (sentence-final )Retain as PUNCT
Ellipsis ()Normalise to single PUNCT

Annotation Process

Team

Annotation was performed entirely manually by a team of Marathi-proficient annotators.

Procedure

  1. 1.Guidelines familiarisation — All annotators trained on written guidelines adapted from the UD annotation manual, supplemented with Marathi-specific decision trees covering postpositional clitics, verbal compounds, and code-mixed tokens
  2. 2.Data division — The full corpus of 32,354 sentences divided into roughly equal portions distributed across team members
  3. 3.Independent tagging — Each annotator tagged their assigned sentences independently
  4. 4.Conflict resolution — Disputed and ambiguous labels resolved through group discussion; majority label adopted
  5. 5.Amendment log — Decisions recorded and used to update shared guidelines iteratively across batches
  6. 6.Final validation — Automatic consistency checker flagged tokens whose label conflicted with the majority label for the same surface form in unambiguous contexts

Primary Sources of Disagreement

  • ADJ/ADV ambiguity in participial constructions — several adjectival stems function as adverbs in pre-verbal position without case agreement, producing surface-identical forms
  • NOUN/PROPN boundaries for institutionalised proper nouns — Marathi lacks mandatory capitalisation, making proper noun detection purely contextual

Sample Annotations

Example 1:
भारत/PROPN  हा/PRON   एक/NUM   सुंदर/ADJ  देश/NOUN  आहे/VERB  ./PUNCT
(India       this      one      beautiful  country   is        .)
→ "India is a beautiful country."

Example 2:
नागपूर/PROPN  येथे/ADP  मोठा/ADJ  कार्यक्रम/NOUN  झाला/VERB  ./PUNCT
(Nagpur        there     big       event            took-place  .)
→ "A big event took place in Nagpur."

Example 3:
राम/PROPN  आणि/CCONJ  सीता/PROPN  वनात/NOUN  गेले/VERB  ./PUNCT
(Ram        and         Sita        forest      went        .)
→ "Ram and Sita went to the forest."

Benchmark Results

Models evaluated on the L3Cube-MahaPOS test set:

ModelAccuracyPrecisionRecallMacro-F1 (15 tags)
HMM72.1%63.4%61.8%63.1%
CRF80.6%72.9%71.5%72.8%
BiLSTM-CRF84.3%76.8%75.4%76.6%
BiLSTM+CharCNN85.7%78.2%77.1%78.3%
MuRIL86.9%79.4%78.8%79.7%
MahaPOS-BERT88.67%88.63%76.49%81.67%
Primary metric is Macro-F1 over 15 tags (POSTP excluded). MahaPOS-BERT 16-tag F1 = 76.57%.

Related Datasets (L3Cube-MahaNLP Family)

DatasetTaskSize
L3Cube-MahaSentSentiment Analysis~16K tweets
L3Cube-MahaNERNamed Entity Recognition~25K sentences
L3Cube-MahaSent-MDMulti-domain Sentiment~60K samples
L3Cube-MahaSocialNERSocial Media NER
L3Cube-MahaPOSPOS Tagging32,354 sentences

Limitations

  • Domain: Formal news text only. Performance may degrade on informal registers, social media, legal text, or scientific Marathi.
  • Code-mixing: English-origin tokens are treated as single-token borrowings. Systematic intra-sentential code-mixing with Hindi/English is not explicitly modelled.
  • POSTP: Only 23 instances corpus-wide. Future versions should either merge POSTP with ADP or collect targeted postpositional constructions to reach statistically viable support.
  • INTJ: ~0.3% of tokens; suffers from data sparsity. Stratified minority-class collection recommended in future work.
  • Proper nouns: Marathi's lack of capitalisation makes PROPN/NOUN disambiguation purely contextual, leading to the lowest F1 among well-supported classes.

Citation

bibtex
@article{ingle2026l3cube,
  title={L3Cube-MahaPOS: A Marathi Part-of-Speech Tagging Dataset and BERT Models},
  author={Ingle, Hariom and Ghode, Ronit and Gondkar, Ishwari and Harad, Jidnyasa and Joshi, Raviraj},
  journal={arXiv preprint arXiv:2606.24825},
  year={2026}
}

Acknowledgements

This work was carried out under the mentorship of L3Cube Labs, Pune. We thank the Universal Dependencies consortium for maintaining the annotation framework we adapted, and the L3Cube-Labs team for making marathi-bert-v2 publicly available. This work is part of the L3Cube-MahaNLP project.