l3cube-pune/marathi-short-topic-bert-pruned-distill
Marathi-Short-Doc-Topic-BERT-Pruned
Marathi-Doc-Topic-BERT model is an l3cube-pune/marathi-topic-all-doc-v2 model pruned and distilled on Marathi documents from the L3Cube-IndicNews Corpus [dataset link]https://github.com/l3cube-pune/indic-nlp. <br> This dataset consists of sub-datasets like LDC (Long Document Classification), LPC (Long Paragraph Classification), and SHC (Short Headlines Classification), each having different document lengths. <br> This model is trained on the SHC dataset.
More details on the dataset, models, and baseline results can be found in our [paper]https://arxiv.org/abs/2401.02254
Citing:
@article{mirashi2024l3cube,
title={L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages},
author={Mirashi, Aishwarya and Sonavane, Srushti and Lingayat, Purva and Padhiyar, Tejas and Joshi, Raviraj},
journal={arXiv preprint arXiv:2401.02254},
year={2024}
}Other document topic models for different Indic languages are listed below: <br> <a href='https://huggingface.co/l3cube-pune/hindi-topic-all-doc'> Hindi-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/bengali-topic-all-doc'> Bengali-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/telugu-topic-all-doc'> Telugu-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/tamil-topic-all-doc'> Tamil-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/gujarati-topic-all-doc'> Gujarati-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/kannada-topic-all-doc'> Kannada-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/odia-topic-all-doc'> Odia-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/malayalam-topic-all-doc'> Malayalam-Doc-Topic-BERT </a> <br> <a href='https://huggingface.co/l3cube-pune/punjabi-topic-all-doc'> Punjabi-Doc-Topic-BERT </a> <br>
