CoRover
Datasets
All datasets matching “CoRover”hplt
HPLT Indic Language Corpus
This repository contains selected Indic-language data from the HPLT Monolingual Dataset 3.0, prepared and hosted by CoRover for large-scale generative language model pretraining and multilingual NLP research.
The dataset contains raw HPLT data for multiple Indian languages and scripts.
Languages
The repository currently contains the following language/script datasets:
Language
Language Code
Script
Directory
Bengali
ben
Bengali… See the full description on the dataset page: https://huggingface.co/datasets/CoRover/hplt.wikipediaThe CoRover Wikipedia Corpus is a cleaned and structured collection of Wikipedia articles prepared for large language model (LLM) pretraining, language modeling, and multilingual NLP research.
