datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Ultimate-Malayalam-Dataset
About
This is a large dataset that contains a lot of malayalam text. The dataset is created by combining many other malayalam datasets, filtering and cleaning them.
The dataset is ideal to pretrain (or maybe even fine tune) a Large Language Model on Malayalam language.
This is the second version of the dataset where I've added some more data and removed a lot of data, which were very short(less than 75 characters). Having a lot off shorter sentences will significantly lower data… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/Ultimate-Malayalam-Dataset.malayalam-sangrahaThis is a cleaned version of the malayalam subset of sangraha dataset. This only contains the human verified part of the dataset(which is high quality data obtained from Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc.)
The csv dataset has around 6.3M rows, accounting to 32.8 GB. I've also removed the doc_id provided in the dataset, making this ideal for pretraining malayalam LLM.
For pretraining, I recommend using this dataset along… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/malayalam-sangraha.
