Arjun-G-Ravi/malayalam-sangraha
This is a cleaned version of the malayalam subset of sangraha dataset. This only contains the human verified part of the dataset(which is high quality data obtained from Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc.) The csv dataset has around 6.3M rows, accounting to 32.8 GB. I've also removed the doc_id provided in the dataset, making this ideal for pretraining malayalam LLM. For pretraining, I recommend using this dataset along… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/malayalam-sangraha.
This repository belongs to Arjun-G-Ravi on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
