CoolFace
Datasetpublic

Arjun-G-Ravi/malayalam-sangraha

This is a cleaned version of the malayalam subset of sangraha dataset. This only contains the human verified part of the dataset(which is high quality data obtained from Indic language PDFs, transcribed data from various Indic language videos, podcasts, movies, courses, etc.) The csv dataset has around 6.3M rows, accounting to 32.8 GB. I've also removed the doc_id provided in the dataset, making this ideal for pretraining malayalam LLM. For pretraining, I recommend using this dataset along… See the full description on the dataset page: https://huggingface.co/datasets/Arjun-G-Ravi/malayalam-sangraha.

sourceHugging Facecc-by-4.0updated 10mo agoView on Hugging Face
1likes26downloads
settings

This repository belongs to Arjun-G-Ravi on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namemalayalam-sangraha
visibilitypublic
licencecc-by-4.0
gatedno
ownerArjun-G-Ravi
Account settings
Arjun-G-Ravi/malayalam-sangraha · CoolFace