CoolFace
Datasetpublicgated

InfoBayAI/Malayalam-Non-STEM-Textbook-Dataset

Dataset Description: This dataset is a large-scale collection of Malayalam Non-STEM textbook data, containing 149 books and 5.60 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Malayalam. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Malayalam-Non-STEM-Textbook-Dataset.

sourceHugging Facecc-by-4.0updated 8d agoView on Hugging Face
0likes13downloads

No commit history came back for main. The revision may not exist, or the source declined the request.