CoolFace
Datasetpublicgated

InfoBayAI/Kannada-Non-STEM-Textbook-Dataset

Dataset Description: This dataset is a large-scale collection of Kannada Non-STEM textbook data, containing 741 books and 40.95 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Kannada. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-Non-STEM-Textbook-Dataset.

sourceHugging Facecc-by-4.0updated 9d agoView on Hugging Face
0likes14downloads
fileBahamani Samrajya.parquet285 KBdownload
fileBeluru TalukinaDevalayagala Kale Mattu Vatushilpa.parquet451 KBdownload
fileKarnataka Lochana.parquet208 KBdownload
fileSampoorna Bharathada Ithihasa.parquet274 KBdownload
fileShala Shikshana Matthu Shikshakaru.parquet389 KBdownload

InfoBayAI/Kannada-Non-STEM-Textbook-Dataset · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.