CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Cognitive-Lab /Aya_Kannadagated Aya_Kannada This Dataset is curated from the original Aya-Collection dataset that was open-sourced by Cohere under the Apache-2.0 license. The Aya Collection is a massive multilingual collection comprising 513 million instances of prompts and completions that cover a wide range of tasks. This collection uses instruction-style templates from fluent speakers and applies them to a curated list of datasets. It also includes translations of instruction-style datasets into 101 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/Aya_Kannada.tabular1M<n<10M0 likes49 downloads3y agoHugging Face02InfoBayAI /Kannada-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Kannada STEM textbook data, containing 127 books and 6.64 million words, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, problem-solving, and concept learning in Kannada. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for deeper… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-STEM-Textbook-Dataset.tabulartext-generation10K<n<100K0 likes14 downloads9d agoHugging Face03InfoBayAI /Kannada-Non-STEM-Textbook-DatasetgatedDataset Description: This dataset is a large-scale collection of Kannada Non-STEM textbook data, containing 741 books and 40.95 million words, designed to support the development and training of advanced NLP systems and AI models for language understanding, reasoning, and general knowledge learning in Kannada. Full Dataset Overview This dataset is part of a large-scale multilingual educational corpus containing over 3+ billion words across 5,000+ subjects, supported by interwoven images for… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Kannada-Non-STEM-Textbook-Dataset.tabular10K<n<100K0 likes14 downloads9d agoHugging Face04SayantanJoker /audio_tts_description_kannadatabular1K<n<10K0 likes12 downloads2y agoHugging Face05cazzz307 /malayalam-kannada-tamil-telugu-samam-datasetgated Samam.net Multilingual Dictionary Dataset Dataset Description This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu. Important Note about Script Usage All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.tabular10K<n<100K0 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.