CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01JKA-NLP /unified-kannada-asr-1.0 Dataset Card for "unified-kannada-asr-1.0" More Information needed audio100K<n<1M1 likes573 downloads3y agoHugging Face02anirudhlakhotia /KannadaPreTrainingtext10M<n<100M0 likes397 downloads3y agoHugging Face03SPRINGLab /IndicTTS_Kannada Kannada Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Kannada Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.audiotext-to-speech1K<n<10K4 likes254 downloads2y agoHugging Face04arpit-tiwari /syspin-kannada-ttsaudio10K<n<100K0 likes177 downloads11mo agoHugging Face05Kannada-LLM-Labs /CulturaX-KnThis is a filtered version of the CulturaX dataset only containing samples of Kannada language. The dataset contains total of 1352142 samples. Dataset Structure: { "text": ..., "timestamp": ..., "url": ..., "source": "mc4" | "OSCAR-xxxx", } Data Sample: {'text': "ಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ ಸ್ಮರಣಾರ್ಥ ; ಉಚಿತ ನೋಟ್ ಬುಕ್ ವಿತರಣೆ | Vartha Bharati- ವಾರ್ತಾ ಭಾರತಿ\nಮುದರಂಗಡಿ ಬಿಜೆಪಿ ಗ್ರಾಪಂ ಸದಸ್ಯರ ವಿರುದ್ಧ ಪ್ರತಿಭಟನೆ\nಹೋಮ್ ಕ್ವಾರಂಟೈನ್ ನಿಯಮ ಉಲ್ಲಂಘನೆ: ಪ್ರಕರಣ ದಾಖಲು\nಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/CulturaX-Kn.texttext-generation1M<n<10M1 likes162 downloads3y agoHugging Face06meharuhanzz /OCR-Bench1000-Kannada OCR-Bench1000-Kannada 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Kannada OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category kannada_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Kannada.imageimage-to-text1K<n<10K0 likes162 downloads10d agoHugging Face07shunyalabs /kannada-speech-datasetaudio100K<n<1M0 likes127 downloads1y agoHugging Face08arpit-tiwari /iisc-mile-kannada-asr-corpusaudio100K<n<1M0 likes122 downloads1y agoHugging Face09cdactvm /kannada_new_dataaudio100K<n<1M0 likes119 downloads2y agoHugging Face10SayantanJoker /original_data_kannada_ttsaudio1K<n<10K0 likes99 downloads2y agoHugging Face11cdactvm /kannada_new_data_v2audio10K<n<100K0 likes61 downloads2y agoHugging Face12charanhu /Kannada-Dataset-v01text100K<n<1M2 likes50 downloads3y agoHugging Face13Cognitive-Lab /Aya_Kannadagated Aya_Kannada This Dataset is curated from the original Aya-Collection dataset that was open-sourced by Cohere under the Apache-2.0 license. The Aya Collection is a massive multilingual collection comprising 513 million instances of prompts and completions that cover a wide range of tasks. This collection uses instruction-style templates from fluent speakers and applies them to a curated list of datasets. It also includes translations of instruction-style datasets into 101 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/Aya_Kannada.tabular1M<n<10M0 likes49 downloads3y agoHugging Face14ainlpml-iitp /ICON26-COILD-INDIC-MT-Kannada-Malayalamgated COILD-INDIC-MT 2026 — Kannada–Malayalam Dataset This dataset is provided for the COILD-INDIC-MT 2026 Shared Task, co-located with ICON 2026. The shared task aims to foster research and innovation in Natural Language Processing (NLP) for Indian Languages. This repository contains data specifically for the: Kannada ↔ Malayalam language pair. 🔐 Access to the Dataset This is a restricted and gated dataset. Access is available only to authorized participants of the… See the full description on the dataset page: https://huggingface.co/datasets/ainlpml-iitp/ICON26-COILD-INDIC-MT-Kannada-Malayalam.texttranslation10K<n<100K0 likes38 downloads16d agoHugging Face15karthik1830 /FineTune_Kannadaaudio10K<n<100K0 likes37 downloads1y agoHugging Face16Sakshamrzt /IndicNLP-Kannadatexttext-classification10K<n<100K0 likes36 downloads2y agoHugging Face17lokesh122151 /kannada-tts-annotatedaudio1K<n<10K0 likes36 downloads5mo agoHugging Face18PharynxAI /IndicVoices-kannada-10000audio10K<n<100K0 likes35 downloads1y agoHugging Face19Cognitive-Lab /Kannada_Bilingual_Instructtext100K<n<1M1 likes34 downloads3y agoHugging Face20adithyal1998Bhat /stt_synthetic_kn-IN_kannadaaudio10K<n<100K0 likes34 downloads1y agoHugging Face21ramachandrajoshi /english-kannada-cleaned English–Kannada Cleaned A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models. Languages: English -> Kannada License: Apache License 2.0 Dataset statistics Train: 8,00,000 sentence pairs Validation: 1,000 sentence pairs Test: 1,000 sentence pairs Total: 5,02,000 sentence pairs These counts exclude per-file CSV headers. Source and provenance The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.text100K<n<1M1 likes34 downloads6mo agoHugging Face22Kannada-LLM-Labs /Wikipedia-Kn Dataset Card for "Wikipedia-Kn" This is a filtered version of the Wikipedia dataset only containing samples of Kannada language. The dataset contains total of 31437 samples. Data Sample: {'id': '832', 'url': 'https://kn.wikipedia.org/wiki/%E0%B2%A1%E0%B2%BF.%E0%B2%B5%E0%B2%BF.%E0%B2%97%E0%B3%81%E0%B2%82%E0%B2%A1%E0%B2%AA%E0%B3%8D%E0%B2%AA', 'title': 'ಡಿ.ವಿ.ಗುಂಡಪ್ಪ', 'text': 'ಡಿ ವಿ ಜಿ(ಮಾರ್ಚ್ ೧೭, ೧೮೮೭ - ಅಕ್ಟೋಬರ್ ೭, ೧೯೭೫) ಎಂಬ ಹೆಸರಿನಿಂದ ಪ್ರಸಿದ್ಧರಾದ ಡಾ. ದೇವನಹಳ್ಳಿ… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/Wikipedia-Kn.texttext-generation10K<n<100K1 likes32 downloads3y agoHugging Face23Anushhh /KannadaPromptBench KannadaPromptBench A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language. Dataset Summary Language: Kannada (kn) Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50) Total: 225 culturally grounded samples Inter-annotator agreement: Cohen's κ > 0.80 Dataset Structure Each sample contains: id, task, input_text, label, difficulty, domain. Citation Please… See the full description on the dataset page: https://huggingface.co/datasets/Anushhh/KannadaPromptBench.texttext-classificationn<1K1 likes32 downloads6mo agoHugging Face24charanhu /Kannada-Dataset-v02text100K<n<1M0 likes30 downloads3y agoHugging Face25saillab /alpaca-kannada-cleanedThis repository contains the dataset used for the TaCo paper. Please refer to the paper for more details: OpenReview If you have used our dataset, please cite it as follows: Citation @inproceedings{upadhayay2024taco, title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes}, author={Bibek Upadhayay and Vahid Behzadan}, booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR}, year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-kannada-cleaned.text10K<n<100K1 likes30 downloads2y agoHugging Face26Tensoic /GPTeacher-Kannadatext10K<n<100K0 likes29 downloads3y agoHugging Face27SPRINGLab /SPRING_INX_Kannada_R1audio10K<n<100K0 likes29 downloads2y agoHugging Face28cdactvm /kannada_new_data_v5audio1K<n<10K0 likes29 downloads2y agoHugging Face29Indic-Benchmark /kannada-arc-c-2.5ktext1K<n<10K0 likes28 downloads3y agoHugging Face30charanhu /kannada-instruct-dataset-390ktext100K<n<1M2 likes28 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.