datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
KannadaPreTrainingIndicTTS_Kannada
Kannada Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Kannada
Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.syspin-kannada-ttsCulturaX-KnThis is a filtered version of the CulturaX dataset only containing samples of Kannada language.
The dataset contains total of 1352142 samples.
Dataset Structure:
{
"text": ...,
"timestamp": ...,
"url": ...,
"source": "mc4" | "OSCAR-xxxx",
}
Data Sample:
{'text': "ಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ ಸ್ಮರಣಾರ್ಥ ; ಉಚಿತ ನೋಟ್ ಬುಕ್ ವಿತರಣೆ | Vartha Bharati- ವಾರ್ತಾ ಭಾರತಿ\nಮುದರಂಗಡಿ ಬಿಜೆಪಿ ಗ್ರಾಪಂ ಸದಸ್ಯರ ವಿರುದ್ಧ ಪ್ರತಿಭಟನೆ\nಹೋಮ್ ಕ್ವಾರಂಟೈನ್ ನಿಯಮ ಉಲ್ಲಂಘನೆ: ಪ್ರಕರಣ ದಾಖಲು\nಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/CulturaX-Kn.OCR-Bench1000-Kannada
OCR-Bench1000-Kannada
1000 synthetic printed-text line images with ground-truth transcriptions,
sampled from a larger locally-held Kannada OCR training corpus.
This is a benchmark/sample release, not the full training set.
Data fields
Field
Description
file_name
relative path to the image (images/...)
text
ground-truth transcription
category
kannada_only / english_only / mixed / numeric_and_symbols
length_bucket
short / medium / long, by character… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Kannada.kannada-speech-datasetiisc-mile-kannada-asr-corpuskannada_new_dataoriginal_data_kannada_ttskannada_new_data_v2Kannada-Dataset-v01Aya_Kannada
Aya_Kannada
This Dataset is curated from the original Aya-Collection dataset that was open-sourced by Cohere under the Apache-2.0 license.
The Aya Collection is a massive multilingual collection comprising 513 million instances of prompts and completions that cover a wide range of tasks. This collection uses instruction-style templates from fluent speakers and applies them to a curated list of datasets. It also includes translations of instruction-style datasets into 101 languages.… See the full description on the dataset page: https://huggingface.co/datasets/Cognitive-Lab/Aya_Kannada.ICON26-COILD-INDIC-MT-Kannada-Malayalam
COILD-INDIC-MT 2026 — Kannada–Malayalam Dataset
This dataset is provided for the COILD-INDIC-MT 2026 Shared Task,
co-located with ICON 2026.
The shared task aims to foster research and innovation in
Natural Language Processing (NLP) for Indian Languages.
This repository contains data specifically for the:
Kannada ↔ Malayalam
language pair.
🔐 Access to the Dataset
This is a restricted and gated dataset.
Access is available only to authorized participants of the… See the full description on the dataset page: https://huggingface.co/datasets/ainlpml-iitp/ICON26-COILD-INDIC-MT-Kannada-Malayalam.FineTune_KannadaIndicNLP-Kannadakannada-tts-annotatedIndicVoices-kannada-10000Kannada_Bilingual_Instructstt_synthetic_kn-IN_kannadaenglish-kannada-cleaned
English–Kannada Cleaned
A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models.
Languages: English -> Kannada
License: Apache License 2.0
Dataset statistics
Train: 8,00,000 sentence pairs
Validation: 1,000 sentence pairs
Test: 1,000 sentence pairs
Total: 5,02,000 sentence pairs
These counts exclude per-file CSV headers.
Source and provenance
The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.Wikipedia-Kn
Dataset Card for "Wikipedia-Kn"
This is a filtered version of the Wikipedia dataset only containing samples of Kannada language.
The dataset contains total of 31437 samples.
Data Sample:
{'id': '832',
'url': 'https://kn.wikipedia.org/wiki/%E0%B2%A1%E0%B2%BF.%E0%B2%B5%E0%B2%BF.%E0%B2%97%E0%B3%81%E0%B2%82%E0%B2%A1%E0%B2%AA%E0%B3%8D%E0%B2%AA',
'title': 'ಡಿ.ವಿ.ಗುಂಡಪ್ಪ',
'text': 'ಡಿ ವಿ ಜಿ(ಮಾರ್ಚ್ ೧೭, ೧೮೮೭ - ಅಕ್ಟೋಬರ್ ೭, ೧೯೭೫) ಎಂಬ ಹೆಸರಿನಿಂದ ಪ್ರಸಿದ್ಧರಾದ ಡಾ. ದೇವನಹಳ್ಳಿ… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/Wikipedia-Kn.KannadaPromptBench
KannadaPromptBench
A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language.
Dataset Summary
Language: Kannada (kn)
Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50)
Total: 225 culturally grounded samples
Inter-annotator agreement: Cohen's κ > 0.80
Dataset Structure
Each sample contains: id, task, input_text, label, difficulty, domain.
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/Anushhh/KannadaPromptBench.Kannada-Dataset-v02alpaca-kannada-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-kannada-cleaned.GPTeacher-KannadaSPRING_INX_Kannada_R1kannada_new_data_v5kannada-arc-c-2.5kkannada-instruct-dataset-390k
