datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
english-kannada-cleaned
English–Kannada Cleaned
A cleaned parallel corpus of English–Kannada sentence pairs suitable for training and evaluating machine translation models.
Languages: English -> Kannada
License: Apache License 2.0
Dataset statistics
Train: 8,00,000 sentence pairs
Validation: 1,000 sentence pairs
Test: 1,000 sentence pairs
Total: 5,02,000 sentence pairs
These counts exclude per-file CSV headers.
Source and provenance
The dataset is provided as UTF-8 CSV files with… See the full description on the dataset page: https://huggingface.co/datasets/ramachandrajoshi/english-kannada-cleaned.KannadaPromptBench
KannadaPromptBench
A benchmark dataset for evaluating prompt strategy sensitivity in Kannada, a low-resource Dravidian language.
Dataset Summary
Language: Kannada (kn)
Tasks: Sentiment Analysis (100), Question Answering (75), Summarization (50)
Total: 225 culturally grounded samples
Inter-annotator agreement: Cohen's κ > 0.80
Dataset Structure
Each sample contains: id, task, input_text, label, difficulty, domain.
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/Anushhh/KannadaPromptBench.kannada-bedtime-tts
Kannada Bedtime Story TTS Dataset
Training data for fine-tuning IndicF5 on Kannada bedtime story narration.
Source
Base: SPRINGLab/IndicTTS_Kannada (800 clips, ~5h 53m)
Synthetic clips: 800 clips generated for bedtime story domain
Files
kannada_finetune/train.csv — Training metadata (text + audio paths)
kannada_finetune/val.csv — Validation split
kannada_manifest.jsonl — Full manifest with text, audio paths, durations
Usage
Used… See the full description on the dataset page: https://huggingface.co/datasets/sush0401/kannada-bedtime-tts.en-kannadakannada_news_classificationKannada-Speech-Dataset
🎧 Kannada Speech Dataset
The Kannada Speech Dataset is a high-quality speech audio dataset designed to deliver structured and reliable audio data for AI and machine learning workflows. It includes 90 hours of audio data across 651 files, available in MP3 and WAV formats, with a total size of 220 MB. This well-organized audio dataset provides balanced and representative voice data, with 48% female and 52% male speakers, and an age range spanning from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Kannada-Speech-Dataset.xnli2.0_train_kannadakannada-cultural-dialogue-datasetxnli2.0_kannadahh_dpo_kannada_translatedkannada-datakannada_qa_datasetmalayalam-kannada-tamil-telugu-samam-dataset
Samam.net Multilingual Dictionary Dataset
Dataset Description
This dataset contains multilingual dictionary entries scraped from samam.net, a comprehensive South Indian language dictionary. The dataset provides translations between Malayalam and three other Dravidian languages: Kannada, Tamil, and Telugu.
Important Note about Script Usage
All text in this dataset is written in Malayalam script, even for non-Malayalam languages. This is a key characteristic of… See the full description on the dataset page: https://huggingface.co/datasets/cazzz307/malayalam-kannada-tamil-telugu-samam-dataset.Quality_English_to_kannada_dataset
Quality English → Kannada Translation Dataset 🇮🇳
This dataset contains high‑quality synthetic bilingual pairs for English to Kannada translation.It was generated using Google Gemini with strict JSON formatting to ensure consistency and clean parsing.
📂 Contents
Total examples: ~3,000
File format: CSV and JSONL
Columns:
input: Prompt in English prefixed with "Translate to Kannada: …"
output: Correct Kannada translation
category: Semantic domain of the sentence… See the full description on the dataset page: https://huggingface.co/datasets/Aryanharitsa/Quality_English_to_kannada_dataset.
