kannada
Datasets
All datasets matching “kannada”unified-kannada-asr-1.0
Dataset Card for "unified-kannada-asr-1.0"
More Information needed
KannadaPreTrainingIndicTTS_Kannada
Kannada Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Kannada monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Kannada
Total Duration: ~7.35 hours (Male: 3.4 hours, Female: 3.95 hours)
Audio Format: WAV
Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Kannada.kannada_newsThe Kannada news dataset contains only the headlines of news article in three categories:
Entertainment, Tech, and Sports.
The data set contains around 6300 news article headlines which collected from Kannada news websites.
The data set has been cleaned and contains train and test set using which can be used to benchmark
classification models in Kannada.syspin-kannada-ttsCulturaX-KnThis is a filtered version of the CulturaX dataset only containing samples of Kannada language.
The dataset contains total of 1352142 samples.
Dataset Structure:
{
"text": ...,
"timestamp": ...,
"url": ...,
"source": "mc4" | "OSCAR-xxxx",
}
Data Sample:
{'text': "ಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ ಸ್ಮರಣಾರ್ಥ ; ಉಚಿತ ನೋಟ್ ಬುಕ್ ವಿತರಣೆ | Vartha Bharati- ವಾರ್ತಾ ಭಾರತಿ\nಮುದರಂಗಡಿ ಬಿಜೆಪಿ ಗ್ರಾಪಂ ಸದಸ್ಯರ ವಿರುದ್ಧ ಪ್ರತಿಭಟನೆ\nಹೋಮ್ ಕ್ವಾರಂಟೈನ್ ನಿಯಮ ಉಲ್ಲಂಘನೆ: ಪ್ರಕರಣ ದಾಖಲು\nಭಟ್ಕಳ : ತಂದೆ ತಾಯಿ… See the full description on the dataset page: https://huggingface.co/datasets/Kannada-LLM-Labs/CulturaX-Kn.
