datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TTS_ODIASPRING_INX_Odia_R1IndicTTS_Odia
Odia Indic TTS Dataset
This dataset is derived from the Indic TTS Database project, specifically using the Odia monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development.
Dataset Details
Language: Odia
Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours)
Audio Format: WAV
Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Odia.original_data_odia_ttsodia_data_v1odia-english-ASROdia-data-collectionThis is a repo created for training purpose for the language "Odia".
dataset's collection, that exists:
Link
odia_data_v2odia_data_v5Odia_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 12,794 hours of processed Odia(Oriya) (OR) dual-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments. The dataset is organized in a dual-channel format… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Odia_Call_Center_Audio_Dataset_Dual_Channel.Odia-Call-Center-Audio-Dataset-Single-ChannelDataset Description:
This dataset is a large-scale collection of 12,794 hours of processed Odia (OR) single-channel call center audio recordings, containing 3,569,083 hours of processed call center audio recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
The dataset captures authentic speech characteristics such as tone variation, pauses, silence patterns, and natural speaking behaviour commonly observed in… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/Odia-Call-Center-Audio-Dataset-Single-Channel.odia-tts-descriptionsVaani-odia-majority-lg-English-no-transcript0TTS_ODIA_500odia-tts-male-audio-44kVaani-odia-majority-lg-English-with-transcriptodia_data_v1odia_asr_collection
Odia ASR Collection
Aggregated Odia audio datasets (supervised and unsupervised) for ASR research.
Contents
oden — source: BBSRguy/ODEN-speech, supervised: None, audio_format: None, rows: 0, status: error, error: An error occurred while generating the dataset
common_voice_odia — source: mozilla-foundation/common_voice_17_0, supervised: None, audio_format: None, rows: 0, status: error, error: Dataset scripts are no longer supported, but found common_voice_17_0.py… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/odia_asr_collection.odia_data_v2odia-2speaker-ttsodia-tts-elevenlabsodia_speech_dataodia_data_v4odia_data_v3
