datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ravnursson_asr
Dataset Card for ravnursson_asr
Dataset Summary
The corpus "RAVNURSSON FAROESE SPEECH AND TRANSCRIPTS" (or RAVNURSSON Corpus for short) is a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications in the language that is spoken at the Faroe Islands (Faroese). It was curated at the Reykjavík University (RU) in 2022.
The RAVNURSSON Corpus is an extract of the "Basic Language Resource Kit 1.0" (BLARK 1.0) [1] developed… See the full description on the dataset page: https://huggingface.co/datasets/carlosdanielhernandezmena/ravnursson_asr.voice-of-care-health-dataset
Voice of Care AI for Global Health Benchmark Dataset
Overview
This dataset contains spoken Hausa Health datasets with rich annotations covering emotion, intent, speaker demographics, and dialect variation, intended for speech and NLP research.
Dataset Summary
Property
Details
Language
Hausa
Modality
Audio + Text
Task(s)
e.g. Speech Recognition, Emotion Detection, Dialect Identification
Version
1.0.0
🛠️ Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Data-Science-Nigeria/voice-of-care-health-dataset.bengali-telecom-customer-care-speech-v2
Bengali Telecom Customer Care Synthetic Speech Dataset v2
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
This is a second version of the Bengali Telecom Customer Care Synthetic Speech Dataset. It follows the same… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech-v2.toy_corpus_asr_caThis is an example of a repository with a standard data loader. The audio files are compressed in tar format. Since this repository contains very few audio files, it can be used to test certain scripts in local machines.
bengali-telecom-customer-care-speech
Bengali Telecom Customer Care Synthetic Speech Dataset
Dataset Description
This dataset contains synthetic Bengali speech generated from telecom and customer-care style text prompts.
The dataset is intended for experiments with:
Bengali ASR/STT
Bengali TTS
Speech-to-text preprocessing
Telecom/customer-care domain adaptation
Synthetic speech research
Important Disclosure
This is a synthetic speech dataset generated using the OmniVoice TTS system in… See the full description on the dataset page: https://huggingface.co/datasets/kawshikbuet17/bengali-telecom-customer-care-speech.toy_corpus_asr_esThis is an example of a repository with a standard data loader. The audio files are compressed in tar format. Since this repository contains very few audio files, it can be used to test certain scripts in local machines.
prueba_parquetThis is an example of a repository with parquet files only.
