synthetic speech
synthetic-speech-indicmultivoice-synthetic-speech
Synthetic Voice Samples · Africa
Synthetic speech. No human speaker was recorded for any clip in this dataset.
Generated with afrispeech-synth: text from
africa-corpus, normalised to a
universal orthography with africa-g2p, spoken by
Google Gemini's Live API.
17,010 clips · 38.8 hours · 566 languages · 30 voices
Every clip is a distinct sentence — no sentence is repeated
Each language is read by up to 30 different voices, one sentence per voice
~1.29 hours per voice… See the full description on the dataset page: https://huggingface.co/datasets/AfriSpeech/multivoice-synthetic-speech.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.romanian_speech_dataset_with_15_percent_6_speakers_synthetic_datasynthetic_speech_commands_PA_taggedromanian_speech_dataset_with_20_percent_4_speakers_synthetic_data
