datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.700h-tr-turkish-text-to-speechSpeech-To-Text-System-Prompts-2
Speech To Text System Prompt Library
This repository provides a collection of system prompts designed to transform and refine text captured using speech-to-text technologies.
By passing STT outputs through large language models with these specialized prompts, you can achieve cleaner, more structured, and purpose-specific text formats.
📋 The Idea
Here is the basic implementation. I don't pretend that this is the stuff of high AI engineering. But it does create quite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Speech-To-Text-System-Prompts-2.Dataset-Text-To-Speech-Indonesia
🎵 Dataset Audio Bahasa Indonesia
Dataset audio berkualitas tinggi untuk Text-to-Speech (TTS) bahasa Indonesia.
Dibuat oleh : Muhammad Arief, S.Kom.Universitas Muhammadiyah SorongTeknik Informatika 2020
📊 Spesifikasi Teknis
Parameter
Nilai
Satuan
Total Durasi
16.38
jam
Jumlah Segmen
4531
file
Durasi Rata-rata
13.01
detik
Sample Rate KHz
22
kHz
Sample Rate Hz
22000
Hz
Bit Depth
PCM_16
PCM
Format
wav
Lossless
🔄 Urutan Pengolahan… See the full description on the dataset page: https://huggingface.co/datasets/X-lord/Dataset-Text-To-Speech-Indonesia.darija_speech_to_textnepali_speech_to_text
Nepali Speech-to-Text Dataset
This repository contains a dataset for Automatic Speech Recognition (ASR) in the Nepali language. The dataset is designed for supervised learning tasks and includes audio files along with their corresponding transcriptions. The audio samples have been collected from various open-source platforms and other publicly available sources on the internet.
Each audio file has an average length of 15 seconds and has been converted into a consistent WAV format… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/nepali_speech_to_text.speech-to-textTamazight-Speech-to-Arabic-Text
Tamazight-Arabic Speech Recognition Dataset
This is the Tamazight-NLP organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset. This dataset contains ~15.5 hours of Tamazight (Tachelhit dialect) speech paired with Arabic transcriptions, designed for automatic speech recognition (ASR) and speech-to-text translation tasks.
Dataset Details
Total Examples: 20,344 audio segments
Training Set: 18,309 examples (~8.9GB)
Test Set: 2,035 examples (~992MB)… See the full description on the dataset page: https://huggingface.co/datasets/Tamazight-NLP/Tamazight-Speech-to-Arabic-Text.text-to-speech-en-IN-checkpointArabic-Text-to-Speechvietnamese-speech-to-text-preprocessed-whisper-large-v3speech_to_text_yixing_dialectvietnamese-speech-to-text-preprocessed-whisper-mediumdarija-speech-to-text
Speech To Text Darija dataset
Reupload of adiren7/darija_speech_to_text
Tamazight-Speech-to-Arabic-Text
Tamazight-Arabic Speech Recognition Dataset
Overview
This is the EMINES organization-hosted version of the Tamazight-Arabic Speech Recognition Dataset, synchronized with the original dataset. It contains ~15.5 hours of Tamazight speech (Tachelhit dialect) paired with Arabic transcriptions, designed for developing ASR and translation systems.
Quick Start
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/EMINES/Tamazight-Speech-to-Arabic-Text.arabic_speech_to_text_20241219_205753_x4mhwqarabic_speech_to_text_20241219_203218_lp9vcvarabic_speech_to_text_20241218_144737_gkopimAfar-language-text-to-speech-TTS
Usage
This dataset is designed to support the development of Text-to-Speech (TTS) systems for the Afar language. It can be integrated into web applications, mobile apps, desktop software, or other platforms that require natural-sounding Afar voice synthesis or accurate spoken language recognition.
For applications involving virtual avatars or voice personas, the following culturally appropriate voice names are recommended:
Female Voices: Emeli, Hanaawi, Kareera, Laysani, Kulsuma… See the full description on the dataset page: https://huggingface.co/datasets/Charif-Ayfarah/Afar-language-text-to-speech-TTS.darija_to_french_speech_to_textarabic_speech_to_text_20241218_174023_sd5hyynepali-speech-to-textHere's a README draft for your Hugging Face dataset:
Nepali Speech-to-Text Dataset
This dataset contains high-quality speech samples in Nepali, originally from OpenSLR SLR43 and Mozilla's Common Voice dataset. It has been cleaned and processed for Automatic Speech Recognition (ASR) tasks. The dataset consists of approximately 3,000 audio samples, each around 30 seconds long, compiled for use in training and testing ASR models.
Dataset Details
Number of samples:… See the full description on the dataset page: https://huggingface.co/datasets/amitpant7/nepali-speech-to-text.TextToSpeechMedConvonepali-speech-to-textarabic_speech_to_text_20241224_135643_72lw7rspeech_to_textarabic_speech_to_text_20241224_134331_eiicpwarabic_speech_to_text_20241218_173148_7rkwsyarabic_speech_to_text_20241222_184338_sngqcgspeech-to-text
