Text to speech
text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.700h-tr-turkish-text-to-speechSpeech-To-Text-System-Prompts-2
Speech To Text System Prompt Library
This repository provides a collection of system prompts designed to transform and refine text captured using speech-to-text technologies.
By passing STT outputs through large language models with these specialized prompts, you can achieve cleaner, more structured, and purpose-specific text formats.
📋 The Idea
Here is the basic implementation. I don't pretend that this is the stuff of high AI engineering. But it does create quite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Speech-To-Text-System-Prompts-2.Dataset-Text-To-Speech-Indonesia
🎵 Dataset Audio Bahasa Indonesia
Dataset audio berkualitas tinggi untuk Text-to-Speech (TTS) bahasa Indonesia.
Dibuat oleh : Muhammad Arief, S.Kom.Universitas Muhammadiyah SorongTeknik Informatika 2020
📊 Spesifikasi Teknis
Parameter
Nilai
Satuan
Total Durasi
16.38
jam
Jumlah Segmen
4531
file
Durasi Rata-rata
13.01
detik
Sample Rate KHz
22
kHz
Sample Rate Hz
22000
Hz
Bit Depth
PCM_16
PCM
Format
wav
Lossless
🔄 Urutan Pengolahan… See the full description on the dataset page: https://huggingface.co/datasets/X-lord/Dataset-Text-To-Speech-Indonesia.darija_speech_to_textnepali_speech_to_text
Nepali Speech-to-Text Dataset
This repository contains a dataset for Automatic Speech Recognition (ASR) in the Nepali language. The dataset is designed for supervised learning tasks and includes audio files along with their corresponding transcriptions. The audio samples have been collected from various open-source platforms and other publicly available sources on the internet.
Each audio file has an average length of 15 seconds and has been converted into a consistent WAV format… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/nepali_speech_to_text.
