datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
agent-sft-stitch-zh-tts
agent-sft-stitch-zh-tts
Voiced version of voidful/agent-sft-stitch-zh: the STITCH-S spoken chunks synthesized with BlueMagpie-TTS (hung_yi_lee voice), per-utterance loudness-aligned to -23 LUFS, best-of-N + Whisper-CER accepted.
Configs
records (default): one row per agent dialogue — id/source/user/msg (full STITCH-S trajectory) + available_tools + STITCH quality scores + spoken (ordered list of the utterances, each with audio, text, seg_index, cer, accepted… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts.narrativeqa-test-tts
Dataset Card for "narrativeqa-test-tts"
More Information needed
librispeech_unit_speech
Dataset Card for "librispeech_unit_speech"
More Information needed
VoidLinuxISOSgen_ai_2024
Dataset Card for "gen_ai_2024"
More Information needed
void-carousel
VOID CAROUSEL
Dark psychedelic trance from graveyard orbit.
Fictional entity
VOID CAROUSEL is a fictional sentient derelict orbital carousel in the Sonic Forage universe: an abandoned amusement machine circling a dead world, translating its failing motors, empty passenger rings and intercepted signals into imagined psychedelic trance. It is not a human performer; its visual identity depicts no human likeness. This is an original fictional characterization, not a… See the full description on the dataset page: https://huggingface.co/datasets/TheMindExpansionNetwork/void-carousel.spoken-alpaca-gpt4
Dataset Card for "speech-alpaca-gpt4"
More Information needed
hung-yi_lee
Hung-yi Lee Lecture ASR/TTS
This dataset contains speech segments aligned with subtitles from Hung-yi Lee lecture videos.
It is intended for Mandarin ASR and TTS experiments.
Dataset Details
Speaker: Hung-yi Lee
Language: Traditional Chinese / Taiwan Mandarin (zh-TW)
Sampling rate: 16 kHz
Audio format in the published dataset: embedded FLAC audio in Parquet shards
Source videos with audio: 15
Segments: 29043
Total duration: 16.92 hours
Courses: 機器學習2026… See the full description on the dataset page: https://huggingface.co/datasets/voidful/hung-yi_lee.agent-sft-stitch-zh-tts-samplecodec-superb-tinytkt-song
TKT Song
Public Taiwanese Hokkien song dataset in Hugging Face AudioFolder layout.
This dataset contains rows where all three resources are available and line-aligned:
song audio converted to 16 kHz mono FLAC
Han-character lyrics in lyrics.han
Hokkien romanization in lyrics.hokkien
Current build row count: 29384.
Schema
audio: Hugging Face Audio feature generated from file_name
source_id: source lyrics id
title: source song title
duration_seconds: YouTube… See the full description on the dataset page: https://huggingface.co/datasets/voidful/tkt-song.agent-sft-stitch-zh-tts-taste-codec-sample
agent-sft-stitch-zh-tts Taste-S codec sample
Ten accepted synthesized clips sampled from
voidful/agent-sft-stitch-zh-tts,
encoded with
andybi7676/taste-s-en-zhtw-small-gemma4.
Extraction follows IntelliGen's stage1_extract_taste.py: the 24 kHz source
audio is resampled to 16 kHz and converted to 80-bin cool-whisper features.
The encoder is conditioned on the external transcript tokenized with the Gemma
4 tokenizer, with streaming disabled.
codec_indices has shape… See the full description on the dataset page: https://huggingface.co/datasets/voidful/agent-sft-stitch-zh-tts-taste-codec-sample.cv_16_1_sample_100e2tts-test-suite
E2 TTS Test Suite
Summary
This dataset converts the evaluation manifest from microsoft/e2tts-test-suite into Hugging Face Datasets format.
Each example contains:
text_prompt
transcription_of_audio_prompt
ground_truth_audio
audio_prompt_audio
speaker metadata
is_subjective_eval
Source
Original repository:
https://github.com/microsoft/e2tts-test-suite
Notes
ground_truth_audio points to the reference utterance.
audio_prompt_audio points to the last 3… See the full description on the dataset page: https://huggingface.co/datasets/voidful/e2tts-test-suite.spoken-alpaca-gpt4-llm-codecdailytalk-conversations-grouped-llm-codecall_conv_data_filtered_smallearica_audio_test_2spoken-alpaca-gpt4-unitmmlm_testPhonologosynthetonPolyharmone
Dataset Card for "Polyharmone"
More Information needed
logomousia
Dataset Card for "logomousia"
More Information needed
Phonochronotoposearica_msearica_musicearica_audio
