datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-speech-human-preferences-315k
Text-to-speech human preferences: 315K votes across 15 models
This gated dataset contains the evaluation record behind Datapoint Audio
Bench: 315,000 eligible pairwise votes comparing 15 text-to-speech
models in a complete round-robin over 300 English prompts. The prompt set
covers eight practical voice-agent categories, and every generated sample is
included as a typed audio record.
The source evaluation collected 357,651 completed responses. The published
benchmark excluded… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-to-speech-human-preferences-315k.Data_voice_AI_human_scam
Vietnamese Deepfake Voice Dataset
Dataset Description
The Vietnamese Deepfake Voice Dataset is a multimodal dataset designed for research on deepfake voice detection and scam call detection in Vietnamese. The dataset contains both authentic human speech and AI-generated speech collected from multiple speech synthesis and voice cloning systems.
The dataset is intended for developing and evaluating machine learning and deep learning models for:
Audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/vietkemmai/Data_voice_AI_human_scam.tts-human-preferences-large
TTS Human Preferences (Large)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the large (2,700-row) subset. See also: small (1,000 rows), medium (2,000 rows).
Dataset Summary
Metric
Value
Total rows
2,700
Annotations per row
15
Total annotations
40,500
Unique… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-large.tts-human-preferences-small
TTS Human Preferences (Small)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the small (1,000-row) subset. Larger versions will follow.
Dataset Summary
Metric
Value
Total rows
1,000
Annotations per row
15
Total annotations
15,000
Unique prompts
1,000
Audio format… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-small.tts-human-preferences-medium
TTS Human Preferences (Medium)
Human preference dataset for text-to-speech (TTS) audio quality evaluation. Each row contains two TTS audio renderings of the same text prompt, along with 15 human preference annotations indicating which audio sounds more natural.
This is the medium (2,000-row) subset. See also: small (1,000 rows). Larger versions will follow.
Dataset Summary
Metric
Value
Total rows
2,000
Annotations per row
15
Total annotations
30,000… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/tts-human-preferences-medium.Noise-from-HumanThis dataset includes noise from Human, such as
BREATH
MUNCHING
