datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
laions-got-talent_wordlevel-annotationvoice-annotation-data-v2
Voice Annotation Data v2
A curated dataset of 18,632 audio samples (9,391 positives + 9,241 negatives) across 58 voice dimensions. Each bucket contains up to 25 positive examples (audio that clearly fits the bucket) and 25 negative examples (audio confirmed to NOT fit the bucket by Gemini 2.0 Flash).
Changes from v1
Positive + Negative pairs: Every bucket now has up to 25 confirmed negative examples alongside 25 positives
EXPL redefined: Content Appropriateness reduced… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/voice-annotation-data-v2.voice-annotation-poc2voice-annotation-pocstarrail-voice-annotation
