datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test4
test4
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: neutral, sad, angry, happy
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral, happy… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test4.test5
test5
This is a merged speech dataset containing 1806 audio segments from 8 source datasets.
Dataset Information
Total Segments: 1806
Speakers: 47
Languages: en
Emotions: happy, neutral, sad, angry
Original Datasets: 8
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion (neutral… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test5.octonet_testru-vc-testDataset to test Russian zero-short voice conversion
Based on test subset of yt-vad-650-clean from https://github.com/GeorgeFedoseev/DeepSpeech
test6
test6
This is a merged speech dataset containing 1994 audio segments from 2 source datasets.
Dataset Information
Total Segments: 1994
Speakers: 3
Languages: en
Emotions: neutral, negative_surprise, positive_surprise, distress, relief, contentment, adoration, interest, confusion, happy, sadness, triumph, fear, disappointment, awe, realization, angry
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test6.TestTest Dataset
testtr12
testtr12
This is a merged speech dataset containing 165 audio segments from 2 source datasets.
Dataset Information
Total Segments: 165
Speakers: 8
Languages: en
Emotions: happy, angry, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/testtr12.indicf5-nfe-9-10-11-stress-testtest3
test3
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: happy, neutral, angry, sad
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test3.testtesttestdata
