datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
test3_suno_resrUrdu-TTS-v2-test3common-voice-test3k
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/DTU54DL/common-voice-test3k.ytb_test3test321
test321
This is a merged speech dataset containing 118 audio segments from 2 source datasets.
Dataset Information
Total Segments: 118
Speakers: 4
Languages: tr
Emotions: happy, angry, sad, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test321.test3
Dataset Card for "test3"
More Information needed
crowdsourced-test3voice-recorder-test3-demotest324234
test324234
This is a merged speech dataset containing 491 audio segments from 2 source datasets.
Dataset Information
Total Segments: 491
Speakers: 2
Languages: tr
Emotions: angry, neutral, happy
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test324234.test3434234
test3434234
This is a merged speech dataset containing 848 audio segments from 2 source datasets.
Dataset Information
Total Segments: 848
Speakers: 2
Languages: tr
Emotions: angry, happy, neutral
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, original sampling rate preserved)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion:… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test3434234.Thai-Voice-Test3
Thanarit/Thai-Voice
Combined Thai audio dataset from multiple sources
Dataset Details
Total samples: 0
Total duration: 0.00 hours
Language: Thai (th)
Audio format: 16kHz mono WAV
Volume normalization: -20dB
Sources
Processed 3 datasets in streaming mode
Source Datasets
GigaSpeech2: Large-scale multilingual speech corpus
ProcessedVoiceTH: Thai voice dataset with processed audio
MozillaCommonVoice: Mozilla Common Voice Thai dataset
Usage… See the full description on the dataset page: https://huggingface.co/datasets/Thanarit/Thai-Voice-Test3.dataset-farma-test3
Dataset Card for "dataset-farma-test3"
More Information needed
SAMLONEv3_20240419223152_Test3test3
test3
This is a merged speech dataset containing 345 audio segments from 2 source datasets.
Dataset Information
Total Segments: 345
Speakers: 7
Languages: en
Emotions: happy, neutral, angry, sad
Original Datasets: 2
Dataset Structure
Each example contains:
audio: Audio file (WAV format, 16kHz sampling rate)
text: Transcription of the audio
speaker_id: Unique speaker identifier (made unique across all merged datasets)
emotion: Detected emotion… See the full description on the dataset page: https://huggingface.co/datasets/Codyfederer/test3.Audio-Correction-Output-Test3whisper-finetune-audio_test3
