datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UZ_voiceuzbek_speech_dataSTT_uzThe dataset is organized into the following directories and files:
audio/
other/: Contains .tar archives like uz_other_0.taruz_other_1.tar
train/: Contains .tar archives like uz_train_0.tar.
validated/: Contains .tar archives like uz_validated_0.tar, uz_validated_1.tar, and uz_validated_2.tar.
test/: Contains individual .wav files.
transcription/: Contains .tsv files including:
other.tsv
train.tsv
validated.tsv
test.tsv
The .tsv files have two columns: file_name and transcription. Each entry… See the full description on the dataset page: https://huggingface.co/datasets/Beehzod/STT_uz.uzbek_stt_datanew_databot_sttuz-data
