datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonvoice_17_tr_fixed
Improving CommonVoice 17 Turkish Dataset
I recently worked on enhancing the Mozilla CommonVoice 17 Turkish dataset to create a higher quality training set for speech recognition models.Here's an overview of my process and findings.
Initial Analysis and Split Organization
My first step was analyzing the dataset organization to understand its structure.Through analysis of filename stems as unique keys, I revealed and documented an important aspect of CommonVoice's design… See the full description on the dataset page: https://huggingface.co/datasets/ysdede/commonvoice_17_tr_fixed.tibetan-audio-to-english-fixed-filtered
Tibetan audio translation Dataset
Dataset Description
Tibetan audio translation Dataset
Dataset Summary
This dataset contains 6,366 audio samples with corresponding transcriptions, totaling approximately 15.8 hours of audio.
Languages
The dataset is in EN (Language code: en).
Dataset Structure
Data Fields
audio: An audio object containing:
path: Path to the audio file (if applicable)
array: Audio waveform as a numpy array… See the full description on the dataset page: https://huggingface.co/datasets/Titung/tibetan-audio-to-english-fixed-filtered.
