datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FeruzaSpeech_to_fine_tuning
FeruzaSpeech_to_fine_tuning
A speech corpus of ⏱️ ~59.1 total hours of Uzbek audio paired with Latin‑script transcripts, intended for fine‑tuning ASR / speech‑to‑text models.
Dataset Details
Dataset Description
This dataset contains recordings of native Uzbek speakers reading a mix of classical literature excerpts and school‑level writing prompts:
001: Choliqushi (a novel by Rashod Nuri Guntekin, trans. by Mirzakalon Ismoiliy; first pub. Sept 1900).
002:… See the full description on the dataset page: https://huggingface.co/datasets/nickoo004/FeruzaSpeech_to_fine_tuning.vocal-burst-annotation-asr-tuning-dataset
Vocal Burst Annotation ASR Tuning Dataset
A synthetic 500,000-sample multilingual dataset for training ASR models with inline vocal burst captioning, speaker diarization, and sentence-level timestamps. Each sample is approximately 1 minute of audio containing speech segments interleaved with vocal bursts (laughs, sighs, coughs, etc.), annotated with precise timing information.
Example Transcript
[nasalized, affirmative hum, steady pitch, moderate intensity]… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-burst-annotation-asr-tuning-dataset.FYP_Fine_Tuning
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/shane062/FYP_Fine_Tuning.Speech_to_fine_tuning
FeruzaSpeech_to_fine_tuning
A speech corpus of ⏱️ ~59.1 total hours of Uzbek audio paired with Latin‑script transcripts, intended for fine‑tuning ASR / speech‑to‑text models.
Dataset Details
Dataset Description
This dataset contains recordings of native Uzbek speakers reading a mix of classical literature excerpts and school‑level writing prompts:
001: Choliqushi (a novel by Rashod Nuri Guntekin, trans. by Mirzakalon Ismoiliy; first pub. Sept 1900).
002:… See the full description on the dataset page: https://huggingface.co/datasets/hostbot77/Speech_to_fine_tuning.
