datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.DA-FNOThis model was trained using the following environment:
Python 3.9
PyTorch 2.2.2
numpy 1.24
CUDA 12.6
comparison: Comparison of the DA-FNO with different theta and Conv-FNO and vanilla FNO models.
final_model: DA-FNO model trained on the entire training set.
dataset: The first 200 samples in the training set.
spectrum_2700: Average spectrum of the training set.
fno-datafNOAhDV0CAIpAmMX
