datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Drivevoxtream2-test
Model Card for VoXtream2 test dataset
This repository contains a test dataset for the VoXtream2 TTS model, referred to as Emilia speaking-rate in the paper.
Audio prompts are derived from the Emilia dataset.
We selected 62 speakers (balanced by gender), half of whom include filler words in the prompt (uniformly distributed across genders).
Each speaker provides three audio prompts with slow (2.0 ± 0.5 SPS), normal (4.0 ± 0.6 SPS), and fast (5.6 ± 0.5 SPS) speaking rates.
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/herimor/voxtream2-test.mdaspcThis repo contains the dataset presented by Almeman et al., 2013. If you use this dataset, please cite the authors.
@INPROCEEDINGS{6487288,
author={Almeman, Khalid and Lee, Mark and Almiman, Ali Abdulrahman},
booktitle={2013 1st International Conference on Communications, Signal Processing, and their Applications (ICCSPA)},
title={Multi dialect Arabic speech parallel corpora},
year={2013},
volume={},
number={},
pages={1-6},
keywords={Speech;Training;Microphones;Cities and… See the full description on the dataset page: https://huggingface.co/datasets/herwoww/mdaspc.id-voicemail-dataset-v2
Dataset Card for "id-voicemail-dataset-v2"
More Information needed
eurospeech-bosnia-herzegovinalj_speech_with_spectogram
Explanation
A small experiment insipred by the Mistral playing DOOM experiment from the Mistral Hackathon
How it works?
Audio -> Waveform Visualization -> Waveform ASCII Art -> Finetune Mistral on ASCII Art to predict text from ASCII Art
Quick video explanation
Example Waveform
Example ASCII Art… See the full description on the dataset page: https://huggingface.co/datasets/0-hero/lj_speech_with_spectogram.Herbeetcelsoaudio_classification_kenya_dataset
Dataset Card for "audio_classification_kenya_dataset"
More Information needed
audio-samples-fixedid-earlymedia-dataset
Dataset Card for "id-earlymedia-dataset"
More Information needed
indonesia-earlymedia-digits-v1
Dataset Card for "indonesia-earlymedia-digits-v1"
More Information needed
hermioneadolescentHerbeetHenriqueexpresso-tagged-hermes-2-pro-llama3-8bafricanvoiceherbertluisaherculesherbetbetovozmgi1-Voice-Data-testname1-Voice-Data-metadata1-Voice-Data-demobosnia-herzegovinaindonesia_earlymedia_dataset
Dataset Card for "indonesia_earlymedia_dataset"
More Information needed
vozherikherbertvipHerbertNarativa1-Voice-Data1-Voice-Data-new-test-5hermio
