datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.spoken-alpaca-gpt04US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.Spoken2TSL
Dataset Description
This dataset is a collection of Turkish to Turkish Sign Language (TSL) grammar version translations. The dataset is designed to facilitate research and development in the field of sign language translation and understanding. It contains pairs of sentences in Turkish and their corresponding TSL translations, which have been curated to follow the grammatical structure of TSL.
Data Collection
The data was collected primarily from the website… See the full description on the dataset page: https://huggingface.co/datasets/ismaildlml/Spoken2TSL.va-spoken-qa-agentvoice-stage2
Stage-2 spoken-QA training corpus (gemma4_talker)
97,697 train / 935 val spoken QA pairs. Questions: REAL user audio from
VoiceAssistant-400K (audio_q.*.tar, filenames match input_audio basenames in the
manifests). Answers: synthesized in ONE fixed agent voice (LibriSpeech
train-clean-100 narrator ref via ResembleAI/Chatterbox; audio_a.*.tar matching
assistant_audio). Manifests carry transcripts, answer text, and GLM-4-Voice speech
tokens for both sides (glm_in_tokens question /… See the full description on the dataset page: https://huggingface.co/datasets/z050209/va-spoken-qa-agentvoice-stage2.spoken-alpaca-gpt4SpokenVisIT
