datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC)
A corpus of read Neapolitan speech for ASR evaluation, with a validated
Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters,
metric implementations, per-clip results, and error annotations.
This release supersedes the earlier 141-clip single-speaker version of this
repository. The earlier release corresponds to Speaker S1 of the present
corpus; the old audioData/ and transcripts.csv are replaced by
data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.ViVoice34
ViVoice-34: Vietnamese Speech Dataset
Dataset Description
ViVoice-34 is a Vietnamese speech dataset featuring recordings from speakers across provinces of Vietnam. This repository contains a preview subset with playable WAV samples and metadata.
The preview data is stored as Parquet with the audio column encoded as Hugging Face Audio, so Dataset Viewer and Data Studio can render an audio player instead of plain file paths.
Data Fields
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-vivoice34/ViVoice34.
