michaelcacioli/Neapolitan-Spoken-Corpus
Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/michaelcacioli/Neapolitan-Spoken-Corpus.
Remove superseded files: old transcripts.csv, requirments.txt, code/, and 141-clip audioData/ (replaced by data/audio and data/metadata.csv)
Full release: expanded NSC (591 clips), lexicon, LOSO splits, LoRA adapters, metrics, annotations, rebuttal analyses
Expansion
Delete audioData/test
Upload 141 files
Create audioData/test
Create transcribe_whisper.py
fix format mistake
Create code/generate_json.py
Create code/evaluate_metrics.py
Create transcripts.csv
Create requirments.txt
Update README.md
initial commit
