baki83/JuzneVesti-SR-Unsloth-Format
Emilia-compatible Serbian speech (JuzneVesti-SR) This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio training pipelines. It exposes the same columns as kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split (with dev named validation). Source Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022). Persistent identifier:… See the full description on the dataset page: https://huggingface.co/datasets/baki83/JuzneVesti-SR-Unsloth-Format.
Emilia-compatible Serbian speech (JuzneVesti-SR)
This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio training pipelines. It exposes the same columns as kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split (with dev named validation).
Source
Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022).
- Persistent identifier: <http://hdl.handle.net/11356/1679>
- Source size: 10,811 entries, 50.55 hours
- Source license: CC BY-SA 4.0
- Content: Serbian interviews from the Južne vesti program
15 minuta - Segment duration in the source: 2-30 seconds
The default conversion retains 3-30 second clips so its duration range matches the referenced Emilia subset. This leaves 10,806 clips: 8,644 train, 1,081 validation, and 1,081 test.
Columns
audio: embedded 16 kHz mono WAVtext: raw transcript with punctuation and case (default conversion)duration: secondsspeaker: host/guest name when a segment contains one speaker role;multiple_speakersotherwiselanguage:srdnsmos: null because the source does not publish DNSMOS valuesphone_count: deterministic Serbian orthographic phoneme estimate_id: stable ID derived from the source audio path
Limitations
The source alignments were created through an automated alignment pipeline and then filtered by transcript/alignment agreement. Some transcription and speaker boundary errors can remain. The domain is interviews/news and does not cover all Serbian regions, speaking styles, ages, or recording conditions. Although the publisher is based in Nis and this is regionally useful material, speakers are not individually labelled by accent; it must not be treated as a verified southern-accent-only corpus. phone_count is an estimate, not a phonetic forced alignment. DNSMOS was not computed.
License
This derivative dataset is distributed under CC BY-SA 4.0. Users must retain attribution to the original authors and source.
