CoolFace
Datasetpublic

baki83/JuzneVesti-SR-Unsloth-Format

Emilia-compatible Serbian speech (JuzneVesti-SR) This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio training pipelines. It exposes the same columns as kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split (with dev named validation). Source Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022). Persistent identifier:… See the full description on the dataset page: https://huggingface.co/datasets/baki83/JuzneVesti-SR-Unsloth-Format.

sourceHugging Facecc-by-sa-4.0updated 4d agoView on Hugging Face
0likes59downloads
Dataset Card

Emilia-compatible Serbian speech (JuzneVesti-SR)

This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio training pipelines. It exposes the same columns as kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split (with dev named validation).

Source

Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022).

  • —Persistent identifier: <http://hdl.handle.net/11356/1679>
  • —Source size: 10,811 entries, 50.55 hours
  • —Source license: CC BY-SA 4.0
  • —Content: Serbian interviews from the Južne vesti program 15 minuta
  • —Segment duration in the source: 2-30 seconds

The default conversion retains 3-30 second clips so its duration range matches the referenced Emilia subset. This leaves 10,806 clips: 8,644 train, 1,081 validation, and 1,081 test.

Columns

  • —audio: embedded 16 kHz mono WAV
  • —text: raw transcript with punctuation and case (default conversion)
  • —duration: seconds
  • —speaker: host/guest name when a segment contains one speaker role; multiple_speakers otherwise
  • —language: sr
  • —dnsmos: null because the source does not publish DNSMOS values
  • —phone_count: deterministic Serbian orthographic phoneme estimate
  • —_id: stable ID derived from the source audio path

Limitations

The source alignments were created through an automated alignment pipeline and then filtered by transcript/alignment agreement. Some transcription and speaker boundary errors can remain. The domain is interviews/news and does not cover all Serbian regions, speaking styles, ages, or recording conditions. Although the publisher is based in Nis and this is regionally useful material, speakers are not individually labelled by accent; it must not be treated as a verified southern-accent-only corpus. phone_count is an estimate, not a phonetic forced alignment. DNSMOS was not computed.

License

This derivative dataset is distributed under CC BY-SA 4.0. Users must retain attribution to the original authors and source.