datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SPC
Dataset Card for "SPC-v2"
More Information needed
spc_r
Dataset Card: Swiss Parliaments Corpus — SPC_R v1.0
Background
The aim is to build a large, high-quality dataset. To get there, we correct pseudo-labeled transcriptions of parliamentary debates with an LLM. The model receives semantically relevant chunks from a manually prepared session protocol as context and then produces the corrected transcription.
We also show that Whisper’s average log probability can be used to predict BLEU. This lets us estimate transcription… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r.SPC
Dataset Card: Swiss Parliaments Corpus — Train v0.9
Summary
The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems.
If you intend to fine‑tune Whisper, we recommend the companion project i4Ds/whisper‑finetune, which is fully compatible with the data structure produced here.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC.spc_r_segmented
i4ds/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap between… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r_segmented.spc_r_whisperspc_r_segmented
eko57/spc_r_segmented
Diarized and segmented speech dataset derived from i4ds/spc_r.
Description
Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline:
Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap.
Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap… See the full description on the dataset page: https://huggingface.co/datasets/eko57/spc_r_segmented.SPC_test
Dataset Card: Swiss Parliaments Corpus — Test
Summary
The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems.
Dataset Details
Maintainer
Curated by: Vincenzo Timmel (@vincenzo.timmel)
Intended Use & Scope
Primary use‑case: Evaluation for Swiss-German STT Systems.… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC_test.spc_r_with_id
