CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Spaiche /SPC Dataset Card for "SPC-v2" More Information needed audio10K<n<100K0 likes659 downloads4y agoHugging Face02i4ds /spc_r Dataset Card: Swiss Parliaments Corpus — SPC_R v1.0 Background The aim is to build a large, high-quality dataset. To get there, we correct pseudo-labeled transcriptions of parliamentary debates with an LLM. The model receives semantically relevant chunks from a manually prepared session protocol as context and then produces the corrected transcription. We also show that Whisper’s average log probability can be used to predict BLEU. This lets us estimate transcription… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r.audioautomatic-speech-recognition10K<n<100K9 likes228 downloads7mo agoHugging Face03i4ds /SPC Dataset Card: Swiss Parliaments Corpus — Train v0.9 Summary The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems. If you intend to fine‑tune Whisper, we recommend the companion project i4Ds/whisper‑finetune, which is fully compatible with the data structure produced here. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC.audioautomatic-speech-recognition10K<n<100K0 likes114 downloads1y agoHugging Face04i4ds /spc_r_segmented i4ds/spc_r_segmented Diarized and segmented speech dataset derived from i4ds/spc_r. Description Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline: Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap. Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap between… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/spc_r_segmented.audioautomatic-speech-recognition100K<n<1M1 likes77 downloads7mo agoHugging Face05i4ds /spc_r_whisperaudio10K<n<100K2 likes52 downloads1y agoHugging Face06eko57 /spc_r_segmented eko57/spc_r_segmented Diarized and segmented speech dataset derived from i4ds/spc_r. Description Each row is a merged speech segment belonging to a single speaker. The source audio and SRT subtitles from i4ds/spc_r were processed with the following pipeline: Diarization -- pyannote/speaker-diarization-3.1 assigned speaker labels to each SRT segment based on temporal overlap. Merging -- Consecutive SRT segments from the same speaker were merged when the silence gap… See the full description on the dataset page: https://huggingface.co/datasets/eko57/spc_r_segmented.audioautomatic-speech-recognition100K<n<1M0 likes50 downloads7mo agoHugging Face07i4ds /SPC_test Dataset Card: Swiss Parliaments Corpus — Test Summary The SPC Train v0.9 release pairs Swiss German speech with Standard German transcriptions, providing a high‑quality resource for training and evaluating automatic speech‑recognition (ASR) or speech‑translation systems. Dataset Details Maintainer Curated by: Vincenzo Timmel (@vincenzo.timmel) Intended Use & Scope Primary use‑case: Evaluation for Swiss-German STT Systems.… See the full description on the dataset page: https://huggingface.co/datasets/i4ds/SPC_test.audio1K<n<10K0 likes36 downloads1y agoHugging Face08talha-waqar17 /spc_r_with_idaudio10K<n<100K0 likes5 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.