CoolFace
Datasetpublic

uam-wmi-asr-eval-labs/2025-zwesui-g03-debata-prezydencka

ZWESUI 2025 - Grupa 3 - debata prezydencka 2025 Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny. Zespol (atrybucja): Grupa 3 (2025) Zrodlo oryginalne: https://huggingface.co/datasets/directtt/polish_presidential_debate Domena: debata prezydencka Opis: Debata prezydencka TVP z 12 maja 2025 — 13 kandydatów, 195 wypowiedzi, mowa spontaniczna.… See the full description on the dataset page: https://huggingface.co/datasets/uam-wmi-asr-eval-labs/2025-zwesui-g03-debata-prezydencka.

sourceHugging Facecc-by-nc-4.0updated 1mo agoView on Hugging Face
0likes23downloads
Dataset Card

ZWESUI 2025 - Grupa 3 - debata prezydencka 2025

Robocza/archiwalna kopia zbioru ewaluacyjnego ASR zbudowanego przez studentow kursu Warsztaty z ewaluacji systemow rozpoznawania mowy (UAM WMI), edycja 2025 (pierwsza), tryb niestacjonarny.

  • —Zespol (atrybucja): Grupa 3 (2025)
  • —Zrodlo oryginalne: https://huggingface.co/datasets/directtt/polishpresidentialdebate
  • —Domena: debata prezydencka
  • —Opis: Debata prezydencka TVP z 12 maja 2025 — 13 kandydatów, 195 wypowiedzi, mowa spontaniczna. Kopia repackagingu BIGOS-format z amu-cai/polishpresidentialdebate-2025 (statyczny Parquet, schemat BIGOS V2).
  • —Licencja zrodla: upstream directtt/polishpresidentialdebate Apache 2.0; repackaging BIGOS-format (amu-cai/polishpresidentialdebate-2025) CC-BY-NC-4.0
  • —Status: kopia publiczna w organizacji kursowej; oryginał publiczny. Planowane włączenie do BIGOS V3 i Polish ASR Leaderboard V2 (do końca 2026).

Kopia utworzona 2026-08-15 (wlaczenie edycji 2025 do organizacji). Konwencja nazwy: {rok}-{tryb}-g{NN}-{domena}.

Oryginalny opis (kopia w formacie BIGOS, amu-cai)

Polish Presidential Debate 2025 (BIGOS-format)

BIGOS-V2-compatible repackaging of directtt/polish_presidential_debate (Apache 2.0, Dudek et al. 2025) for use with the BIGOS V3 platform ASR evaluation pipeline. 195 samples (13 candidates × ~15 utterances each) from the TVP Polish Presidential Debate on 12 May 2025.

Source

  • —Upstream corpus: directtt/polish_presidential_debate (Apache 2.0)
  • —Pinned revision: ae67d748a50d
  • —Audio source: DEBATA PREZYDENCKA TVP | 12.05.2025
  • —Authors: Dudek Marcel, Jerzykiewicz Sebastian, Rybczyński Jędrzej, Solarski Antoni (2025)
  • —Repackaging: Michał Junczyk, AMU CAI (2026-05). BIGOS V2 schema mapping + Parquet conversion, no audio modification.

Why a separate repo

The upstream directtt/polish_presidential_debate ships as a script-based HF dataset (loader at polish_presidential_debate.py), which requires trust_remote_code=True to load. The BIGOS platform's ADR-040 policy blocks that flag for safety. This repository serves identical audio + transcripts as a static Parquet dataset, removing the policy block and enabling reproducible loading via plain load_dataset().

Sample distribution

13 candidates × 15 utterances each = 195 samples. Each candidate is a separate dataset subset value, enabling natural stratified sampling in the BIGOS scenario loader.

Schema (BIGOS V2 compatible)

ColumnTypeDescription
audionamestringUnique sample identifier (e.g., bartoszewicz_artur_0)
audioAudio(16kHz)WAV audio bytes
ref_origstringGold reference transcription
datasetstringSubset name = candidate (e.g., bartoszewicz_artur)
speaker_idstringSame as dataset (each candidate is one speaker)
speaker_genderstringM / F
speaker_agestringSpeaker age (integer-as-string)
speech_typestringread / spontaneous
sourcestringyoutube (TVP broadcast originally posted on YouTube)
localestringalways pl

License

  • —Original audio + transcripts: Apache 2.0 (per upstream directtt/polish_presidential_debate)
  • —BIGOS schema mapping (this repo's curation): CC-BY-NC-4.0

When using this dataset, please cite the original authors:

bibtex
@misc{polish_presidential_debate,
  title = {Polish Presidential Debate ASR Dataset},
  author = {Dudek Marcel, Jerzykiewicz Sebastian, Rybczyński Jędrzej, Solarski Antoni},
  year = {2025}
}

Usage

python
from datasets import load_dataset

ds = load_dataset("amu-cai/polish_presidential_debate-2025", split="test")
print(ds[0]["audioname"], ds[0]["ref_orig"][:80])

BIGOS evaluation pipeline

bash
./scripts/benchmark/run_long_eval.sh --scenario debate-test --sync-cache