datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
p1-segments
DR P1 speech segments
Dataset
Danish speech clips from DR P1, in mono 16 kHz OGG/Opus, with verbatim text, timing, and speaker metadata. Transcript text and speaker attribution may contain automated errors.
Source
The recordings cover roughly 2006–2022 and come from DR P1 recordings in kb.dk’s DR archive. Audio is sourced through the pinned syvai/p1 revision 449b9c2294026df6d0d37538f279fdec03f565ff. Transcripts were generated with ElevenLabs… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1-segments.p1
DR P1 Audio Archive
Danish public radio (DR) P1 audio recordings sourced from the kb.dk DR-arkivet (Royal Danish Library DR archive), covering roughly 2006–2022.
Format
Audio: Opus, 24 kbps, mono, in OGG container (transcoded from DR's mp3 archive)
Parquet shards (~500 items each), small row groups for streaming compatibility
Sortable by year / month / start_time
Schema
Each row is one broadcast item with the full audio bytes inline plus rich… See the full description on the dataset page: https://huggingface.co/datasets/syvai/p1.
