datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lunde_nor_nob_reading_optimisedTest only - not for training.
First version - 0.1 of lunde_nor_nob_reading_optimised
This dataset does not contain any audio data.
Export Details
Train samples: 10040932
Validation samples: 0
Test samples: 0
Dataset created using search datasets:lunde_nor_nob_reading_optimised.
gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814
This dataset combines six independently aligned source archives. Each row embeds its selected MP3 in the audio Parquet column. SBPN-derived word timestamps are observational and do not control chunk edges or the Demucs vote. Accepted hard-word verbalizations are projected back to the original written forms; pronunciation_alignment_dictionary_json records the winning spoken form. Non-music tags are preserved… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-sbpn-fresh-diarization-demucs-optimized-terminal-l4-20260814.
