united-nations/transcription-results
UN Transcription Benchmark Results Evaluation results for speech-to-text systems on UN Security Council and General Assembly meeting recordings, assessed against official UN verbatim records. See united-nations/transcription-corpus for the underlying audio and ground truth data. Metrics WER: Word Error Rate (reference = verbatim record, no normalization) normalized_wer: WER after lowercasing, punctuation removal, and filler word removal CER: Character Error Rate… See the full description on the dataset page: https://huggingface.co/datasets/united-nations/transcription-results.
UN Transcription Benchmark Results
Evaluation results for speech-to-text systems on UN Security Council and General Assembly meeting recordings, assessed against official UN verbatim records.
See united-nations/transcription-corpus for the underlying audio and ground truth data.
Metrics
- WER: Word Error Rate (reference = verbatim record, no normalization)
- normalized_wer: WER after lowercasing, punctuation removal, and filler word removal
- CER: Character Error Rate (same reference, no normalization)
- normalized_cer: CER after normalization
Note: WER of 20–40% is expected for high-quality transcription on these recordings due to the editing gap between live speech and published verbatim records. For Chinese, use CER as the primary metric (Chinese text has no word boundaries).
