MonlamAI/Bo-voice-v1.0.0
Tibetan STT Benchmark Model Card Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles. ### Dataset Overview Snapshot Date: 15 July 2024, 02:47:06 PM Total Samples: 8,367 audio-transcript pairs. Verification: Every transcript has been reviewed by at least one… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-voice-v1.0.0.
Tibetan STT Benchmark Model Card
Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles.
### Dataset Overview
- Snapshot Date: 15 July 2024, 02:47:06 PM
- Total Samples: 8,367 audio-transcript pairs.
- Verification: Every transcript has been reviewed by at least one expert in addition to the original transcriber to ensure ground-truth precision.
### Domain Distribution
The dataset is categorized into eight departments to test model robustness against varied vocabularies and audio qualities.
### Quality Grading System
Each entry includes a grade column, reflecting the depth of human verification and transcription quality.
### Evaluation Metrics
To benchmark performance on Bo-voice-v1.0.0, we recommend reporting:
- Word Error Rate (WER): The primary metric for transcription accuracy.
- Character Error Rate (CER): Vital for assessing Tibetan script-level precision.
- Domain-Specific WER: To identify performance gaps in specific categories (e.g., Natural Speech vs. News).
Iterative Purpose: This benchmark is designed to track progress. As models improve, the "Grade 3" verified subset remains the ultimate gold standard for evaluating state-of-the-art performance.
