CoolFace
Datasetpublic

MonlamAI/Bo-voice-v1.0.0

Tibetan STT Benchmark Model Card Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles. ### Dataset Overview Snapshot Date: 15 July 2024, 02:47:06 PM Total Samples: 8,367 audio-transcript pairs. Verification: Every transcript has been reviewed by at least one… See the full description on the dataset page: https://huggingface.co/datasets/MonlamAI/Bo-voice-v1.0.0.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes27downloads
Dataset Card

Tibetan STT Benchmark Model Card

Bo-voice-v1.0.0 is a high-fidelity benchmark for Tibetan Speech-to-Text (STT) technology. It provides a rigorous, multi-domain evaluation set to measure Automatic Speech Recognition (ASR) performance across diverse acoustic environments and speaking styles.


### Dataset Overview

  • Snapshot Date: 15 July 2024, 02:47:06 PM
  • Total Samples: 8,367 audio-transcript pairs.
  • Verification: Every transcript has been reviewed by at least one expert in addition to the original transcriber to ensure ground-truth precision.

### Domain Distribution

The dataset is categorized into eight departments to test model robustness against varied vocabularies and audio qualities.

Dept CodeDomain DescriptionCount
STT_CSChildren's Speech1,367
STT_ABAudio Book1,000
STT_HSHistory1,000
STT_MVTibetan Movies1,000
STT_NSNatural Speech1,000
STT_NWNews1,000
STT_PCPodcast1,000
STT_TTTibetan Teachings1,000

### Quality Grading System

Each entry includes a grade column, reflecting the depth of human verification and transcription quality.

GradeMeaningDescription
1TranscribedInitial transcription completed.
2Team Lead ReviewedVerified by a Team Lead for linguistic consistency.
3QC VerifiedAudited by the Quality Control team for maximum accuracy.

### Evaluation Metrics

To benchmark performance on Bo-voice-v1.0.0, we recommend reporting:

  1. 1.Word Error Rate (WER): The primary metric for transcription accuracy.
  2. 2.Character Error Rate (CER): Vital for assessing Tibetan script-level precision.
  3. 3.Domain-Specific WER: To identify performance gaps in specific categories (e.g., Natural Speech vs. News).
Iterative Purpose: This benchmark is designed to track progress. As models improve, the "Grade 3" verified subset remains the ultimate gold standard for evaluating state-of-the-art performance.