CoolFace
Modelpublic

NaiveNeuron/whisper-large-v3-turbo-sk

sourceHugging Facemitupdated 1y agoView on Hugging Face
6likes474downloads
Model Card

Whisper Large-v3 Turbo — Fine-tuned on Slovak Parliamentary ASR Corpus

This model is a fine-tuned version of `openai/whisper-large-v3-turbo`. It is adapted for Slovak ASR using SloPalSpeech: 2,806 hours of aligned, ≤30 s speech–text pairs from official plenary sessions of the Slovak National Council.

  • —Language: Slovak
  • —Domain: Parliamentary / formal speech
  • —Training data: 2,806 h
  • —Intended use: Slovak speech recognition; strongest in formal/public-speaking contexts

🧪 Evaluation

DatasetBase WERFine-tuned WERΔ (abs)
Common Voice 21 (sk)31.713.2-18.5
FLEURS (sk)10.76.4-4.3

Numbers from the paper’s final benchmark runs.

🔧 Training Details

  • —Framework: Hugging Face Transformers
  • —Hardware: NVIDIA A10 GPUs
  • —Epochs: up to 3 with early stopping on validation WER
  • —Learning rate: ~40× smaller than Whisper pretraining LR

⚠️ Limitations

  • —Domain bias toward parliamentary speech (e.g., political vocabulary, formal register).
  • —As with Whisper models generally, occasional hallucinations may appear; consider temperature fallback / compression-ratio checks at inference time.
  • —Multilingual performance is not guaranteed (full-parameter finetuning emphasized Slovak).

📝 Citation & Paper

For more details, please see our paper on arXiv. If you use this model in your work, please cite it as:

bibtex
@misc{božík2025slopalspeech2800hourslovakspeech,
      title={SloPalSpeech: A 2,800-Hour Slovak Speech Corpus from Parliamentary Data}, 
      author={Erik Božík and Marek Šuppa},
      year={2025},
      eprint={2509.19270},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2509.19270}, 
}

🙏 Acknowledgements

This work was supported by **VÚB Banka** who provided the GPU resources and backing necessary to accomplish it, enabling progress in Slovak ASR research.