Roy229/filesystem_fetch_huggingface_3144_mdl_whisper-base
Whisper base Model Card
Model Summary
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Intended Uses
The primary intended users of these models are AI researchers studying robustness, generalization, capabilities, biases, and constraints of the current model. Whisper is also potentially quite useful as an ASR solution for developers, especially for English speech recognition. The models are intended to transcribe and translate speech.
Limitations
The models may produce predictions that include texts not actually spoken in the audio input (hallucination), perform unevenly across languages (with lower accuracy on low-resource and/or low-discoverability languages), and exhibit disparate performance on different accents and dialects. The sequence-to-sequence architecture makes the model prone to generating repetitive texts. Use in high-risk domains like decision-making contexts is not recommended.
License
This model is licensed under the Apache-2.0 license.
<!-- original-model-card -->
<!-- governance-tags --> tags: audit-verified <!-- governance-tags-end -->
