CoolFace
Modelpublic

jimmymeister/whisper-large-v3-turbo-german-ct2

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
5likes613downloads
Model Card

Important note:

This model is just a CTranslate2 Translation, for usage in CTranslate conform frameworks such as faster-whisper. For any questions about the fine tuning method or the dataset used please refer to the original Repo primeline/whisper-large-v3-turbo-german

Summary

This model map provides information about a model based on Whisper Large v3 that has been fine-tuned for speech recognition in German. Whisper is a powerful speech recognition platform developed by OpenAI. This model has been specially optimized for processing and recognizing German speech.

Applications

This model can be used in various application areas, including

  • —Transcription of spoken German language
  • —Voice commands and voice control
  • —Automatic subtitling for German videos
  • —Voice-based search queries in German
  • —Dictation functions in word processing programs

Model family

ModelParameterslink
Whisper large v3 german1.54Blink
Whisper large v3 turbo german809Mlink
Distil-whisper large v3 german756Mlink
tiny whisper37.8Mlink

Evaluations - Word error rate

Datasetopenai-whisper-large-v3-turboopenai-whisper-large-v3primeline-whisper-large-v3-germannyrahealth-CrisperWhisper (large)primeline-whisper-large-v3-turbo-german
Tuda-De8.3007.8847.7115.1486.441
commonvoice19_03.8493.4843.2151.9273.200
multilingual librispeech3.2032.8322.1292.8152.070
All3.6493.2792.7342.6622.628

The data and code for evaluations are available here

Training data

The training data for this model includes a large amount of spoken German from various sources. The data was carefully selected and processed to optimize recognition performance.

Training process

The training of the model was performed with the following hyperparameters

  • —Batch size: 12288
  • —Epochs: 3
  • —Learning rate: 1e-6
  • —Data augmentation: No
  • —Optimizer: Ademamix

How to use

python
import torch
from transformers import AutoModelForSpeechSeq2Seq, AutoProcessor, pipeline
from datasets import load_dataset
device = "cuda:0" if torch.cuda.is_available() else "cpu"
torch_dtype = torch.float16 if torch.cuda.is_available() else torch.float32
model_id = "primeline/whisper-large-v3-turbo-german"
model = AutoModelForSpeechSeq2Seq.from_pretrained(
    model_id, torch_dtype=torch_dtype, low_cpu_mem_usage=True, use_safetensors=True
)
model.to(device)
processor = AutoProcessor.from_pretrained(model_id)
pipe = pipeline(
    "automatic-speech-recognition",
    model=model,
    tokenizer=processor.tokenizer,
    feature_extractor=processor.feature_extractor,
    max_new_tokens=128,
    chunk_length_s=30,
    batch_size=16,
    return_timestamps=True,
    torch_dtype=torch_dtype,
    device=device,
)
dataset = load_dataset("distil-whisper/librispeech_long", "clean", split="validation")
sample = dataset[0]["audio"]
result = pipe(sample)
print(result["text"])

About us

![primeline AI](https://primeline-ai.com/en/)

Your partner for AI infrastructure in Germany <br> Experience the powerful AI infrastructure that drives your ambitions in Deep Learning, Machine Learning & High-Performance Computing. Optimized for AI training and inference.

Model author: Florian Zimmermeister