CoolFace
Modelpublic

zeromodels/whisper_medium

sourceHugging Faceapache-2.0updated 27d agoView on Hugging Face
0likes32downloads
Model Card

*See [our collection](https://huggingface.co/collections/zeromodels/whisper-6a8eaf34918224be11047f86) for all versions of Whisper.*

Run Whisper with Keras 3: JAX, PyTorch, or TensorFlow

![GitHub](https://github.com/IMvision12/ZeroModels) ![Docs](https://imvision12.github.io/ZeroModels/whisper/) ![Collection](https://huggingface.co/collections/zeromodels/whisper-6a8eaf34918224be11047f86)

zeromodels/whisper_medium

Paper: Robust Speech Recognition via Large-Scale Weak Supervision (arXiv:2212.04356) · HF Papers

Whisper is a multilingual encoder-decoder ASR model trained on large-scale weak supervision. Use task="transcribe" to keep the source language or task="translate" to render English. Pass language=None to let the model detect the spoken language. Output is cased and punctuated.

For more details on the model, please go to the upstream model card.

Pure-Keras 3 conversion of `openai/whisper-medium` for zeromodels. One implementation runs unmodified on TensorFlow / Torch / JAX.

This is an ASR checkpoint (WhisperConditionalGenerate, 769M).

✨ Quick start

python
import os
os.environ["KERAS_BACKEND"] = "torch"  # or "jax" / "tensorflow"

import soundfile as sf
from zeromodels.models.whisper import (
    WhisperProcessor,
    WhisperConditionalGenerate,
)

model = WhisperConditionalGenerate.from_weights("zeromodels/whisper_medium")
processor = WhisperProcessor.from_weights("zeromodels/whisper_medium")

audio, sr = sf.read("your_audio.wav", dtype="float32")  # 16 kHz mono
# task="transcribe" keeps the source language; "translate" -> English.
text = model.generate(audio, processor, language="en", task="transcribe")
print(repr(text[0]))

Load any Whisper variant the same way with from_weights("zeromodels/<variant>"):

VariantHubNotes
whisper_tiny`zeromodels/whisper_tiny`39M
whisper_base`zeromodels/whisper_base`74M
whisper_small`zeromodels/whisper_small`244M
whisper_medium`zeromodels/whisper_medium`769M
whisper_large`zeromodels/whisper_large`1.55B
whisper_large_v2`zeromodels/whisper_large_v2`1.55B
whisper_large_v3`zeromodels/whisper_large_v3`128 mel bins
whisper_large_v3_turbo`zeromodels/whisper_large_v3_turbo`4 decoder layers

Tips

  • Set KERAS_BACKEND before importing Keras / zeromodels.
  • Prefer WhisperProcessor.from_weights(...) so mel bins match the variant (v3 uses 128).
  • Clips are padded to a 30 s window; chunk longer audio yourself.
  • See Whisper docs and Loading Weights.
  • Community / upstream safetensors still work via the hf: prefix, e.g. WhisperConditionalGenerate.from_weights("hf:openai/whisper-medium").

Special Thanks

A huge thank you to the OpenAI Whisper authors for creating and releasing these models.

License: Apache 2.0.