CoolFace
Modelpublic

rtrthryjh/Kokoro-82M-Backpack-TTS

sourceHugging Faceapache-2.0updated 9d agoView on Hugging Face
0likes21downloads
Model Card

Kokoro 82M — Backpack Voice Package

Lightweight local text-to-speech with a curated set of voices. This package stages immutable upstream artifacts for Backpack's voice runtime layer. It does not replace the chat model selected by the user.

Package

FieldValue
Capabilitytext-to-speech
Inputtext
Outputaudio
Parameters82000000
Runtimekokoro
Configured runtime revisionkokoro==0.9.4;en_core_web_sm==3.8.0

| Format | pytorch | | Precision | F32 | | Primary artifact | kokoro-v1_0.pth | | Total packaged-file size | 314.1 MiB | | Recommended RAM | 1.88 GB | | Languages | multilingual |

Validation status

The packager verified the immutable revision, selected-file inventory, non-empty files, hashes, configuration JSON, and the primary artifact container/header. It also loaded the package and passed deterministic audio inference with the configured runtime.

IntegrityMetadataRuntime loadAudio inferenceTokenizer
passedpassedpassedpassedpassed

Run with Kokoro

Install kokoro==0.9.4, construct KModel from config.json and kokoro-v1_0.pth, then pass one of the packaged voices/*.pt files to KPipeline.

Provenance

  • —Upstream: hexgrad/Kokoro-82M
  • —Immutable revision: f3ff3571791e39611d31c381e3a41a3af07b4987
  • —License: apache-2.0
  • —Backpack copied the selected upstream artifacts without modifying model weights.
  • —Backpack did not train this model and does not claim ownership of it.

Files and checksums

  • —config.json — 2.3 KiB — 5abb01e2403b072bf03d04fde160443e209d7a0dad49a423be15196b9b43c17f
  • —kokoro-v1_0.pth — 312.1 MiB — 496dba118d1a58f5f3db2efc88dbdc216e0483fc89fe6e47ee1f2c53f18ad1e4
  • —voices/af_heart.pt — 511.2 KiB — 0ab5709b8ffab19bfd849cd11d98f75b60af7733253ad0d67b12382a102cb4ff
  • —voices/am_michael.pt — 511.2 KiB — 9a443b79a4b22489a5b0ab7c651a0bcd1a30bef675c28333f06971abbd47bd37
  • —voices/bf_emma.pt — 511.2 KiB — d0a423deabf4a52b4f49318c51742c54e21bb89bbbe9a12141e7758ddb5da701
  • —voices/bm_george.pt — 511.2 KiB — f1bc812213dc59774769e5c80004b13eeb79bd78130b11b2d7f934542dab811b

Review the upstream model card and license before use or redistribution. Speech systems can mis-transcribe, synthesize misleading content, or behave differently across languages and accents.