ghananlpcommunity/voxcpm-ghana-latents
VoxCPM Ghana — Precomputed AudioVAE Latents The exact training-ready data used to fine-tune ghananlpcommunity/voxcpm-ghana: precomputed VoxCPM-0.5B AudioVAE latents (16 kHz) for 42 Ghanaian languages + filtered Ghanaian English, with language-tagged transcripts. Drop-in for VoxCPM fine-tuning — no audio decoding or VAE encoding needed at train time. 1,756,157 clips · ~3,400 h · 16 kHz 42 Ghanaian languages (incl. Twi split: twi-asante, twi-akuapem) + en AudioVAE from… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm-ghana-latents.
VoxCPM Ghana — Precomputed AudioVAE Latents
The exact training-ready data used to fine-tune `ghananlpcommunity/voxcpm-ghana`: precomputed VoxCPM-0.5B AudioVAE latents (16 kHz) for 42 Ghanaian languages + filtered Ghanaian English, with language-tagged transcripts. Drop-in for VoxCPM fine-tuning — no audio decoding or VAE encoding needed at train time.
- 1,756,157 clips · ~3,400 h · 16 kHz
- 42 Ghanaian languages (incl. Twi split:
twi-asante,twi-akuapem) +en - AudioVAE from `openbmb/VoxCPM-0.5B`: 64-dim, hop 640 (~25 fps)
Format (parquet shards, CODE__*.parquet)
Language tags
Each transcript starts with <|lang:CODE|> , where CODE is the ISO-639-3 code (e.g. ewe, dag, hau, fat), with Twi split into twi-asante / twi-akuapem, and en for English. The model learns the tag as text (VoxCPM is tokenizer-free).
dataset_id map
Source
Derived from `ghananlpcommunity/ghana-speech` and `ghananlpcommunity/ghana-english-tts-filtered`.
