ghanaopenai/voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune, ready for training with the official train_voxcpm_finetune.py (train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript. Each language is a dataset subset: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune, ready for training with the official train_voxcpm_finetune.py (train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds = load_dataset("ghanaopendata/voxcpm2-ghana-english-ipa-latents", "English_eng", split="train")- 52,855 clips · ~201 h across 1 subsets
- Ghanaian English (
English_eng) - AudioVAE from VoxCPM-2: 64-dim, hop 640 (~25 fps)
Subsets (config_name = language code)
English_eng
Phonemization
Transcripts in text are IPA phonemes (space-separated, punctuation kept).
- Transcripts are read off the audio by ASR, in the
ghana-english-g2pconvention — see below. - The 42 Ghanaian languages are at
ghanaopendata/voxcpm2-ghana-speech-ipa-latents. - Audio for these same clips:
ghanaopendata/ghana-english-speech-ipa.
Inference note: a TTS fine-tuned on these latents was trained on IPA text. To synthesize new utterances, phonemize the input text into the same space-separated IPA convention (ghana-g2pfor Ghanaian languages,ghana-english-g2pfor English) and pass that IPA string to the model. Note the training labels are ASR-read pronunciations while a g2p phonemizer produces predicted ones, so the two conventions agree on inventory but can differ on individual words.
Format (each subset: parquet shards CODE__*.parquet)
text holds IPA rather than orthography because the finetune script re-tokenizes text directly. raw_text keeps the orthography available for inspection, filtering, and training a grapheme model on the same latents.
Source
Derived from `ghanaopendata/ghana-english-tts-clean2` (first 200 h).
How the IPA was produced
Read off the audio by `ghananlpcommunity/ghana-english-phoneme-asr`, a CTC recogniser finetuned from the 42-language Ghanaian model on ghana-english-g2p targets.
That detail matters for anyone synthesising from this data. An earlier version of this dataset used the Ghanaian recogniser directly, which transcribes English in its own convention — 76% UER against ghana-english-g2p, with symbols like ð and iː largely absent. Since inference has no ASR, only a g2p, those labels could not be reproduced at synthesis time. The English recogniser was built specifically to close that gap: it reads what the speaker actually said, but writes it in the g2p convention, reaching 16.7% UER against g2p on held-out clips. The residual is mostly genuine pronunciation difference rather than convention mismatch, which is the point.
Transcription used per-utterance input normalisation and 6-second windowing (the encoder collapses silently on long audio; these clips average 13.8 s).
Pipeline: `GhanaNLP/ghana-speech-english-ipa-latents-data-prep`
