CoolFace
Datasetpublic

ghanaopenai/voxcpm2-ghana-english-ipa-latents

VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune, ready for training with the official train_voxcpm_finetune.py (train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript. Each language is a dataset subset: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
0likes151downloads
Dataset Card

VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents

Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune, ready for training with the official train_voxcpm_finetune.py (train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.

Each language is a dataset subset:

python
from datasets import load_dataset
ds = load_dataset("ghanaopendata/voxcpm2-ghana-english-ipa-latents", "English_eng", split="train")
  • —52,855 clips · ~201 h across 1 subsets
  • —Ghanaian English (English_eng)
  • —AudioVAE from VoxCPM-2: 64-dim, hop 640 (~25 fps)

Subsets (config_name = language code)

  • —English_eng

Phonemization

Transcripts in text are IPA phonemes (space-separated, punctuation kept).

  • —Transcripts are read off the audio by ASR, in the ghana-english-g2p convention — see below.
  • —The 42 Ghanaian languages are at ghanaopendata/voxcpm2-ghana-speech-ipa-latents.
  • —Audio for these same clips: ghanaopendata/ghana-english-speech-ipa.
Inference note: a TTS fine-tuned on these latents was trained on IPA text. To synthesize new utterances, phonemize the input text into the same space-separated IPA convention (ghana-g2p for Ghanaian languages, ghana-english-g2p for English) and pass that IPA string to the model. Note the training labels are ASR-read pronunciations while a g2p phonemizer produces predicted ones, so the two conventions agree on inventory but can differ on individual words.

Format (each subset: parquet shards CODE__*.parquet)

columntypenotes
idstringclip id, joinable back to the source audio dataset (a few source shards repeat ids, so it is not a unique key)
featbinaryfp16 AudioVAE latent, shape [feat_t, 64]. Reconstruct: np.frombuffer(feat, np.float16).reshape(feat_t, 64)
feat_tint32number of latent frames (~25 per second)
textstringIPA phoneme transcript (space-separated, punctuation kept) — the training input
raw_textstringthe original written transcript in the language's orthography
text_idslist[int32]Llama token ids of text (informational; train re-tokenizes)
dataset_idint32language id 42 (continues the numbering of the 42 Ghanaian configs, so the two datasets can be trained together)
splitstringtrain / dev (first 40 clips per language)

text holds IPA rather than orthography because the finetune script re-tokenizes text directly. raw_text keeps the orthography available for inspection, filtering, and training a grapheme model on the same latents.

Source

Derived from `ghanaopendata/ghana-english-tts-clean2` (first 200 h).

How the IPA was produced

Read off the audio by `ghananlpcommunity/ghana-english-phoneme-asr`, a CTC recogniser finetuned from the 42-language Ghanaian model on ghana-english-g2p targets.

That detail matters for anyone synthesising from this data. An earlier version of this dataset used the Ghanaian recogniser directly, which transcribes English in its own convention — 76% UER against ghana-english-g2p, with symbols like ð and iː largely absent. Since inference has no ASR, only a g2p, those labels could not be reproduced at synthesis time. The English recogniser was built specifically to close that gap: it reads what the speaker actually said, but writes it in the g2p convention, reaching 16.7% UER against g2p on held-out clips. The residual is mostly genuine pronunciation difference rather than convention mismatch, which is the point.

Transcription used per-utterance input normalisation and 6-second windowing (the encoder collapses silently on long audio; these clips average 13.8 s).

Pipeline: `GhanaNLP/ghana-speech-english-ipa-latents-data-prep`