voxcpm2
Datasets
All datasets matching “voxcpm2”voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-speech-ipa-latents.voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-speech-ipa-latents.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.voxcpm2-native-generated-audio-user-ref
VoxCPM2 Native Generated Audio (User Ref)
Raw audio files generated from the native VoxCPM2 path in sglang-omni using a user-provided reference clip.
Contents
9 generated .wav files
metadata.json with prompt text, mode, status, size, and latency
Source Reference Audio
Reference clip used for the reference-mode generations:
https://huggingface.co/datasets/adarshxs/voxcpm2-native-test-samples/resolve/main/data/audio.wav
Files
ref_expressive.wav… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-native-generated-audio-user-ref.voxcpm2-onnx-modelsvoxcpm2-gguf-models
VoxCPM2 GGUF models
This Kaggle dataset contains VoxCPM2 GGUF model files for the VoxCPM.cpp/GGUF backend.
Source repository: bluryar/VoxCPM-GGUF
Revision: main
Layout: gguf
Required backend: gguf-voxcpm-cpp
Default model file: voxcpm2-f16.gguf
Files: 3
The runner expects the full directory to be mounted as a Kaggle input dataset. HF/Nano-vLLM/vLLM-Omni layouts are discovered by config.json, model.safetensors, and audiovae.pth or audiovae.safetensors; GGUF layouts are… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/voxcpm2-gguf-models.
