datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-speech-ipa-latents.voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-speech-ipa-latents.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.voxcpm2-native-generated-audio-user-ref
VoxCPM2 Native Generated Audio (User Ref)
Raw audio files generated from the native VoxCPM2 path in sglang-omni using a user-provided reference clip.
Contents
9 generated .wav files
metadata.json with prompt text, mode, status, size, and latency
Source Reference Audio
Reference clip used for the reference-mode generations:
https://huggingface.co/datasets/adarshxs/voxcpm2-native-test-samples/resolve/main/data/audio.wav
Files
ref_expressive.wav… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-native-generated-audio-user-ref.voxcpm2-onnx-modelsvoxcpm2-gguf-models
VoxCPM2 GGUF models
This Kaggle dataset contains VoxCPM2 GGUF model files for the VoxCPM.cpp/GGUF backend.
Source repository: bluryar/VoxCPM-GGUF
Revision: main
Layout: gguf
Required backend: gguf-voxcpm-cpp
Default model file: voxcpm2-f16.gguf
Files: 3
The runner expects the full directory to be mounted as a Kaggle input dataset. HF/Nano-vLLM/vLLM-Omni layouts are discovered by config.json, model.safetensors, and audiovae.pth or audiovae.safetensors; GGUF layouts are… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/voxcpm2-gguf-models.meddies-consultant-voxcpm2
Meddies Consultant VoxCPM2 Prepared Text
This public dataset contains the full Vietnamese text-preparation output from
Meddies/meddies-consultant,
prepared for downstream VoxCPM2 speech synthesis.
Contents
prepared.parquet: 1,105,937 accepted VoxCPM-ready generation units.
prepared.jsonl: the same accepted rows in an inspectable JSONL form.
rejects.jsonl: 9,128 rejected units with raw text and rejection details.
summary.json: source, acceptance, and rejection… See the full description on the dataset page: https://huggingface.co/datasets/christian-hoang-04/meddies-consultant-voxcpm2.voxcpm2-python-modelsvoxcpm2-sglang-omni-samples
VoxCPM2 Audio Samples (SGLang Omni)
Audio samples generated by VoxCPM2 running on SGLang Omni.
Samples
File
Language
Input Text
Duration
en_intro.wav
English
Hello! Welcome to VoxCPM2, a text to speech model developed by OpenBMB, now running on SGLang Omni.
6.88s
en_long.wav
English
Artificial intelligence is transforming how we interact with technology. From voice assistants to autonomous vehicles, the possibilities are endless. What excites me most is how… See the full description on the dataset page: https://huggingface.co/datasets/adarshxs/voxcpm2-sglang-omni-samples.voxcpm2-gguf-runtime-portable
VoxCPM2 runtime payload
This private Kaggle dataset is generated by Phorcys.Tools.VoxCPM2RuntimeUploader for PHRunner.Kaggle.Service.VoxCPM2.
Runtime flavor: LinuxCuda
Python tag: python3.10
Generated UTC: 2026-09-09T13:31:02.3368864+00:00
The dataset intentionally contains runtime artifacts, not model weights. Keep the official openbmb/VoxCPM2 checkpoint snapshot in a separate private Kaggle dataset, for example kaggle-pool-account/voxcpm2-python-models.
Top-level runtime… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/voxcpm2-gguf-runtime-portable.voxcpm2-gguf-runtime-cpu-portable
VoxCPM2 runtime payload
This private Kaggle dataset is generated by Phorcys.Tools.VoxCPM2RuntimeUploader for PHRunner.Kaggle.Service.VoxCPM2.
Runtime flavor: LinuxCpu
Python tag: python3.10
Generated UTC: 2026-09-09T14:18:24.2697296+00:00
The dataset intentionally contains runtime artifacts, not model weights. Keep the official openbmb/VoxCPM2 checkpoint snapshot in a separate private Kaggle dataset, for example kaggle-pool-account/voxcpm2-python-models.
Top-level runtime… See the full description on the dataset page: https://huggingface.co/datasets/stokiz/voxcpm2-gguf-runtime-cpu-portable.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-english-ipa-latents.voxcpm2-native-test-samplesvoxcpm2-synthetic-en-v1
Inflect VoxCPM2 Synthetic English v1
Synthetic English speech dataset generated with openbmb/VoxCPM2.
Dataset Summary
This is a multi-voice synthetic English speech dataset prepared for:
TTS fine-tuning
voice-cloning research
synthesis benchmarking
stability, robustness, and post-processing experiments
The audio in this release is synthetic. It is not a corpus of naturally recorded human speech.
Dataset Structure
train/metadata.csv: public release manifest… See the full description on the dataset page: https://huggingface.co/datasets/owensong/voxcpm2-synthetic-en-v1.voxcpm2-french-samplesvoxcpm2-toolsswamiji-voxcpm2-ft-refsswamiji-voxcpm2-refsVoxCPM2-DatasetVoxCPM2-TTS-Dataset
