datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-speech-ipa-latents.layergen-eval-latents
LayerGen — Eval-Set Latents (VAE latents + baked text embeddings)
Pre-encoded evaluation-set inputs for the LayerGen layer-decomposition / harmonization models,
so inference can run anywhere (off-AIP) without the raw video → VAE-encode → umT5-encode pipeline.
Each *.parquet is one clip and is fully self-contained:
column group
contents
{composite,mask,fg,bg}_latent_bytes (+ _shape, _dtype)
4-stream Wan-VAE latents, 81f/21 latent-T, fp16, [16,21,60,104]… See the full description on the dataset page: https://huggingface.co/datasets/cs-mshah/layergen-eval-latents.voxcpm-ghana-latents
VoxCPM Ghana — Precomputed AudioVAE Latents
The exact training-ready data used to fine-tune
ghananlpcommunity/voxcpm-ghana:
precomputed VoxCPM-0.5B AudioVAE latents (16 kHz) for 42 Ghanaian languages +
filtered Ghanaian English, with language-tagged transcripts. Drop-in for
VoxCPM fine-tuning — no audio decoding or VAE encoding needed at train time.
1,756,157 clips · ~3,400 h · 16 kHz
42 Ghanaian languages (incl. Twi split: twi-asante, twi-akuapem) + en
AudioVAE from… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm-ghana-latents.ego10k-vjepa-latents
Ego10k V-JEPA Latents Dataset
This dataset contains compressed, highly-informative Video Joint Embedding Predictive Architecture (V-JEPA) latents extracted from Ego-centric industrial manufacturing videos.
Dataset Structure
The dataset is partitioned into roughly 1GB .parquet chunks using PyArrow.
Data Source and Preprocessing
The latent embeddings in this dataset were systematically extracted from the Ego10k Master Dataset provided by build.ai. The… See the full description on the dataset page: https://huggingface.co/datasets/rookierufus/ego10k-vjepa-latents.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.humanego_serve_bread_lingbot_lerobot_with_latents
HumanEgo Serve Bread LingBot LeRobot With Latents
This dataset contains LeRobot-format robot demonstrations for the task:
pick up the bread and place it on the plate
The repository has two standalone LeRobot-style roots:
humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos.
humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos.
Each split includes:
data/: episode parquet files.
videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.latentsimagenet-latents-e2e-invae-f16d32voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-speech-ipa-latents.capstone_sakuga_vae_latentssit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199LatentSkill
LatentSkill Data
This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents.
Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill
Contents
skill_pretrain/
train.jsonl
val.jsonl
skill_ift/
train.json
search_test/
2wikimultihopqa_test.jsonl
bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.imagenet-latents-sdvae-ft-mse-f8d4-resolution512libero-spatial-latents-v2sit-latents-ode-heun-1000-class-0_1000-samples-segment-400-499cs2-10k-vjepa2-latents-300
CS2-10k V-JEPA2 latent cache (300 matches)
Frozen facebook/vjepa2-vitl-fpc64-256 embeddings of single-POV Counter-Strike 2
gameplay, from 300 matches of RekaAI/CS2-10k
(mirage + dust2). This is the training substrate for the
scale300 hierarchical world model.
Rebuilding it from source takes ~82 hours of wall-clock (elapsed_min 4942.7),
almost all of it network-bound, which is why it is published here.
Contents
file
shape / rows
notes
latents.npy
[6,942… See the full description on the dataset page: https://huggingface.co/datasets/cbctr/cs2-10k-vjepa2-latents-300.sit-latents-ode-heun-segment-0-99ffhq_with_llava_shorter_captions_flux_latentsvoxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-english-ipa-latents.ava-FLUX.1-latents-5k
Dataset Card for AVA FLUX.1-schnell Latents 5k
5k latents from FLUX.1-schnell VAE for the AVA dataset
Dataset Details
Dataset Description
Curated by: Dave Lage
License: Apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Uses
Latents are a sample from the AVA dataset. These latents were created using the FLUX.1-schnell VAE model. Use of these latents is intended for research purposes only. Useful for… See the full description on the dataset page: https://huggingface.co/datasets/rockerBOO/ava-FLUX.1-latents-5k.eurospeech-latents-bf16LATENT-SWITCH-69K
LATENT-SWITCH-69K
This dataset contains processed sft samples for LaTER latent reasoning training.
Dataset Summary
Samples: 69,745
Format: Parquet
File: sft_train.parquet
Columns: 31
Generated at: 2026-04
Source preprocessing mode: sft
Token counter mode: Hugging Face tokenizer
Reference tokenizer path used during preprocessing: https://huggingface.co/Qwen/Qwen3-14B
Data Files
File
Description
sft_train.parquet
Main SFT training split in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Tioe/LATENT-SWITCH-69K.ava-FLUX.1-latents-10k
Dataset Card for AVA FLUX.1-schnell VAE Latents 10k
9.7k latents from FLUX.1-schnell VAE for the AVA dataset
Dataset Details
Dataset Description
Curated by: Dave Lage
License: Apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Uses
Latents are a sample from the AVA dataset. These latents were created using the FLUX.1-schnell VAE model. Use of these latents is intended for research purposes only.… See the full description on the dataset page: https://huggingface.co/datasets/rockerBOO/ava-FLUX.1-latents-10k.sit-latents-ode-heun-1000-test2-class-0_1000-samples-segment-10-19sit-latents-ode-heun-1000-class-0_1000-samples-segment-600-699sit-latents-ode-heun-1000-class-0_1000-samples-segment-500-599cifar64-latentssit-latents-ode-heun-1000-class-0_1000-samples-segment-900-999sit-latents-100-multiply-1000
