datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-speech-ipa-latents.layergen-eval-latents
LayerGen — Eval-Set Latents (VAE latents + baked text embeddings)
Pre-encoded evaluation-set inputs for the LayerGen layer-decomposition / harmonization models,
so inference can run anywhere (off-AIP) without the raw video → VAE-encode → umT5-encode pipeline.
Each *.parquet is one clip and is fully self-contained:
column group
contents
{composite,mask,fg,bg}_latent_bytes (+ _shape, _dtype)
4-stream Wan-VAE latents, 81f/21 latent-T, fp16, [16,21,60,104]… See the full description on the dataset page: https://huggingface.co/datasets/cs-mshah/layergen-eval-latents.Latent-Earth
Latent Earth: An Atlas of Architecture in Flux.2
200,000 images of 40,000 places on Earth, each rendered by a single
image model in a single state of its training, with five internal
representations recorded for every image while it was being generated.
Nothing else enters. Each prompt contains only a place's name; no
photographs, no maps, no climate records correct what the model proposes.
This is therefore not a depiction of the world but a probe of the model: a
survey of what… See the full description on the dataset page: https://huggingface.co/datasets/Punktiert/Latent-Earth.got-activations-llama3.1-405b-base
meta-llama/Llama-3.1-405B — Activation Dataset
Cached activations extracted from meta-llama/Llama-3.1-405B (revision unknown).
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-125
16384
-
12
-
Prompts: 7660
Format version: 1.1
Load with lmprobe
from lmprobe import pull_dataset, load_activation_dataset
# Option 1: Pull into local cache (enables probe training without re-extraction)… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-llama3.1-405b-base.ego10k-vjepa-latents
Ego10k V-JEPA Latents Dataset
This dataset contains compressed, highly-informative Video Joint Embedding Predictive Architecture (V-JEPA) latents extracted from Ego-centric industrial manufacturing videos.
Dataset Structure
The dataset is partitioned into roughly 1GB .parquet chunks using PyArrow.
Data Source and Preprocessing
The latent embeddings in this dataset were systematically extracted from the Ego10k Master Dataset provided by build.ai. The… See the full description on the dataset page: https://huggingface.co/datasets/rookierufus/ego10k-vjepa-latents.voxcpm-ghana-latents
VoxCPM Ghana — Precomputed AudioVAE Latents
The exact training-ready data used to fine-tune
ghananlpcommunity/voxcpm-ghana:
precomputed VoxCPM-0.5B AudioVAE latents (16 kHz) for 42 Ghanaian languages +
filtered Ghanaian English, with language-tagged transcripts. Drop-in for
VoxCPM fine-tuning — no audio decoding or VAE encoding needed at train time.
1,756,157 clips · ~3,400 h · 16 kHz
42 Ghanaian languages (incl. Twi split: twi-asante, twi-akuapem) + en
AudioVAE from… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm-ghana-latents.SEMM-Latent-Telemetry
SEMM-Latent-Telemetry
Bare-metal hardware telemetry and SNN latent space routing data for neuromorphic quantization research. This dataset documents the discovery of Semantic Attractor Clustering — that a Spiking Neural Network physically routes different semantic concepts (abstract language vs code syntax vs math logic) into distinct, repeatable biological pathways when L2 Normalization is applied to LLM embeddings.
Hub ID: rmems/SEMM-Latent-TelemetryNames: SEMM = Spiking… See the full description on the dataset page: https://huggingface.co/datasets/rmems/SEMM-Latent-Telemetry.midashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.got-activations-qwen2.5-0.5b
Qwen/Qwen2.5-0.5B — Activation Dataset
Cached activations extracted from Qwen/Qwen2.5-0.5B (revision 060db6499f32faf8b98477b0a26969ef7d8b9987).
Full-sequence activations (24 layers, 896 dim, float16) and top-100 logits from Qwen/Qwen2.5-0.5B on 7,660 Geometry of Truth statements. Per-layer sharding (v1.2) with independent shard boundaries.
Contents
Tensor
Layers
Dim
Pooling
Shards
Row Bytes
hidden_layers
0-23
896
-
1
-
logits_topk
-
k=100
last_token
1
1200… See the full description on the dataset page: https://huggingface.co/datasets/latent-lab/got-activations-qwen2.5-0.5b.latent-3d-cachelatent-image-training
squiggles (metadata-fix)
OC-map FEM rebuild at 35 pixels per wavelength, with corrected geometries,
Helmholtz residuals, and the resolved JCMsuite .jcm / .jcmp files used
for each solve.
Configs
metadata (default)
One row per structure folder (sample_XXXX). Geometry comes from published
optical-constant maps (not the old nested-interface metadata).
validation
One row per FEM incidence (theta in {0, 45}). Self-contained pixel map:… See the full description on the dataset page: https://huggingface.co/datasets/als-rixs/latent-image-training.humanego_serve_bread_lingbot_lerobot_with_latents
HumanEgo Serve Bread LingBot LeRobot With Latents
This dataset contains LeRobot-format robot demonstrations for the task:
pick up the bread and place it on the plate
The repository has two standalone LeRobot-style roots:
humanego_serve_bread_lingbot_eef_train: 55 episodes, 41,603 frames, 55 videos.
humanego_serve_bread_lingbot_eef_val: 6 episodes, 5,533 frames, 6 videos.
Each split includes:
data/: episode parquet files.
videos/: MP4 videos for observation.images.ego_rgb.… See the full description on the dataset page: https://huggingface.co/datasets/Coffeecoderss/humanego_serve_bread_lingbot_lerobot_with_latents.IndicVoice-latent-NEW-parquetmindcube-latent-data
MindCube reasoning traces (text)
Self-distilled map-then-reason chain-of-thought traces for the
MindCube spatial-VLM benchmark. This repo ships plain text
only — the raw reasoning traces. It contains no pre-compressed / tokenized targets, so it is
useful as-is for any reasoning-distillation setup.
Contents
file
rows
what
native_maptrace_full.jsonl
7,474
Frozen Qwen2.5-VL-3B-Instruct, run greedily on MindCube spatial questions (the aug_cgmap_ffr_out… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/mindcube-latent-data.voxcpm2-ghana-speech-ipa-latents
VoxCPM2 Ghana — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghananlpcommunity/voxcpm2-ghana-speech-ipa-latents.latentssit-latents-ode-heun-1000-class-0_1000-samples-segment-100-199capstone_sakuga_vae_latentsimagenet-latents-e2e-invae-f16d32kubric_pairs_latentsort5_20260828_172447This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/sort5_20260828_172447.sort6_20260828_173921This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/sort6_20260828_173921.sort3_20260828_165120This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/sort3_20260828_165120.robocasa_pickplace_countertosink1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.robot0_eye_in_hand": {
"dtype": "video",
"shape": [
256,
256,
3
],
"names": [
"height",
"width",
"channel"
],
"video_info": {… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/robocasa_pickplace_countertosink1.sort2_20260828_164039This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/sort2_20260828_164039.latent-taxonomy-samplessort11_20260829_113725This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/latentforce/sort11_20260829_113725.flux2-klein-latent-trajectories
FLUX.2 Klein Latent Trajectories
This dataset contains synthetic text-to-image generations produced with a local FLUX.2 Klein 4B snapshot, together with the prompts, generated WebP images, and captured intermediate diffusion latent states.
Contents
262,144 generated examples.
512 Torch shard files under shards/, with 512 samples per shard.
One JSON sidecar per shard with generation metadata.
Image resolution: 512 x 512.
Image format: WebP, quality 90.
Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/ryanhlewis/flux2-klein-latent-trajectories.LatentSkill
LatentSkill Data
This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents.
Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill
Contents
skill_pretrain/
train.jsonl
val.jsonl
skill_ift/
train.json
search_test/
2wikimultihopqa_test.jsonl
bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.voxcpm2-ghana-english-ipa-latents
VoxCPM2 Ghanaian English — Precomputed AudioVAE Latents
Precomputed VoxCPM-2 AudioVAE latents for a Ghanaian multilingual TTS fine-tune,
ready for training with the official train_voxcpm_finetune.py
(train_manifest: ghana-latents). No audio decoding or VAE encoding needed at train
time — the latent feat is fed straight to the VoxCPM-2 model with the IPA transcript.
Each language is a dataset subset:
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/voxcpm2-ghana-english-ipa-latents.
