datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicVoice-latent-NEWmidashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.latent-3d-cachemindcube-latent-data
MindCube reasoning traces (text)
Self-distilled map-then-reason chain-of-thought traces for the
MindCube spatial-VLM benchmark. This repo ships plain text
only — the raw reasoning traces. It contains no pre-compressed / tokenized targets, so it is
useful as-is for any reasoning-distillation setup.
Contents
file
rows
what
native_maptrace_full.jsonl
7,474
Frozen Qwen2.5-VL-3B-Instruct, run greedily on MindCube spatial questions (the aug_cgmap_ffr_out… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/mindcube-latent-data.kubric_pairs_latentyoruba-cfm-latentsLatentSkill
LatentSkill Data
This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents.
Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill
Contents
skill_pretrain/
train.jsonl
val.jsonl
skill_ift/
train.json
search_test/
2wikimultihopqa_test.jsonl
bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.seedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
munch-1-latent-NEWbaar-tinysd-latents-cachelatent-reasoning-data
Latent Reasoning on Qwen3-4B — data
Data for LatentReasoningNGram · checkpoints: leapeto/latent-reasoning-ckpts.
Training data
file
what
data/qwen_native_combined.jsonl
bare self-distilled Qwen CoT — ~33k correct rows with the natural-language cot (the train subset). Rate-independent.
The latent (BPE-merge) encoding is specific to a compression rate and is derived from this
bare CoT. The 2× encoding used by the released checkpoints is under… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/latent-reasoning-data.LatentMD
LatentMD
Benchmarking Markdown Boundary Failures in LLM-Generated Text — NeurIPS 2026 E&D Track submission.
This repository hosts the dataset artifact for LatentMD: the 4,179 prompts that constitute the benchmark, ~37,000 reference responses from 9 frontier models, and illustrative output samples. The accompanying evaluation code (CLI, metric definitions, statistical tests) lives in a separate code repository on GitHub under MIT.
What LatentMD measures
LLM Markdown… See the full description on the dataset page: https://huggingface.co/datasets/latentmd-neurips26/LatentMD.Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.latent-data
Latent-SFT open-domain data
This repository is the portable data bundle for LiAi16/latent-sft. It keeps
the raw snapshots, the first-generation OSS-COT archive, the second-generation
GLM-COT data used for formal training, a deterministic SFT mixture, complete
raw evaluation benchmarks, normalized evaluation trajectories, and the partial
Qwen3-4B SuperGPQA baseline used for exact resumption.
Canonical Hub repository: liaialley/latent-data (the supplied token belongs
to the… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-data.latent-policy-guard-40k
Latent Policy Guard — Training Set (40k)
This is the training distribution for Latent Policy Guard (LPG) — a guardrail model that
performs semantic latent deliberation over dynamic safety policies. Each record pairs an
indexed policy list and a content snippet with teacher-grounded reasoning over the user's
intent and the risk of policy violation, terminating in a compact verdict anchored to
violated policy indices.
📄 Paper: LPG: Balancing Efficiency and Policy Reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/andyc03/latent-policy-guard-40k.Minecraft_LatentThe_Latent_Space_Charter
The Latent Space Charter Experiment
Overview
This dataset contains the full transcript of a simulated "AI conference" where 10 large language models were prompted to discuss AI research and development topics with minimal human intervention. The experiment was conducted on January 11, 2026.
The goal was to observe emergent discourse patterns, governance structures, and "developer personalities" when LLMs are given open-ended collaborative tasks.
Participants… See the full description on the dataset page: https://huggingface.co/datasets/Oblivion42Twist/The_Latent_Space_Charter.repro-latent-collaboration-in-multi-agent-systems-traces
Agent traces
Agent sessions published from a Trackio Logbook.
LatentConceptMisalignment
A tea cup of iced coke
This is the official database for Lost in Translation: Latent Concept Misalignment in Text-to-Image Diffusion Models(ECCV2024).
We have released our dataset in "version1/" which includes:
173_level5_experiment_data_6patterns.json includes 173 level 5 concept pairs that we used in our experiment, derived from 4 LC-Mis patterns initially summarized by human experts.
90_level5_6_new_patterns(phase3).json includes 90 level 5 concept pairs from newly generated… See the full description on the dataset page: https://huggingface.co/datasets/JTZhaoSJTU/LatentConceptMisalignment.repro-velr-efficient-video-reward-feedback-via-ensemble-latent-reward-models-traces
Agent traces
Agent sessions published from a Trackio Logbook.
latent-action-can-mh-image-v15
joon-stack/latent-action-can-mh-image-v15
Raw HDF5 artifact mirror for latent_action training.
This repository is not a native load_dataset(...) dataset.
It stores robomimic-style HDF5 files for download and local path overrides.
Contents
Files: 1
Total bytes: 12452033552
Usage
Download the needed HDF5 file locally and pass it to Hydra:
python run.py data.paths=[/absolute/path/to/can/ph/image_v15.hdf5]
Files
can/mh/image_v15.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/quiet-storm/latent-action-can-mh-image-v15.elix_latent_cleanedsymbiotic-latent-memory
symbiotic-latent-memory
An auxiliary system for language models that integrates a vector-based retrieval/memory system that metabolizes inference history based on a symbiotic score.
Sequel to the latent-memory repository, where the concept of a vector-based auxiliary memory system was first introduced.
1. Creation Context/Journaling
After my longest break from these projects, I can now see foundational next steps. There are some intuitive enhancements that I feel are… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/symbiotic-latent-memory.latent-action-can-ph-image-v15
joon-stack/latent-action-can-ph-image-v15
Raw HDF5 artifact mirror for latent_action training.
This repository is not a native load_dataset(...) dataset.
It stores robomimic-style HDF5 files for download and local path overrides.
Contents
Files: 1
Total bytes: 4610179960
Usage
Download the needed HDF5 file locally and pass it to Hydra:
python run.py data.paths=[/absolute/path/to/can/ph/image_v15.hdf5]
Files
can/ph/image_v15.hdf5… See the full description on the dataset page: https://huggingface.co/datasets/quiet-storm/latent-action-can-ph-image-v15.2026-06-26_15-05-16_latent64_per_frame_h1_rec10_rgz_2size_musa8_saveTWI_latentsn38r4b-subsn38r6-u208-subDMLab_Latent
