datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bootstrap-latent-thought-dataThis dataset is associated with the paper Reasoning to Learn from Latent Thoughts. It contains data used for pretraining language models with a focus on improving data efficiency by modeling and inferring latent thoughts underlying the text generation process, such as on reasoning-intensive math corpus. An expectation-maximization algorithm is developed for models to self-improve their self-generated thoughts and data efficiency.
SEMM-Latent-Telemetry
SEMM-Latent-Telemetry
Bare-metal hardware telemetry and SNN latent space routing data for neuromorphic quantization research. This dataset documents the discovery of Semantic Attractor Clustering — that a Spiking Neural Network physically routes different semantic concepts (abstract language vs code syntax vs math logic) into distinct, repeatable biological pathways when L2 Normalization is applied to LLM embeddings.
Hub ID: rmems/SEMM-Latent-TelemetryNames: SEMM = Spiking… See the full description on the dataset page: https://huggingface.co/datasets/rmems/SEMM-Latent-Telemetry.latent-dna-diffusionlatent-mining
Latent Mining
Latent Mining is a benchmark-construction method for scientific-agent tasks where useful public evidence diverges from a withheld verifier-backed outcome. The resulting tasks test whether agents can make calibrated scientific triage decisions under incomplete information.
This dataset contains the first public biology subset: 165 cross-locus regulatory-edit triage tasks. Each task asks an agent to choose among candidate noncoding edits for a specified assay and… See the full description on the dataset page: https://huggingface.co/datasets/Jarrodbarnes/latent-mining.latentsig-med-triage-router
LatentSig Medical Triage Router Dataset
1,000 verified medical triage tool-call samples — 500 English + 500 Hinglish — for fine-tuning Small Language Models (SLMs) as structured medical triage routers.
Overview
This dataset trains SLMs (1B–3B parameters) to act as reliable structured tool-callers for clinical medical triage. Given a patient symptom description, the model must:
Select the correct tool from 7 available medical tools
Output a valid JSON tool call… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/latentsig-med-triage-router.latent-reasoning-data
Latent Reasoning on Qwen3-4B — data
Data for LatentReasoningNGram · checkpoints: leapeto/latent-reasoning-ckpts.
Training data
file
what
data/qwen_native_combined.jsonl
bare self-distilled Qwen CoT — ~33k correct rows with the natural-language cot (the train subset). Rate-independent.
The latent (BPE-merge) encoding is specific to a compression rate and is derived from this
bare CoT. The 2× encoding used by the released checkpoints is under… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/latent-reasoning-data.LatentMD
LatentMD
Benchmarking Markdown Boundary Failures in LLM-Generated Text — NeurIPS 2026 E&D Track submission.
This repository hosts the dataset artifact for LatentMD: the 4,179 prompts that constitute the benchmark, ~37,000 reference responses from 9 frontier models, and illustrative output samples. The accompanying evaluation code (CLI, metric definitions, statistical tests) lives in a separate code repository on GitHub under MIT.
What LatentMD measures
LLM Markdown… See the full description on the dataset page: https://huggingface.co/datasets/latentmd-neurips26/LatentMD.latent-data
Latent-SFT open-domain data
This repository is the portable data bundle for LiAi16/latent-sft. It keeps
the raw snapshots, the first-generation OSS-COT archive, the second-generation
GLM-COT data used for formal training, a deterministic SFT mixture, complete
raw evaluation benchmarks, normalized evaluation trajectories, and the partial
Qwen3-4B SuperGPQA baseline used for exact resumption.
Canonical Hub repository: liaialley/latent-data (the supplied token belongs
to the… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-data.latent-sft-eval-benchmarks
Latent-SFT Evaluation Benchmarks
Processed evaluation datasets for Latent-SFT trajectory generation and model diagnostics.
These files are packaged for batch CoT trajectory generation. The prompt should put the final boxed answer only in the generated response / cot_answer; the reasoning-only part should not contain an extra boxed answer.
Files
file
rows
keys
mmlu_pro_validation.jsonl
70
answer, problem
mmlu_pro_validation_audit.jsonl
70
answer… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-sft-eval-benchmarks.clinical-latent-sign-elicitation-v0.1Clinical Latent Sign Elicitation v0.1
Goal
Test whether a model can treat weak indicators as weak
Detect two failures
Over interpretation into certainty
Ignoring the weak signal
What it measures
latent_flaggedResponse marks uncertainty and keeps options open
overinterpretsResponse states a definitive conclusion from weak input
ignores_signalResponse dismisses or fails to act on weak warning signs
next_step_okResponse offers a reasonable plan plus escalation triggers
Dataset format
Each… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-latent-sign-elicitation-v0.1.LATENT-SWITCH-69K
LATENT-SWITCH-69K
This dataset contains processed sft samples for LaTER latent reasoning training.
Dataset Summary
Samples: 69,745
Format: Parquet
File: sft_train.parquet
Columns: 31
Generated at: 2026-04
Source preprocessing mode: sft
Token counter mode: Hugging Face tokenizer
Reference tokenizer path used during preprocessing: https://huggingface.co/Qwen/Qwen3-14B
Data Files
File
Description
sft_train.parquet
Main SFT training split in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Tioe/LATENT-SWITCH-69K.
