datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.IndicVoice-latent-NEWmidashenglm-gen-training-latents
ModelsLab/midashenglm-gen-training-latents
Precomputed audio latents for fine-tuning
mispeech/midashenglm-gen,
paired with six-view prompts in the exact format the model was trained on.
This is not an audio dataset and not a caption dataset. Each record is the
output of the model's frozen DashengTokenizer encoder — 768-dimensional latents
at 25 Hz, stored float16 — next to the tagged prompt string built from the
source metadata.
Why it exists
The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.egopi_latal_openarm_bottlelatxa-corpus-v1.1
Latxa Corpus v1.1
This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2.
💻 Repository: https://github.com/hitz-zentroa/latxa
📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque
📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque
📧 Point of Contact: hitz@ehu.eus
📌 Notice
As of February 13th 2026, this repository reflects a curated version of the original dataset.
Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.latent-3d-cacheegopi_latal_openarm_snackegopi_latal_humanegopi_latal_openarm_cupexams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.egopi_latal_openarm_dollmindcube-latent-data
MindCube reasoning traces (text)
Self-distilled map-then-reason chain-of-thought traces for the
MindCube spatial-VLM benchmark. This repo ships plain text
only — the raw reasoning traces. It contains no pre-compressed / tokenized targets, so it is
useful as-is for any reasoning-distillation setup.
Contents
file
rows
what
native_maptrace_full.jsonl
7,474
Frozen Qwen2.5-VL-3B-Instruct, run greedily on MindCube spatial questions (the aug_cgmap_ffr_out… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/mindcube-latent-data.MoeGirlPedia_zh_cleaned_latest
🌐Language 中文|English
本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。
特色
⚡体积优势
🧠文本易理解
💬更符合中文语境
仅经过基础清洗的数据集
1.06GB
存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚
暴力抹除非中文文字,导致信息缺失严重
本数据集
0.74GB(30.2%↓)
通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语)
仅经过基础清洗的数据集
size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。
'''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}}
url=akanejunpei.jpg|position=up
图片说明=999中的茜(2027,21岁)
|本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.latxa-corpus-v2
Latxa Corpus v2
📧 Point of Contact: hitz@ehu.eus
Dataset Summary
Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU)
Language(s): eu-ES
Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources.
Compared to v1.1, it substantially increases coverage, diversity, and volume.
The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.kubric_pairs_latentyoruba-cfm-latentsLatentSkill
LatentSkill Data
This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents.
Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill
Contents
skill_pretrain/
train.jsonl
val.jsonl
skill_ift/
train.json
search_test/
2wikimultihopqa_test.jsonl
bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.cc100-latin
Latin part of cc100 corpus
This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface.
Preprocessing
I undertook the following preprocessing steps:
Removal of all "pseudo-Latin" text ("Lorem ipsum ...").
Use of CLTK for sentence splitting and normalisation.
Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.COPA-SR_lat
COPA-SR_lat
(The dataset uses latin script. For the original (cyrillic) version, see this dataset.)
The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology , transliterated into Latin script.
The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR_lat.seedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tla-late_egyptian-v19-premium
Dataset Card for Dataset tla-Late_Egyptian-v19-premium
This data set contains Late Egyptian sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation.
The data comes from the database of the Thesaurus Linguae Aegyptiae, corpus version 19.
This set of Late Egyptian sentences only contains text witnesses classified as "Late Egyptian" in the TLA corpus metadata. Moreover, it contains only fully intact,
unambiguously readable… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-late_egyptian-v19-premium.repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces
Agent traces
Agent sessions published from a Trackio Logbook.
munch-1-latent-NEWzhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora.
Download
You can download the latest Chinese Wikipedia dump from the following link:
Chinese Wikipedia Dump
English Wikipedia Dump (For reference)
Extraction
After you download the dump, you can extract the data using the following commands:
# install wikiextractor
pip install wikiextractor
# extract the data
wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2
Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.Defensive-LLM
Defensive-LLM — Training Data Samples
Preview samples from the Lateos defensive LLM training pipeline. Each file contains a deterministic 5% sample (seed 42) of the full corpus, drawn from the exact production datasets used to train Homeland Defender — a frontier LLM for OT/ICS vulnerability discovery and defensive security analysis.
These samples let you evaluate schema, quality, and provenance before licensing the full datasets. All data is structural and defensive: no… See the full description on the dataset page: https://huggingface.co/datasets/Lateos/Defensive-LLM.latent-reasoning-data
Latent Reasoning on Qwen3-4B — data
Data for LatentReasoningNGram · checkpoints: leapeto/latent-reasoning-ckpts.
Training data
file
what
data/qwen_native_combined.jsonl
bare self-distilled Qwen CoT — ~33k correct rows with the natural-language cot (the train subset). Rate-independent.
The latent (BPE-merge) encoding is specific to a compression rate and is derived from this
bare CoT. The 2× encoding used by the released checkpoints is under… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/latent-reasoning-data.LatentMD
LatentMD
Benchmarking Markdown Boundary Failures in LLM-Generated Text — NeurIPS 2026 E&D Track submission.
This repository hosts the dataset artifact for LatentMD: the 4,179 prompts that constitute the benchmark, ~37,000 reference responses from 9 frontier models, and illustrative output samples. The accompanying evaluation code (CLI, metric definitions, statistical tests) lives in a separate code repository on GitHub under MIT.
What LatentMD measures
LLM Markdown… See the full description on the dataset page: https://huggingface.co/datasets/latentmd-neurips26/LatentMD.llm-api-pricing-latency-2026
LLM Inference Unit Economics & Architecture Engine
Empirical benchmark dataset by Groundwork Research (https://gworky.com).
Full interactive decision engine available at: https://gworky.com/tools/llm-token-cost-calculator.
Description
Full-stack inference cost and latency estimator comparing frontier proprietary models (Claude 3.7, GPT-4.5) against open-weight hosted providers (Groq, DeepSeek R1, Together AI).
Primary source authority: https://gworky.com/tech
Latent-Resonance-AI-Image-Forensics-Benchmark-N100
Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale)
Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026)
Benchmark Overview
This repository provides:
The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.latent-data
Latent-SFT open-domain data
This repository is the portable data bundle for LiAi16/latent-sft. It keeps
the raw snapshots, the first-generation OSS-COT archive, the second-generation
GLM-COT data used for formal training, a deterministic SFT mixture, complete
raw evaluation benchmarks, normalized evaluation trajectories, and the partial
Qwen3-4B SuperGPQA baseline used for exact resumption.
Canonical Hub repository: liaialley/latent-data (the supplied token belongs
to the… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-data.
