CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WindyVerse /Handwritten-Latex-Datasets Dataset This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets. Dataset source Collected in various junior high schools and high schools, handwritten by students. Usage The label is stored at json folder and scanned hand-writted pictures are stored at pic folder. Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.imageimage-to-text1K<n<10K1 likes4.8k downloads3y agoHugging Face02zuhri025 /IndicVoice-latent-NEWtext100K<n<1M0 likes905 downloads5mo agoHugging Face03ModelsLab /midashenglm-gen-training-latents ModelsLab/midashenglm-gen-training-latents Precomputed audio latents for fine-tuning mispeech/midashenglm-gen, paired with six-view prompts in the exact format the model was trained on. This is not an audio dataset and not a caption dataset. Each record is the output of the model's frozen DashengTokenizer encoder — 768-dimensional latents at 25 Hz, stored float16 — next to the tagged prompt string built from the source metadata. Why it exists The encoder is frozen… See the full description on the dataset page: https://huggingface.co/datasets/ModelsLab/midashenglm-gen-training-latents.tabulartext-to-audion<1K0 likes649 downloads1mo agoHugging Face04junbrro /egopi_latal_openarm_bottletabularn<1K0 likes625 downloads2mo agoHugging Face05HiTZ /latxa-corpus-v1.1 Latxa Corpus v1.1 This is the training corpus of Latxa v1.1, a family of large language models for Basque based on Llama 2. 💻 Repository: https://github.com/hitz-zentroa/latxa 📒 Blog Post: Latxa: An Open Language Model and Evaluation Suite for Basque 📖 Paper: Latxa: An Open Language Model and Evaluation Suite for Basque 📧 Point of Contact: hitz@ehu.eus 📌 Notice As of February 13th 2026, this repository reflects a curated version of the original dataset. Some data… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v1.1.textfill-mask1M<n<10M2 likes568 downloads7mo agoHugging Face06henryyzhaoo /latent-3d-cacheimagen<1K0 likes568 downloads4mo agoHugging Face07junbrro /egopi_latal_openarm_snacktabularn<1K0 likes532 downloads2mo agoHugging Face08junbrro /egopi_latal_humantabularn<1K0 likes519 downloads2mo agoHugging Face09junbrro /egopi_latal_openarm_cuptabularn<1K0 likes512 downloads2mo agoHugging Face10natnitaract /exams-basic-and-quantum-cryptography-and-security-latex Open Problem Exams: Cryptography and Security (LaTeX) A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions. Dataset Overview Institution Files Topics Questions Caltech & TU Delft 8 38 145 EPFL 6 19 86 ETH Zurich 1 14 37 MIT 3 33 79 Total 18 104 347 Difficulty Distribution Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.textn<1K1 likes470 downloads6mo agoHugging Face11junbrro /egopi_latal_openarm_dolltabularn<1K0 likes457 downloads2mo agoHugging Face12leapeto /mindcube-latent-data MindCube reasoning traces (text) Self-distilled map-then-reason chain-of-thought traces for the MindCube spatial-VLM benchmark. This repo ships plain text only — the raw reasoning traces. It contains no pre-compressed / tokenized targets, so it is useful as-is for any reasoning-distillation setup. Contents file rows what native_maptrace_full.jsonl 7,474 Frozen Qwen2.5-VL-3B-Instruct, run greedily on MindCube spatial questions (the aug_cgmap_ffr_out… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/mindcube-latent-data.tabularvisual-question-answering100K<n<1M0 likes353 downloads16d agoHugging Face13YCWTG /MoeGirlPedia_zh_cleaned_latest 🌐Language 中文|English 本数据集由2025年10月萌娘百科的快照经过清洗得来,专用于预训练等文本生成相关的模型训练。 特色 ⚡体积优势 🧠文本易理解 💬更符合中文语境 仅经过基础清洗的数据集 1.06GB 存在复杂的网址链接残留的html标记正文内容被清除后残存的标题牛皮癣一样的引文注脚 暴力抹除非中文文字,导致信息缺失严重 本数据集 0.74GB(30.2%↓) 通过多重工序清洗基本不存在难以理解的文本内容保留部分英文以及少量其他语言文字(如日语) 仅经过基础清洗的数据集 size=66px|color=#8230FF|她已经不是我所认识的那个-{zh-hans:茜;zh-hant:仓式茜}-了。 '''仓式 茜'''(Kurashiki Akane)是由Spike Chunsoft所创作的系列游戏'''《极限脱出》'''及其衍生作品的主要角色之一。{{ZETOP}} url=akanejunpei.jpg|position=up 图片说明=999中的茜(2027,21岁) |本名=仓式 茜(くらしき… See the full description on the dataset page: https://huggingface.co/datasets/YCWTG/MoeGirlPedia_zh_cleaned_latest.texttext-generation100K<n<1M4 likes344 downloads11mo agoHugging Face14HiTZ /latxa-corpus-v2 Latxa Corpus v2 📧 Point of Contact: hitz@ehu.eus Dataset Summary Curated by: HiTZ Research Center & IXA Research group (University of the Basque Country UPV/EHU) Language(s): eu-ES Latxa Corpus v2 is a large-scale monolingual Basque corpus, created by combining curated crawls, public datasets, institutional data, and newly collected resources. Compared to v1.1, it substantially increases coverage, diversity, and volume. The final corpus is deduplicated, filtered, and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/latxa-corpus-v2.textfill-mask1M<n<10M1 likes320 downloads7mo agoHugging Face15sunovivid /kubric_pairs_latenttabularn<1K0 likes246 downloads8mo agoHugging Face16FloatinggOnion /yoruba-cfm-latentstextn<1K0 likes180 downloads4mo agoHugging Face17AofaYu71 /LatentSkill LatentSkill Data This dataset repository contains the data released for LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents. Code: https://github.com/yuaofan0-oss/LatentSkillPaper: https://arxiv.org/abs/2606.06087Checkpoint repository: https://huggingface.co/AofaYu71/LatentSkill Contents skill_pretrain/ train.jsonl val.jsonl skill_ift/ train.json search_test/ 2wikimultihopqa_test.jsonl bamboogle_test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/AofaYu71/LatentSkill.tabularquestion-answering100K<n<1M4 likes174 downloads2mo agoHugging Face18pstroe /cc100-latin Latin part of cc100 corpus This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface. Preprocessing I undertook the following preprocessing steps: Removal of all "pseudo-Latin" text ("Lorem ipsum ..."). Use of CLTK for sentence splitting and normalisation. Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.textn<1K9 likes150 downloads4y agoHugging Face19classla /COPA-SR_lat COPA-SR_lat (The dataset uses latin script. For the original (cyrillic) version, see this dataset.) The COPA-SR dataset (Choice of plausible alternatives in Serbian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology , transliterated into Latin script. The dataset consists of 1,000 premises (My body cast a shadow over the grass), each given a question (What is the cause? / What happened as a result?), and two choices (The sun was… See the full description on the dataset page: https://huggingface.co/datasets/classla/COPA-SR_lat.tabulartext-classification1K<n<10K0 likes138 downloads3y agoHugging Face20luyu1021 /seedance_general_all_dance_scm_latent_lmdb Seedance General-All + Dance SCM Latent LMDB This dataset stores precomputed SCM latents used for TurboT2AV training. Source mapping: seedance_general_all_dance_mapping.csv Successful latent samples: 44,305 Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007 Video latent shape per sample: (1, 16, 128, 16, 24) Audio latent shape per sample: (1, 127, 128) The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.tabulartext-to-video10K<n<100K0 likes130 downloads3mo agoHugging Face21thesaurus-linguae-aegyptiae /tla-late_egyptian-v19-premium Dataset Card for Dataset tla-Late_Egyptian-v19-premium This data set contains Late Egyptian sentences in hieroglyphs and transliteration, with lemmatization, with POS glossing and with a German translation. The data comes from the database of the Thesaurus Linguae Aegyptiae, corpus version 19. This set of Late Egyptian sentences only contains text witnesses classified as "Late Egyptian" in the TLA corpus metadata. Moreover, it contains only fully intact, unambiguously readable… See the full description on the dataset page: https://huggingface.co/datasets/thesaurus-linguae-aegyptiae/tla-late_egyptian-v19-premium.texttranslation1K<n<10K4 likes123 downloads2y agoHugging Face22Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes115 downloads2mo agoHugging Face23zuhri025 /munch-1-latent-NEWtext10K<n<100K0 likes111 downloads5mo agoHugging Face24Jax-dan /zhwiki-latestThis repository demonstrates access to the latest Chinese Wikipedia corpora. Download You can download the latest Chinese Wikipedia dump from the following link: Chinese Wikipedia Dump English Wikipedia Dump (For reference) Extraction After you download the dump, you can extract the data using the following commands: # install wikiextractor pip install wikiextractor # extract the data wikiextractor --json -o <output_dir> zhwiki-latest-pages-articles.xml.bz2 Then, you… See the full description on the dataset page: https://huggingface.co/datasets/Jax-dan/zhwiki-latest.textfill-mask1M<n<10M0 likes102 downloads1y agoHugging Face25Lateos /Defensive-LLM Defensive-LLM — Training Data Samples Preview samples from the Lateos defensive LLM training pipeline. Each file contains a deterministic 5% sample (seed 42) of the full corpus, drawn from the exact production datasets used to train Homeland Defender — a frontier LLM for OT/ICS vulnerability discovery and defensive security analysis. These samples let you evaluate schema, quality, and provenance before licensing the full datasets. All data is structural and defensive: no… See the full description on the dataset page: https://huggingface.co/datasets/Lateos/Defensive-LLM.tabulartext-generationn<1K1 likes96 downloads2mo agoHugging Face26leapeto /latent-reasoning-data Latent Reasoning on Qwen3-4B — data Data for LatentReasoningNGram · checkpoints: leapeto/latent-reasoning-ckpts. Training data file what data/qwen_native_combined.jsonl bare self-distilled Qwen CoT — ~33k correct rows with the natural-language cot (the train subset). Rate-independent. The latent (BPE-merge) encoding is specific to a compression rate and is derived from this bare CoT. The 2× encoding used by the released checkpoints is under… See the full description on the dataset page: https://huggingface.co/datasets/leapeto/latent-reasoning-data.texttext-generation10K<n<100K0 likes93 downloads3mo agoHugging Face27latentmd-neurips26 /LatentMD LatentMD Benchmarking Markdown Boundary Failures in LLM-Generated Text — NeurIPS 2026 E&D Track submission. This repository hosts the dataset artifact for LatentMD: the 4,179 prompts that constitute the benchmark, ~37,000 reference responses from 9 frontier models, and illustrative output samples. The accompanying evaluation code (CLI, metric definitions, statistical tests) lives in a separate code repository on GitHub under MIT. What LatentMD measures LLM Markdown… See the full description on the dataset page: https://huggingface.co/datasets/latentmd-neurips26/LatentMD.texttext-generation10K<n<100K0 likes82 downloads5mo agoHugging Face28elenagroundwork /llm-api-pricing-latency-2026 LLM Inference Unit Economics & Architecture Engine Empirical benchmark dataset by Groundwork Research (https://gworky.com). Full interactive decision engine available at: https://gworky.com/tools/llm-token-cost-calculator. Description Full-stack inference cost and latency estimator comparing frontier proprietary models (Claude 3.7, GPT-4.5) against open-weight hosted providers (Groq, DeepSeek R1, Together AI). Primary source authority: https://gworky.com/tech tabulartabular-classificationn<1K0 likes62 downloads21d agoHugging Face29DebdipCS /Latent-Resonance-AI-Image-Forensics-Benchmark-N100 Latent Resonance: SOTA Empirical AI Image Forensics Benchmark (N=100 & N=1,000 Scale) Author: Debdip Bandyopadhyay (Independent AI Researcher, Kolkata, India; M.Tech, IIT Jodhpur, AI & Data Science)Preprint & Paper: Latent Resonance: Zero-Shot Autoencoder Inversion and Azimuthal Spectral Forensics for Diffusion Image Attribution (IEEE Flagship / CERN Zenodo 2026) Benchmark Overview This repository provides: The official verified $N=100$ ground-truth image… See the full description on the dataset page: https://huggingface.co/datasets/DebdipCS/Latent-Resonance-AI-Image-Forensics-Benchmark-N100.imageimage-classificationn<1K0 likes60 downloads11d agoHugging Face30liaialley /latent-data Latent-SFT open-domain data This repository is the portable data bundle for LiAi16/latent-sft. It keeps the raw snapshots, the first-generation OSS-COT archive, the second-generation GLM-COT data used for formal training, a deterministic SFT mixture, complete raw evaluation benchmarks, normalized evaluation trajectories, and the partial Qwen3-4B SuperGPQA baseline used for exact resumption. Canonical Hub repository: liaialley/latent-data (the supplied token belongs to the… See the full description on the dataset page: https://huggingface.co/datasets/liaialley/latent-data.tabulartext-generation10K<n<100K0 likes59 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.