CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Pageshift-Entertainment /LongPage Overview 🚀📚 The first comprehensive dataset for training AI models to write complete novels with sophisticated reasoning. 🧠 Hierarchical Reasoning Architecture — Multi-layered planning traces including character archetypes, story arcs, world rules, and scene breakdowns. A complete cognitive roadmap for long-form narrative construction. 📖 Complete Novel Coverage — From 40,000 to 600,000+ tokens per book, spanning novellas to epic series with consistent quality throughout. ⚡… See the full description on the dataset page: https://huggingface.co/datasets/Pageshift-Entertainment/LongPage.texttext-generation1K<n<10K154 likes599 downloads8mo agoHugging Face02pagarsky /agent-trace AgentTrace AgentTrace is an open dataset of tool-using language-model agent traces with execution telemetry. Each trace records model-generation steps, tool calls, wall-clock timing, OS-level resource usage, tool inputs and outputs, reasoning content, and reproducibility metadata. The repository contains the dataset, collection code, analysis scripts, and the deterministic NL2Bash fixture needed to replay the local command-line tasks. Links GitHub:… See the full description on the dataset page: https://huggingface.co/datasets/pagarsky/agent-trace.tabulartext-generation1K<n<10K0 likes159 downloads5mo agoHugging Face03KalnRangelov /landing-page-training-data Landing Page Training Data Synthetic training data for fine-tuning LLMs to generate HTML landing pages. Generated using DeepSeek V3 (685B parameters). Dataset Structure data_sm/ # Small dataset ├── train.jsonl # 50 examples ├── valid.jsonl # 5 examples └── test.jsonl # 5 examples data_md/ # Medium dataset ├── train.jsonl # 500 examples ├── valid.jsonl # 5 examples └── test.jsonl # 5 examples Format Each line is a JSON object in… See the full description on the dataset page: https://huggingface.co/datasets/KalnRangelov/landing-page-training-data.texttext-generationn<1K1 likes122 downloads7mo agoHugging Face04PiotrSty /plwiki-talk-pages Polish Wikipedia talk pages (plwiki, namespaces 1+5) Article-talk and project-talk pages from the pinned plwiki-20260901-pages-meta-current dump (https://dumps.wikimedia.org/plwiki/20260901/). The source selects namespaces 1 and 5; zero collisions of source-prefixed record IDs do not establish disjointness from the target wikipedia shard. Target-wide text deduplication remains pending. This is conversational written Polish: editorial disputes, coordination, questions and… See the full description on the dataset page: https://huggingface.co/datasets/PiotrSty/plwiki-talk-pages.texttext-generation10K<n<100K0 likes98 downloads15d agoHugging Face05freococo /thahabiorg_pages ❤️ Credits Original Source All original books, bibliographic metadata, and online library organization originate from: Thahabi Shamela Books Team https://thahabi.org All intellectual credit for the collection belongs to the Thahabi Shamela Books Team, the respective authors, editors, publishers, and other rights holders. Dataset Creation The page extraction, large-scale scraping pipeline, data cleaning, deduplication, validation, Apache Parquet… See the full description on the dataset page: https://huggingface.co/datasets/freococo/thahabiorg_pages.texttext-generation10M<n<100M0 likes65 downloads3mo agoHugging Face06ctoraman /front-page-news(This dataset contains raw text, which are unlabeled.) 5,844 Turkish news articles obtained from the Milliyet newspaper between 9 September 2009 and 31 October 2009 Github Repo: https://github.com/BilkentInformationRetrievalGroup/BilFront2009 If you would like to use any material in this repository, please cite the following paper: Toraman, C., & Can, F. (2015). A front-page news-selection algorithm based on topic modelling using raw text. Journal of Information Science, 41(5)… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/front-page-news.texttext-generationn<1K0 likes63 downloads3y agoHugging Face07pageman /virginia-woolf-monologue-chunks Virginia Woolf Monologue Chunks Dataset This dataset contains 6 semantically chunked text segments derived from a contemporary monologue based on Virginia Woolf's seminal essay "A Room of One's Own" (1929). It comes pre-loaded with vector embeddings from three different models, making it a ready-to-use resource for a variety of NLP tasks. In addition to the dataset itself, this repository includes a comprehensive embedding analysis, detailed statistics, and 7 visualizations to help… See the full description on the dataset page: https://huggingface.co/datasets/pageman/virginia-woolf-monologue-chunks.tabulartext-generationn<1K0 likes56 downloads11mo agoHugging Face08pagantibet /normalisation-S2S-training Tibetan Normalisation - S2S Training Data A large-scale parallel training dataset for Tibetan text normalisation, containing approximately 2 million line pairs mapping diplomatic (non-standard, abbreviated) Tibetan manuscript text to Standard Classical Tibetan. This dataset was used to train the sequence-to-sequence normalisation models (tokenised S2S model and non-tokenised S2S model) released as part of the PaganTibet project. The dataset combines a manually curated gold-standard… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/normalisation-S2S-training.texttext-generation1M<n<10M0 likes55 downloads6mo agoHugging Face09miracleyin /pdfsys-page-v2-demo pdfsys.page/v2 — 格式演示数据集 pdfsys.page/v2 是 pdfsystem_mnbvc 的 L2 发布格式,为 MNBVC 中文语料的 PB 级 PDF 流水线设计。 这是一个格式演示,不是训练语料。 25 页、18 份文档,只够说明 schema 长什么样、 三种视图怎么取。真实语料是 21.8 万份 PDF 的量级。 来源提示:这里的 PDF 页来自 OmniDocBench 与 olmOCR-bench 两个公开 benchmark,逐份的上游许可未经核实。放出来是为了说明数据格式,不是为了再分发这些 文档本身——要拿去用请自行确认源文档的许可。详见文末「来源与许可」。 一句话设计 一行一页,主键 (doc_id, page_index) ——这个身份来自 PDF 本身,不是模型造出来的; 页文本里内联图标记来承载图文交错;模型派生的结构是旁边一列可丢弃的增强; 图像像素要么是裁剪图、要么是整页光栅,二选一。 里面有什么 config 行数… See the full description on the dataset page: https://huggingface.co/datasets/miracleyin/pdfsys-page-v2-demo.tabularimage-to-textn<1K0 likes47 downloads29d agoHugging Face10pagantibet /Tibetan-KenLM-ACTibtrainingdata Tibetan Normalisation - KenLM ACTib Training Data A large-scale corpus of Standard Classical Tibetan text prepared specifically for training character-level KenLM n-gram language models for use in the PaganTibet normalisation pipeline. The dataset contains approximately 17.7 million lines of cleaned, line-split ACTib text in two versions: non-tokenised and tokenised (via a customised version of the Botok Tibetan tokeniser). These two files are the direct training corpora for… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-KenLM-ACTibtrainingdata.texttext-generation10M<n<100M0 likes43 downloads6mo agoHugging Face11cakradana-app /cakradana-kpu-filings-14k-pages 🏛️ Cakradana — KPU Campaign-Finance Filings Indonesian election campaign-finance filings, for document OCR and structured extraction 📋 Table of Contents 🎯 What this is 📦 What is in here 🚀 Loading 🗂️ Schema 🔒 Redaction ⚠️ Limitations 📜 Licence and provenance 🎯 What this is Indonesian candidates and parties must file campaign-finance reports — LADK, LPSDK and LPPDK — with the KPU (Komisi Pemilihan Umum, the General Elections… See the full description on the dataset page: https://huggingface.co/datasets/cakradana-app/cakradana-kpu-filings-14k-pages.image-to-text10K<n<100K0 likes38 downloads1mo agoHugging Face12atrost /dsv4-flash-tmax-git-pager-recovery-23 DeepSeek V4 Flash TMax Git Pager Recovery This dataset contains 23 reward-one SFT trajectories across 19 TMax tasks generated by DeepSeek-V4-Flash-0731. Every row was manually audited against the raw terminal recording and contains a real foreground Git pager/less interaction, an executed recovery action, shell-prompt restoration, and subsequent working shell use. Composition 6 original parser-clean full last-episode exports. 17 additional manually confirmed… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery-23.tabulartext-generationn<1K0 likes37 downloads2mo agoHugging Face13kogai /landing-pages-styling-css-only-v2v3-merged merged_v2_v3_dedup This dataset is the deduplicated merged training set built from the repository's landing_page_v2 and landing_page_v3 pipelines. It contains chat-format rows for CSS generation on landing pages: messages + metadata Each sample keeps: messages: the training conversation, usually a user prompt plus the assistant CSS response. metadata: generation and provenance fields such as source pipeline, style phrase, recipe, and other analysis attributes.… See the full description on the dataset page: https://huggingface.co/datasets/kogai/landing-pages-styling-css-only-v2v3-merged.texttext-generation10K<n<100K0 likes30 downloads3mo agoHugging Face14atrost /dsv4-flash-tmax-git-pager-recovery DeepSeek V4 Flash TMax Git Pager Recovery This dataset contains 6 manually audited, SFT-ready terminal-agent trajectories generated by DeepSeek-V4-Flash-0731 in public TMax environments. The primary subset is deliberately narrow: the agent must actually enter a Git pager or foreground TUI, execute a useful recovery action, return to a shell prompt, and finish the task with reward 1.0. Source and collection Environment/task source: TMaxxx/TMax-15K-Harbor, pinned… See the full description on the dataset page: https://huggingface.co/datasets/atrost/dsv4-flash-tmax-git-pager-recovery.tabulartext-generationn<1K0 likes28 downloads2mo agoHugging Face15pagantibet /Tibetan-normalisation-testdata Tibetan Normalisation - Test Data A collection of evaluation datasets for Classical Tibetan text normalisation, containing three distinct test sets designed to assess normalisation systems under different conditions: a manually curated gold-standard set of diplomatic manuscript text, and two synthetic sets of Standard Classical Tibetan text with OCR-based noise applied. Together these test sets allow evaluation across a spectrum from clean, realistic manuscript normalisation to more… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-normalisation-testdata.texttext-generation1K<n<10K0 likes27 downloads6mo agoHugging Face16cetusian /atlas-pages Atlas Pages Atlas Pages is a synthetic instruction dataset of ~7,000 expert-level concept explanations, generated by Claude Haiku and curated for fine-tuning small language models into precise, warm, human-friendly explainers. It is the training backbone of Pocket Atlas — a fine-tuned Qwen3.5 model that explains any idea clearly, concisely, and with genuine warmth. What's inside Each example teaches a model to explain a concept using a strict 5-part structure: What… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/atlas-pages.texttext-generation1K<n<10K0 likes23 downloads7mo agoHugging Face17pagantibet /Tibetan-abbreviation-dictionary Tibetan Normalisation - Abbreviation Dictionary A custom-built dictionary of ~10,400 Classical Tibetan abbreviation–expansion pairs, mapping abbreviated or contracted diplomatic forms to their Standard Classical Tibetan equivalents. This dictionary is used as the basis for the rule-based normalisation component of the PaganTibet inference pipeline, and also as a data augmentation resource during training. Abbreviations are one of the most systematic and frequent sources of deviation… See the full description on the dataset page: https://huggingface.co/datasets/pagantibet/Tibetan-abbreviation-dictionary.texttext-generation10K<n<100K0 likes14 downloads6mo agoHugging Face18EkBass /fiwiki-20250220-pages-articles Finnish Wikipedia Articles (20.02.2025) - JSON Dataset On my journey towards building my first small Finnish cognitive intelligence model, the first step was to parse the entire Finnish Wikipedia in a structured way. I believe the end result is quite solid. There are some inaccuracies in the dataset, but 95% of it is in good Finnish and structurally sound. Instructions for Extracting Finnish Wikipedia Follow these steps to convert Finnish Wikipedia into a… See the full description on the dataset page: https://huggingface.co/datasets/EkBass/fiwiki-20250220-pages-articles.texttext-generation100K<n<1M2 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.