datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
venra
VeNRA: Financial Hallucination Detection Dataset
VeNRA (Verification & Reasoning Audit) is a specialized dataset designed to train "Judge" models to detect hallucinations in Financial RAG (Retrieval-Augmented Generation) systems.
Unlike datasets that rely on "generative hallucinations" (asking an LLM to invent errors), VeNRA adopts an Adversarial Simulation philosophy. We scientifically reconstruct the specific cognitive failures that RAG systems exhibit in production by applying… See the full description on the dataset page: https://huggingface.co/datasets/pagand/venra.landing-page-training-data
Landing Page Training Data
Synthetic training data for fine-tuning LLMs to generate HTML landing pages. Generated using DeepSeek V3 (685B parameters).
Dataset Structure
data_sm/ # Small dataset
├── train.jsonl # 50 examples
├── valid.jsonl # 5 examples
└── test.jsonl # 5 examples
data_md/ # Medium dataset
├── train.jsonl # 500 examples
├── valid.jsonl # 5 examples
└── test.jsonl # 5 examples
Format
Each line is a JSON object in… See the full description on the dataset page: https://huggingface.co/datasets/KalnRangelov/landing-page-training-data.Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1
Dataset Card for Volmarr's Norse Paganism Fine-Tuning Dataset v1
A comprehensive JSONL dataset of approximately 1000 high-quality training pairs designed for fine-tuning large language models on authentic Norse Paganism (Ásatrú/Heathenry) topics. Each pair features user queries about key concepts—such as introduction to Norse Paganism, cosmology, deities, creation myths, Ragnarok, religious practices, runes, sacred sites, and more—paired with detailed, lore-accurate responses in the… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v1.Wikipedia-Pagesrepro-towards-optimal-robustness-in-learning-augmented-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
tvtropes.self.demonstrating.character.pages
TvTropes Self-Demonstrating Character Pages
Dataset Summary
Semi-cleaned dataset with all 'in-character' pages from TvTropes.
Inherits cc-by-nc-sa-3.0 license from tvtropes.
Languages
English
Dataset Structure
Jsonl format. {"text": "EXAMPLE OF TEXT"}
Dataset Creation
Content of all visible tags from the links on https://tvtropes.org/pmwiki/pmwiki.php/SelfDemonstrating/CharacterPages that contain SelfDemonstrating taken.
Cleaned by… See the full description on the dataset page: https://huggingface.co/datasets/MDGraff/tvtropes.self.demonstrating.character.pages.pageguide_hide_dataLink to the project: https://pageguide.github.io/
Link to the paper: https://huggingface.co/papers/2604.23772
Link to the code: https://github.com/tin-xai/pageguide
repro-la-paging-traces
Agent traces
Agent sessions published from a Trackio Logbook.
atlas-pages
Atlas Pages
Atlas Pages is a synthetic instruction dataset of ~7,000 expert-level concept explanations, generated by Claude Haiku and curated for fine-tuning small language models into precise, warm, human-friendly explainers.
It is the training backbone of Pocket Atlas — a fine-tuned Qwen3.5 model that explains any idea clearly, concisely, and with genuine warmth.
What's inside
Each example teaches a model to explain a concept using a strict 5-part structure:
What… See the full description on the dataset page: https://huggingface.co/datasets/cetusian/atlas-pages.landingboost-landing-page-trust-benchmark
LandingBoost Landing Page Trust Bottleneck Benchmark
LandingBoost is an AI landing page audit tool for SaaS founders. It reviews clarity, relevance, trust, CTA strength, proof, page order, and conversion friction, then recommends one prioritized first edit before a redesign, paid traffic, or an A/B test.
This public dataset contains privacy-safe aggregate results from a June 25, 2026 LandingBoost export. It does not publish customer identities, URLs, screenshots, page copy, or… See the full description on the dataset page: https://huggingface.co/datasets/yusuke0714/landingboost-landing-page-trust-benchmark.diataxis-pagestrolldom_and_magick_practices_in_norse_paganism_volume1
Trolldom and Magick Practices in Norse Paganism - Volume 1
This dataset consists of synthetic conversational dialogues inspired by Norse paganism, mythology, and ancient magick practices (trolldom and seidhr). It simulates exchanges between seekers of wisdom (users) and knowledgeable seers or völvas (assistants), discussing topics such as rune magic, galdr (chanted spells), offerings to gods like Odin, Freyja, Thor, and Frigg, wards against spirits, omens, cleansing rituals, and… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/trolldom_and_magick_practices_in_norse_paganism_volume1.squad-v2-reference-taskhumanizer_v2indexed_pagesThis is a growing dataset of indexed pages, made for anyone trying to build a mini-search engine locally.
Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v2
Dataset Card for Volmarr's Norse Paganism Fine-Tuning Dataset v2
A comprehensive JSONL dataset of approximately 1000 high-quality training pairs designed for fine-tuning large language models on authentic Norse Paganism (Ásatrú/Heathenry) topics. Each pair features user queries about key concepts—such as introduction to Norse Paganism, cosmology, deities, creation myths, Ragnarok, religious practices, runes, sacred sites, and more—paired with detailed, lore-accurate responses in the… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/Volmarrs_Norse_Paganism_Fine-Tuning_Dataset_v2.4-page-Plus-talespinsquad-v2-copy-taskvietnamese-financial-report-page-classifier
Vietnamese Financial Report Page Classifier Dataset
Dataset này dùng để huấn luyện mô hình phân loại từng trang trong báo cáo tài chính doanh nghiệp Việt Nam đã được OCR/convert sang Markdown.
Dataset Format
Mỗi dòng là một JSON object:
{"text": "...", "label": "balance_sheet", "important": 3}
Fields
text: nội dung một trang báo cáo tài chính ở dạng Markdown/OCR text.
label: loại trang.
important: mức độ quan trọng của trang.
Labels
cover
toc
company_information… See the full description on the dataset page: https://huggingface.co/datasets/Zenng2812/vietnamese-financial-report-page-classifier.medtutor_embed_page_exampleMozila-Common-Voice-page-Emakhuwa-localizationchinese_sentence_segmentation_page
