datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
megawika-report-generation
Dataset Card for MegaWika for Report Generation
Dataset Summary
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span
50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a
non-English language, an automated English translation is provided.
This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.corral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.financial-reports
Description
Topic: Financial Reports
Domains: Finance, Accounting, Economics
Focus: Synthetic raw financial reports for analysis and training
Number of Entries: 1000
Dataset Type: Raw Dataset
Model Used: bedrock/us.amazon.nova-pro-v1:0
Language: English
Generated by: SynthGenAI Package
corral-QAs-reports
Corral – QA Reports
Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.McKinsey-Reportsmeta-llama/synthetic-data-kit
https://github.com/meta-llama/synthetic-data-kit
McKinsey reports
https://www.mckinsey.com/featured-insights/insights-store
Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at:
https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
📄 100 Samples of Synthetic Automated Penetration Test Reports
This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including:
Reconnaissance / Discovery Phase
Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.OMB-Circular-A-136-Financial-Reporting-Requirements
OMB Circular A-136 Financial Reporting Requirement
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains 250 document-grounded question-and-answer records based on the August 23, 2005 revision of OMB Circular A-136, Financial Reporting Requirements.
The Circular consolidated and updated Office of Management and Budget guidance governing federal agency financial statements, interim statements, Performance and… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A-136-Financial-Reporting-Requirements.ReportBench
ReportBench — Dataset Card
Overview
ReportBench is a comprehensive benchmark for evaluating the factual quality and citation behavior of Deep Research agents. Leveraging expert-authored survey papers as ground truth, ReportBench reverse-engineers domain-specific prompts and provides automated tools to assess both cited and non-cited content.
[GitHub] / [Paper]
ReportBench addresses this need by:
Leveraging expert surveys: Uses high-quality, peer-reviewed survey papers… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-BandAI/ReportBench.sustainability-report-emissions-instruction-styleThe sustainability-report-emissions dataset converted into instruction-style JSONL format for direct consumption by SFTTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The dataset generation scripts are at this GitHub repo. An… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-instruction-style.oasst1-contamination-report
Contamination Report — OpenAssistant/oasst1
What this is
A row-level audit of OpenAssistant/oasst1 (revision
fdf72ae0827c1cda404aff25b6603abec9e3399b) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/oasst1-contamination-report.Chameleon-Radiology-Reportspubmed_case_reports
PubMed Case Reports
A collection of 13,989 full-text case reports from the PubMed Central (PMC) Open Access subset, spanning 2005–2025. Each article includes structured metadata, abstract, full body text, and section-level annotations. This dataset is designed for medical NLP, clinical reasoning, and biomedical text mining.
Dataset Description
Summary
This dataset comprises case reports published in peer-reviewed medical journals, sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/awinml/pubmed_case_reports.high-quality-english-sentences-contamination-report
Contamination Report — agentlans/high-quality-english-sentences
What this is
A row-level audit of agentlans/high-quality-english-sentences (revision
main) for exact 13-gram overlap with standard benchmark test sets
(gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new
artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone
training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-contamination-report.french-senate-session-reports
🏛️ French Senate Session Reports Dataset
A dataset of parliamentary debates and sessions reports from the French Senate.508,647,861 tokens of high-quality French text transcribed manually from Senate Sessions
Description
This dataset consists of all session reports from the French Senate debates, crawled from the official website senat.fr. It provides high-quality text data of parliamentary discussions, covering a wide range of political, economic, and social topics… See the full description on the dataset page: https://huggingface.co/datasets/TheJeanneCompany/french-senate-session-reports.financial-reports-extractive-summarization_eval
Financial Reports Extractive Summarization Evaluation Dataset
Validation and test splits for evaluating models on Arabic financial reports extractive summarization.
Dataset Structure
Format: Simple prompt-answer pairs
Validation: ~20 examples (10%)
Test: ~20 examples (10%)
Language: Arabic
Domain: Financial reports and market news
Fields
id: Unique identifier
prompt: The summarization prompt
full_text: Complete financial report
answer: Ground… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_eval.sustainability-report-emissions-dpoThe sustainability-report-emissions dataset converted into preferences-style JSONL format for DPO training. It can be directly used by DPOTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The chosen output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The rejected output is randomly… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-dpo.corral-intervention-reports
Corral – Intervention Ablation Reports
Intervention ablation reports for multiple LLM agents across all 8 Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the final evaluation reports of the intervention ablation study conducted across multiple LLM agents and all 8 Corral environments.
The intervention ablation examines… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports.TCGA_Reports_ja_structured_qwen38_27b
TCGA Reports Japanese Structured Dataset with Qwen3.8-27B
The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。
既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。
日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。
元データ
本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。
TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.traffic-accidents-reports-kd-smollm2-360M-7k
Accident Reporting KD Dataset (One-Paragraph)
Short description.A training/evaluation dataset for generating one-paragraph accident/incident reports from structured facts.This dataset mixes gold human targets from zBotta/traffic-accidents-reports-5k with teacher-generated soft targets produced by the model zBotta/smollm2-accident-reporter-360m to support knowledge distillation (KD) of a smaller student.
Output style: a single paragraph, neutral tone, covering What, When, Where, Who… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-kd-smollm2-360M-7k.traffic-accidents-reports-5k
Accident Reports 5k (One-Paragraph)
Rows: 5kSplits: train (~ 4500), eval (~ 500), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report)
Schema
input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions
target (string): One single-paragraph, neutral report (no line breaks)
file type: parquet dataset (training ready)
Languages
English
License
MIT
Intended use
Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-5k.clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1
What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty.
Quad coupling nodes
Emerging safety event pattern
Reporting or entry latency
Concomitant medication or behavior confound
Governance decision timing such as DSMB, batch release, or safety review
Input
One vignette
OutputReturn strict JSON only.
Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.ReportBench-Multilingual
ReportBench Multilingual
This dataset adds prompt-level multilingual translations to the 100 benchmark prompts in ByteDance-BandAI/ReportBench.
The translated prompt set covers eight languages:
en
zh
es
it
ar
bn
ja
el
What is included
The default all configuration exposes the full benchmark rows as a parquet file and preserves the upstream benchmark fields, including ground_truth and metadata fields such as title, abstract, authors, arxiv_id, and application_domain.… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/ReportBench-Multilingual.traffic-accidents-reports-800
Accident Reports 800 (One-Paragraph)
Rows: 800Splits: train (~ 630), eval (~ 70), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report)
Schema
input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions
target (string): One single-paragraph, neutral report (no line breaks)
file type: parquet dataset (training ready)
Languages
English
License
MIT
Intended use
Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-800.geo-research-report-2026
Moonify GEO Research Report 2026 (Agosto)
Descrizione del dataset
Questo dataset contiene la versione strutturata in 66 record del "Moonify GEO Research Report 2026 – Agosto", il documento di ricerca interno pubblicato da Moonify S.r.l. (ID documento MNF-GEO-2026-001) sullo stato della Generative Engine Optimization (GEO).
Il report copre il periodo osservato dicembre 2022 – agosto 2026 e raccoglie esperimenti, scoperte proprietarie, principi teorici, un framework… See the full description on the dataset page: https://huggingface.co/datasets/lunocode/geo-research-report-2026.medical-reports-simplification-dataset
🏥 Medical Reports Simplification Dataset
📋 Description
Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients.
🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public.
🔧 Génération du Dataset
Génération : Gemini 2.5 Pro (Preview)
Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.financial-reports-extractive-summarization_train
Financial Reports Extractive Summarization Training Dataset
Training split of the Arabic financial reports extractive summarization dataset in conversational format.
Dataset Structure
Format: Conversational (human-agent pairs)
Size: ~160 training examples (80% of total)
Language: Arabic
Domain: Financial reports and market news
Features
id: Unique identifier
conversations: Human prompt and agent summary
report_type: Type of financial report… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_train.protocols_and_reportstw-sinica-report
Dataset Card for tw-sinica-report
本資料集收集中華民國台灣中央研究院的研究報告,旨在協助語言模型學習繁體中文的專業領域知識。
Dataset Details
Dataset Description
本資料集彙整自中華民國臺灣之中央研究機構──中央研究院所發表之各類研究報告與學術成果,涵蓋人文、社會、自然與應用科學等多元領域。資料內容均以繁體中文撰寫,具有高度語言規範性與領域專業性,適合作為大型語言模型進行繁體中文專業知識建構與語言表達訓練之基礎素材。
本資料集之編撰目標在於補強語言模型於繁體中文專業文本之理解與生成能力,特別聚焦於提升模型處理學術用語、論述邏輯與正式文體之表現。透過納入中央研究院具權威性的研究成果,期能協助模型獲得更深層的語義推理能力與本地化知識背景,並進一步應用於教育、研究、科技應用等場域。
如需進一步瞭解資料來源,請參閱中央研究院官方網站。
Curated by: Huang Liang Hsun、Min Yi Chen & Wei Hao Lu… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-sinica-report.traffic-accidents-reports-5k
Accident Reports 5k (One-Paragraph)
Rows: 5kSplits: train (~ 4500), eval (~ 500), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report)
Schema
input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions
target (string): One single-paragraph, neutral report (no line breaks)
file type: parquet dataset (training ready)
Languages
English
License
MIT
Intended use
Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/dhana1012/traffic-accidents-reports-5k.krx_report
