CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hltcoe /megawika-report-generation Dataset Card for MegaWika for Report Generation Dataset Summary MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.textsummarization100K<n<1M6 likes859 downloads3y agoHugging Face02jablonkagroup /corral_runs_reports Corral – Evaluation Score Reports Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments. The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.tabulartext-generationn<1K0 likes543 downloads3mo agoHugging Face03Shekswess /financial-reports Description Topic: Financial Reports Domains: Finance, Accounting, Economics Focus: Synthetic raw financial reports for analysis and training Number of Entries: 1000 Dataset Type: Raw Dataset Model Used: bedrock/us.amazon.nova-pro-v1:0 Language: English Generated by: SynthGenAI Package texttext-generation1K<n<10K1 likes218 downloads1y agoHugging Face04jablonkagroup /corral-QAs-reports Corral – QA Reports Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.texttext-generation1K<n<10K0 likes213 downloads3mo agoHugging Face05OrcinusOrca /McKinsey-Reportsmeta-llama/synthetic-data-kit https://github.com/meta-llama/synthetic-data-kit McKinsey reports https://www.mckinsey.com/featured-insights/insights-store texttext-generation10K<n<100K0 likes183 downloads1y agoHugging Face06CJJones /Synthetic_PenTest_ReportsThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com Want more? 🚀 Get the AI Startup Bundle from Gumroad. 📄 100 Samples of Synthetic Automated Penetration Test Reports This dataset contains 100+ realistic, synthetic penetration testing reportsstructured to simulate professional internal security assessments. Each record models the full flow of a pentest engagement, including: Reconnaissance / Discovery Phase Vulnerability Assessment… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_PenTest_Reports.texttext-classification10K<n<100K3 likes83 downloads7mo agoHugging Face07leeroy-jankins /OMB-Circular-A-136-Financial-Reporting-Requirements OMB Circular A-136 Financial Reporting Requirement Maintainer: Terry Eppler Owner: US Federal Government Dataset Summary This dataset contains 250 document-grounded question-and-answer records based on the August 23, 2005 revision of OMB Circular A-136, Financial Reporting Requirements. The Circular consolidated and updated Office of Management and Budget guidance governing federal agency financial statements, interim statements, Performance and… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/OMB-Circular-A-136-Financial-Reporting-Requirements.documentquestion-answering0 likes60 downloads2mo agoHugging Face08ByteDance-BandAI /ReportBench ReportBench — Dataset Card Overview ReportBench is a comprehensive benchmark for evaluating the factual quality and citation behavior of Deep Research agents. Leveraging expert-authored survey papers as ground truth, ReportBench reverse-engineers domain-specific prompts and provides automated tools to assess both cited and non-cited content. [GitHub] / [Paper] ReportBench addresses this need by: Leveraging expert surveys: Uses high-quality, peer-reviewed survey papers… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance-BandAI/ReportBench.texttext-generationn<1K0 likes59 downloads1y agoHugging Face09nopperl /sustainability-report-emissions-instruction-styleThe sustainability-report-emissions dataset converted into instruction-style JSONL format for direct consumption by SFTTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The dataset generation scripts are at this GitHub repo. An… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-instruction-style.texttext-generation1K<n<10K1 likes55 downloads3y agoHugging Face10liodon-ai /oasst1-contamination-report Contamination Report — OpenAssistant/oasst1 What this is A row-level audit of OpenAssistant/oasst1 (revision fdf72ae0827c1cda404aff25b6603abec9e3399b) for exact 13-gram overlap with standard benchmark test sets (gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/oasst1-contamination-report.texttext-generationn<1K0 likes53 downloads3d agoHugging Face11satvikt04 /Chameleon-Radiology-Reportstexttext-classification10K<n<100K7 likes48 downloads7mo agoHugging Face12awinml /pubmed_case_reports PubMed Case Reports A collection of 13,989 full-text case reports from the PubMed Central (PMC) Open Access subset, spanning 2005–2025. Each article includes structured metadata, abstract, full body text, and section-level annotations. This dataset is designed for medical NLP, clinical reasoning, and biomedical text mining. Dataset Description Summary This dataset comprises case reports published in peer-reviewed medical journals, sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/awinml/pubmed_case_reports.texttext-generation10K<n<100K0 likes45 downloads3mo agoHugging Face13liodon-ai /high-quality-english-sentences-contamination-report Contamination Report — agentlans/high-quality-english-sentences What this is A row-level audit of agentlans/high-quality-english-sentences (revision main) for exact 13-gram overlap with standard benchmark test sets (gsm8k, hellaswag, humaneval, mmlu). This is not a filtered copy of the source — it's a new artifact: a list of which rows overlap which benchmark, plus summary statistics, so anyone training on the source can decide how to handle it.… See the full description on the dataset page: https://huggingface.co/datasets/liodon-ai/high-quality-english-sentences-contamination-report.texttext-generationn<1K0 likes41 downloads18d agoHugging Face14TheJeanneCompany /french-senate-session-reports 🏛️ French Senate Session Reports Dataset A dataset of parliamentary debates and sessions reports from the French Senate.508,647,861 tokens of high-quality French text transcribed manually from Senate Sessions Description This dataset consists of all session reports from the French Senate debates, crawled from the official website senat.fr. It provides high-quality text data of parliamentary discussions, covering a wide range of political, economic, and social topics… See the full description on the dataset page: https://huggingface.co/datasets/TheJeanneCompany/french-senate-session-reports.texttext-classification1K<n<10K4 likes38 downloads2y agoHugging Face15SahmBenchmark /financial-reports-extractive-summarization_eval Financial Reports Extractive Summarization Evaluation Dataset Validation and test splits for evaluating models on Arabic financial reports extractive summarization. Dataset Structure Format: Simple prompt-answer pairs Validation: ~20 examples (10%) Test: ~20 examples (10%) Language: Arabic Domain: Financial reports and market news Fields id: Unique identifier prompt: The summarization prompt full_text: Complete financial report answer: Ground… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_eval.tabularsummarizationn<1K0 likes37 downloads9mo agoHugging Face16nopperl /sustainability-report-emissions-dpoThe sustainability-report-emissions dataset converted into preferences-style JSONL format for DPO training. It can be directly used by DPOTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The chosen output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The rejected output is randomly… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-dpo.texttext-generation1K<n<10K1 likes31 downloads3y agoHugging Face17jablonkagroup /corral-intervention-reports Corral – Intervention Ablation Reports Intervention ablation reports for multiple LLM agents across all 8 Corral environments 📋 Dataset Summary This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the final evaluation reports of the intervention ablation study conducted across multiple LLM agents and all 8 Corral environments. The intervention ablation examines… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-intervention-reports.tabulartext-generationn<1K0 likes31 downloads3mo agoHugging Face18morizon /TCGA_Reports_ja_structured_qwen38_27b TCGA Reports Japanese Structured Dataset with Qwen3.8-27B The Cancer Genome Atlas(TCGA)由来の英語病理報告書を日本語へ翻訳し、その日本語病理報告書から主要な病理情報を9項目へ構造化したデータセットです。 既存の morizon/TCGA_Reports_ja_structured と同じ100症例を使用し、生成モデルを Qwen/Qwen3.8-27B に変更して再生成しています。 日本語訳には morizon/TCGA_Reports_ja_qwen38_27b と同じ生成結果を使用しています。 元データ 本データセットでは、The Cancer Genome Atlas(TCGA)の病理報告書をもとに作成されたTCGA-Reportsを使用しています。 TCGAは、複数のがん種についてゲノム情報や臨床情報などを収集した大規模ながん研究プロジェクトです。… See the full description on the dataset page: https://huggingface.co/datasets/morizon/TCGA_Reports_ja_structured_qwen38_27b.texttext-generationn<1K0 likes30 downloads1mo agoHugging Face19DSTI /traffic-accidents-reports-kd-smollm2-360M-7k Accident Reporting KD Dataset (One-Paragraph) Short description.A training/evaluation dataset for generating one-paragraph accident/incident reports from structured facts.This dataset mixes gold human targets from zBotta/traffic-accidents-reports-5k with teacher-generated soft targets produced by the model zBotta/smollm2-accident-reporter-360m to support knowledge distillation (KD) of a smaller student. Output style: a single paragraph, neutral tone, covering What, When, Where, Who… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-kd-smollm2-360M-7k.texttext-generation1K<n<10K1 likes29 downloads1y agoHugging Face20DSTI /traffic-accidents-reports-5k Accident Reports 5k (One-Paragraph) Rows: 5kSplits: train (~ 4500), eval (~ 500), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report) Schema input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions target (string): One single-paragraph, neutral report (no line breaks) file type: parquet dataset (training ready) Languages English License MIT Intended use Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-5k.texttext-generation1K<n<10K0 likes28 downloads1y agoHugging Face21ClarusC64 /clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1Clarus Clinical Quad Coupling Safety Signal Latency Reporting Lag Conmed Confound v0.1 What this dataset isThis dataset tests whether a model can detect latent safety signals when four interacting nodes create uncertainty. Quad coupling nodes Emerging safety event pattern Reporting or entry latency Concomitant medication or behavior confound Governance decision timing such as DSMB, batch release, or safety review Input One vignette OutputReturn strict JSON only. Required output… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-quad-safety-signal-latency-reporting-lag-conmed-confound-v0.1.texttext-generationn<1K0 likes28 downloads7mo agoHugging Face22JRQi /ReportBench-Multilingual ReportBench Multilingual This dataset adds prompt-level multilingual translations to the 100 benchmark prompts in ByteDance-BandAI/ReportBench. The translated prompt set covers eight languages: en zh es it ar bn ja el What is included The default all configuration exposes the full benchmark rows as a parquet file and preserves the upstream benchmark fields, including ground_truth and metadata fields such as title, abstract, authors, arxiv_id, and application_domain.… See the full description on the dataset page: https://huggingface.co/datasets/JRQi/ReportBench-Multilingual.texttext-generationn<1K0 likes27 downloads6mo agoHugging Face23DSTI /traffic-accidents-reports-800 Accident Reports 800 (One-Paragraph) Rows: 800Splits: train (~ 630), eval (~ 70), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report) Schema input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions target (string): One single-paragraph, neutral report (no line breaks) file type: parquet dataset (training ready) Languages English License MIT Intended use Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/DSTI/traffic-accidents-reports-800.texttext-generationn<1K1 likes25 downloads1y agoHugging Face24lunocode /geo-research-report-2026 Moonify GEO Research Report 2026 (Agosto) Descrizione del dataset Questo dataset contiene la versione strutturata in 66 record del "Moonify GEO Research Report 2026 – Agosto", il documento di ricerca interno pubblicato da Moonify S.r.l. (ID documento MNF-GEO-2026-001) sullo stato della Generative Engine Optimization (GEO). Il report copre il periodo osservato dicembre 2022 – agosto 2026 e raccoglie esperimenti, scoperte proprietarie, principi teorici, un framework… See the full description on the dataset page: https://huggingface.co/datasets/lunocode/geo-research-report-2026.texttext-generationn<1K0 likes24 downloads1mo agoHugging Face25Sadou /medical-reports-simplification-dataset 🏥 Medical Reports Simplification Dataset 📋 Description Dataset créé avec Gemini 2.5 Pro (Preview) pour entraîner des modèles à simplifier les rapports médicaux complexes en explications compréhensibles pour les patients. 🎯 Objectif : Démocratiser l'accès à l'information médicale en rendant les rapports techniques accessibles au grand public. 🔧 Génération du Dataset Génération : Gemini 2.5 Pro (Preview) Validation : Contrôle qualité automatisé… See the full description on the dataset page: https://huggingface.co/datasets/Sadou/medical-reports-simplification-dataset.texttext-generationn<1K0 likes19 downloads1y agoHugging Face26SahmBenchmark /financial-reports-extractive-summarization_train Financial Reports Extractive Summarization Training Dataset Training split of the Arabic financial reports extractive summarization dataset in conversational format. Dataset Structure Format: Conversational (human-agent pairs) Size: ~160 training examples (80% of total) Language: Arabic Domain: Financial reports and market news Features id: Unique identifier conversations: Human prompt and agent summary report_type: Type of financial report… See the full description on the dataset page: https://huggingface.co/datasets/SahmBenchmark/financial-reports-extractive-summarization_train.tabularsummarizationn<1K0 likes17 downloads9mo agoHugging Face27TristanGraeble /protocols_and_reportstexttext-generationn<1K0 likes12 downloads2y agoHugging Face28lianghsun /tw-sinica-reportgated Dataset Card for tw-sinica-report 本資料集收集中華民國台灣中央研究院的研究報告,旨在協助語言模型學習繁體中文的專業領域知識。 Dataset Details Dataset Description 本資料集彙整自中華民國臺灣之中央研究機構──中央研究院所發表之各類研究報告與學術成果,涵蓋人文、社會、自然與應用科學等多元領域。資料內容均以繁體中文撰寫,具有高度語言規範性與領域專業性,適合作為大型語言模型進行繁體中文專業知識建構與語言表達訓練之基礎素材。 本資料集之編撰目標在於補強語言模型於繁體中文專業文本之理解與生成能力,特別聚焦於提升模型處理學術用語、論述邏輯與正式文體之表現。透過納入中央研究院具權威性的研究成果,期能協助模型獲得更深層的語義推理能力與本地化知識背景,並進一步應用於教育、研究、科技應用等場域。 如需進一步瞭解資料來源,請參閱中央研究院官方網站。 Curated by: Huang Liang Hsun、Min Yi Chen & Wei Hao Lu… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-sinica-report.tabulartext-generationn<1K0 likes6 downloads1y agoHugging Face29dhana1012 /traffic-accidents-reports-5k Accident Reports 5k (One-Paragraph) Rows: 5kSplits: train (~ 4500), eval (~ 500), test (~ 100)Task: text2text-generation (5W1H → single-paragraph incident report) Schema input (string): 5W1H lines: What, When, Where, Who, How, Why, ContingencyActions target (string): One single-paragraph, neutral report (no line breaks) file type: parquet dataset (training ready) Languages English License MIT Intended use Fine-tuning small… See the full description on the dataset page: https://huggingface.co/datasets/dhana1012/traffic-accidents-reports-5k.texttext-generation1K<n<10K0 likes6 downloads8mo agoHugging Face30BigShort /krx_reportgatedtexttext-generation1K<n<10K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.