CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hltcoe /megawika-report-generation Dataset Card for MegaWika for Report Generation Dataset Summary MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.textsummarization100K<n<1M6 likes861 downloads3y agoHugging Face02elihoole /asrs-aviation-reports Dataset Card for ASRS Aviation Incident Reports Dataset Summary This dataset collects 47,723 aviation incident reports published in the Aviation Safety Reporting System (ASRS) database maintained by NASA. Supported Tasks and Leaderboards 'summarization': Dataset can be used to train a model for abstractive and extractive summarization. The model performance is measured by how high the output summary's ROUGE score for a given narrative account of an aviation… See the full description on the dataset page: https://huggingface.co/datasets/elihoole/asrs-aviation-reports.textsummarization10K<n<100K11 likes338 downloads4y agoHugging Face03treychase /mlb-daily-reporttextn<1K0 likes327 downloads4d agoHugging Face04DataNeed /company-reports Company Reports Dataset Description This dataset contains ESG (Environmental, Social, and Governance) sustainability reports from various companies. It includes data like company details, report categories, textual analysis of the reports, and more. Dataset Structure id: Unique identifier for each report entry. document_category: Classification of the document (e.g., ESG sustainability report). year: Publication year of the report. company_name: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/DataNeed/company-reports.texttext-classification1K<n<10K7 likes262 downloads3y agoHugging Face05csoai /measured-vs-reported Measured vs reported — the empty table, on purpose An honesty artefact, and deliberately close to empty. overlap.json would map our measured Elo against third-party reported numbers, but its state reads "UNKNOWN — no verified cross-platform Elo for our fleet models yet (honest, not fabricated)", cells is [], and the gate is stated in the file: a reported cell is populated only when we hold a cited, attributed number for the same model we measured. Until that holds, the table… See the full description on the dataset page: https://huggingface.co/datasets/csoai/measured-vs-reported.textothern<1K0 likes198 downloads9d agoHugging Face06OrcinusOrca /McKinsey-Reportsmeta-llama/synthetic-data-kit https://github.com/meta-llama/synthetic-data-kit McKinsey reports https://www.mckinsey.com/featured-insights/insights-store texttext-generation10K<n<100K0 likes172 downloads1y agoHugging Face07dongbobo /annoy-datasync-license-reporttextn<1K0 likes98 downloads8mo agoHugging Face08Kira-Floris /gov-report-qs-llama2-format Government Report Question Answering Dataset in LLAMA2 Format Dataset Description This dataset is a LLAMA2 formatted dataset of the GovReport Dataset which is a report dataset, consisting of reports written by government research agencies including Congressional Research Service and US Government Accountability Office. The purpose of creating this dataset is to provide those trying to finetune LLAMA2 and other LLM models for Government domain a formatted and easier to use… See the full description on the dataset page: https://huggingface.co/datasets/Kira-Floris/gov-report-qs-llama2-format.textquestion-answering10K<n<100K2 likes84 downloads3y agoHugging Face09AriaAICompany /threat-intel-reports ThreatIntel synthetic reports 32 synthetic English and Persian CTI notes for the ThreatIntel extraction demo. Seed 5. Organization dataset and collection item are public. Live Gradio (AriaAICompany/threat-intel or alirezaaminzadeh/threat-intel) is created by scripts/publish.py after the daily Space-creation cap resets. This is fixture data (level 1). It does not prove operational extraction quality on real vendor reports. Reports are original laboratory text. They are not copies… See the full description on the dataset page: https://huggingface.co/datasets/AriaAICompany/threat-intel-reports.tabulartoken-classificationn<1K0 likes84 downloads2d agoHugging Face10patrickocal /gov_report_kgtext10K<n<100K1 likes77 downloads3y agoHugging Face11MongoDB /fake_tech_companies_market_reportstextn<1K0 likes71 downloads2y agoHugging Face12jsmarkschoon /10K_Report_Query_Tooltextn<1K0 likes67 downloads6mo agoHugging Face13ChrisRPL /satellite-civilian-conflict-disruption-reporter-v1 Satellite Civilian Conflict Disruption Reporter v1 Dataset ID: ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1 Status This is a valid diagnostic reporter-schema dataset, not the current Blackline Atlas canonical model gate. The canonical compact calibration/gold dataset remains ChrisRPL/satellite-disruption-triage-aux-v2-2. Use this dataset for future schema-simplification experiments only after respecting the mixed source licenses. Do not treat the associated… See the full description on the dataset page: https://huggingface.co/datasets/ChrisRPL/satellite-civilian-conflict-disruption-reporter-v1.imageimage-to-textn<1K0 likes66 downloads5mo agoHugging Face14findzebra /case-reports FindZebra case reports A collection of 3344 case reports fetched from the PubMed API for the Fabry, Gaucher and Familial amyloid cardiomyopathy (FAC) diseases. Articles are labelled using a text segmentation model described in "FindZebra online search delving into rare disease case reports using natural language processing". text1K<n<10K4 likes57 downloads3y agoHugging Face15baker-street /maib-incident-reports-5K MAIB Incident Type Dataset The MAIB Incident Type Dataset contains short textual descriptions of marine accidents and incidents reported by the UK Marine Accident Investigation Branch (MAIB).Each record includes a short narrative and a corresponding incident-type label (e.g. Grounding / Stranding, Fire / Explosion, Collision).This dataset enables research and experimentation in maritime safety text classification and domain-specific NLP. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/baker-street/maib-incident-reports-5K.texttext-classification1K<n<10K0 likes55 downloads11mo agoHugging Face16nopperl /sustainability-report-emissions-instruction-styleThe sustainability-report-emissions dataset converted into instruction-style JSONL format for direct consumption by SFTTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The dataset generation scripts are at this GitHub repo. An… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-instruction-style.texttext-generation1K<n<10K1 likes54 downloads3y agoHugging Face17Roy229 /fsfh6410-triage-report-0c136bn<1K0 likes51 downloads29d agoHugging Face18AdamLucek /apple-environmental-report-QA-retrieval Apple's 2024 Environmental Report QA Pairs 4300 question and relevant text chunks made from Apple's 2024 Environmental Report. Chunking was done with a token based recursive chunker at 800 token chunk size with a 400 token overlap resulting in 215 chunks. 20 question labels per chunk were synthetically generated using gpt-4o-mini with the attached prompt and a temperature of 1.0. Entries were shuffled and split into an 80/20 Train/Validation split resulting in:Training set size:… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/apple-environmental-report-QA-retrieval.textquestion-answering1K<n<10K0 likes50 downloads2y agoHugging Face19tgrex6 /mimic-cxr-reports-summarizationtext10K<n<100K4 likes46 downloads2y agoHugging Face20referencesource /chemical-regulatory-reporting-thresholds US federal chemical regulatory reporting thresholds by program (CERCLA, EPCRA, CAA) Canonical, always-current version: https://referencesource.org/chemical-regulatory-reporting-thresholds/ Machine-readable: https://referencesource.org/chemical-regulatory-reporting-thresholds/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-05 Stale after: 2027-08-05 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records:… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/chemical-regulatory-reporting-thresholds.text1K<n<10K0 likes46 downloads27d agoHugging Face21ZeroCommand /test-giskard-reporttabularn<1K0 likes43 downloads3y agoHugging Face22buttersx /bug-report-fixtexttext-classification1K<n<10K1 likes40 downloads7mo agoHugging Face23referencesource /epcra-tier-ii-state-reporting-thresholds EPCRA Tier II hazardous chemical inventory reporting thresholds by state Canonical, always-current version: https://referencesource.org/epcra-tier-ii-state-reporting-thresholds/ Machine-readable: https://referencesource.org/epcra-tier-ii-state-reporting-thresholds/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-18 Stale after: 2027-08-18 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 19 The… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/epcra-tier-ii-state-reporting-thresholds.textn<1K0 likes40 downloads27d agoHugging Face24sarahooker /global-news-reports-v1 This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. global_news_reports This dataset comprises news articles covering major international events, with a significant focus on natural disasters in Turkey and Syria, geopolitical tensions involving Russia and Iran, and human interest stories. The texts provide detailed accounts of casualties, political responses, and survivor experiences drawn from various global sources. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/global-news-reports-v1.textn<1K0 likes39 downloads6mo agoHugging Face25referencesource /notifiable-disease-reporting-timeframes US state notifiable disease reporting timeframes by condition Canonical, always-current version: https://referencesource.org/notifiable-disease-reporting-timeframes/ Machine-readable: https://referencesource.org/notifiable-disease-reporting-timeframes/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-16 Stale after: 2027-08-16 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 334 State-by-state… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/notifiable-disease-reporting-timeframes.textn<1K0 likes39 downloads27d agoHugging Face26dnamodel /xraydar-reports X-Raydar Annotated Radiology Reports Manually annotated chest X-ray radiology reports for multi-label classification and span-level segmentation. Each report is annotated with 45 radiological finding categories at both the report level and the token/span level. Website: x-raydar.info X-ray classifier: dnamodel/xraydar-cv Report classifier: dnamodel/xraydar-nlp CV code: gmontana/xraydar-cv NLP code: gmontana/xraydar-nlp This dataset was used to train and evaluate the RoBERTaX model… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/xraydar-reports.texttoken-classification10K<n<100K0 likes38 downloads6mo agoHugging Face27Joakimpalm-Zen /Qwen3-speculative-pair-report Qwen3 speculative pair report: 0.6B draft + 8B target, measured acceptance Research evidence dataset. No model weights. Part of the collection Xyntetik Research: Runner Compatibility Reports on this account, produced with Xyntetik Runner. Dataset summary Question tested. What the acceptance rate of a Qwen3-0.6B draft against a Qwen3-8B target actually is across draft depths, whether the engine's printed tokens-per-round figure can be tuned on (it cannot), and… See the full description on the dataset page: https://huggingface.co/datasets/Joakimpalm-Zen/Qwen3-speculative-pair-report.textn<1K0 likes38 downloads7d agoHugging Face28lunocode /geo-research-report-2026 Moonify GEO Research Report 2026 (Agosto) Descrizione del dataset Questo dataset contiene la versione strutturata in 66 record del "Moonify GEO Research Report 2026 – Agosto", il documento di ricerca interno pubblicato da Moonify S.r.l. (ID documento MNF-GEO-2026-001) sullo stato della Generative Engine Optimization (GEO). Il report copre il periodo osservato dicembre 2022 – agosto 2026 e raccoglie esperimenti, scoperte proprietarie, principi teorici, un framework… See the full description on the dataset page: https://huggingface.co/datasets/lunocode/geo-research-report-2026.texttext-generationn<1K0 likes36 downloads1mo agoHugging Face29sarahooker /global-news-reports This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform. global_news_reports This dataset comprises news articles covering major international events, with a significant focus on natural disasters in Turkey and Syria, geopolitical tensions involving Russia and Iran, and human interest stories. The texts provide detailed accounts of casualties, political responses, and survivor experiences drawn from various global sources. Each… See the full description on the dataset page: https://huggingface.co/datasets/sarahooker/global-news-reports.textn<1K0 likes35 downloads6mo agoHugging Face30nopperl /sustainability-report-emissions-dpoThe sustainability-report-emissions dataset converted into preferences-style JSONL format for DPO training. It can be directly used by DPOTrainer, axolotl, etc. The prompt consists of an instruction and text extracted from relevant pages of a sustainability report. The chosen output is generated using the Mixtral-8x7B-v0.1 model and consists of a JSON string containing the scope 1, 2 and 3 emissions as well as the ids of pages containing this information. The rejected output is randomly… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/sustainability-report-emissions-dpo.texttext-generation1K<n<10K1 likes31 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.