datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.esg_reports_v2
Vidore Benchmark 2 - ESG Restaurant Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "french" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_v2.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.economics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
lk-tourism-weekly-reports-docslk-tourism-weekly-reports-chunkscbsl-annual-reports-chunkscbsl-annual-reports-docssyntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
lk-tourism-monthly-reports-chunkslk-tourism-monthly-reports-docsoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.lk-dmc-situation-reports-chunkscorral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.lk-dmc-situation-reports-docssyntheticDocQA_government_reports_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Government Reports that allow ViDoRe to benchmark administrative/legal documents.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('government reports'). From these documents, we randomly sampled 1000 pages.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_government_reports_test.popcorn-reportsasrs-aviation-reports
Dataset Card for ASRS Aviation Incident Reports
Dataset Summary
This dataset collects 47,723 aviation incident reports published in the Aviation Safety Reporting System (ASRS) database maintained by NASA.
Supported Tasks and Leaderboards
'summarization': Dataset can be used to train a model for abstractive and extractive summarization. The model performance is measured by how high the output summary's ROUGE score for a given narrative account of an aviation… See the full description on the dataset page: https://huggingface.co/datasets/elihoole/asrs-aviation-reports.annual_reportsThe dataset is a comprehensive collection of financial documents from corporate annual reports.
Reference: https://www.annualreports.com/
syntheticDocQA_government_reports_test_tesseracteconomics_reports_eng_v2
Vidore Benchmark 2 - World Economics report Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to World economic reports. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_eng_v2.company-reports
Company Reports Dataset
Description
This dataset contains ESG (Environmental, Social, and Governance) sustainability reports from various companies. It includes data like company details, report categories, textual analysis of the reports, and more.
Dataset Structure
id: Unique identifier for each report entry.
document_category: Classification of the document (e.g., ESG sustainability report).
year: Publication year of the report.
company_name: Name of the… See the full description on the dataset page: https://huggingface.co/datasets/DataNeed/company-reports.esg_reports_eng_v2
Vidore Benchmark 2 - ESG Restaurant Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
Each query is in french.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports of fast food companies. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_eng_v2.corral-QAs-reports
Corral – QA Reports
Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.IMF-Reports
IMF Technical Assistance Reports — Recommendation Process Corpus
A page-grounded research corpus of 780 IMF technical-assistance report
records. It contains source PDFs, layout-aware Markdown, page-level text, extracted
visuals, metadata, observations, recommendations, and labeled links between observations
and recommendations.
Required acknowledgement
All research, publications, datasets, models, applications, or other work derived from
this corpus should… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/IMF-Reports.financial-reports
Description
Topic: Financial Reports
Domains: Finance, Accounting, Economics
Focus: Synthetic raw financial reports for analysis and training
Number of Entries: 1000
Dataset Type: Raw Dataset
Model Used: bedrock/us.amazon.nova-pro-v1:0
Language: English
Generated by: SynthGenAI Package
Bugzilla_Eclipse_Bug_Reports_Dataset
Special Thanks
Special thanks to Lamkanfi, Ahmed; Pérez, Javier; and Demeyer, Serge for their contributions. Please cite their paper, as this dataset is the processed part of their dataset.
Citation
@INPROCEEDINGS{6624028,
author={Lamkanfi, Ahmed and Pérez, Javier and Demeyer, Serge},
booktitle={2013 10th Working Conference on Mining Software Repositories (MSR)},
title={The Eclipse and Mozilla defect tracking dataset: A genuine dataset for mining bug information}… See the full description on the dataset page: https://huggingface.co/datasets/AliArshad/Bugzilla_Eclipse_Bug_Reports_Dataset.McKinsey-Reportsmeta-llama/synthetic-data-kit
https://github.com/meta-llama/synthetic-data-kit
McKinsey reports
https://www.mckinsey.com/featured-insights/insights-store
