datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multimodal-ct-radiology-reports
Perle AI Multi-phase CECT and CT with Radiology Reports
Summary
A de-identified CT dataset from Perle AI, paired with the original radiology reports. It supports work on multi-modal medical imaging: phase or pathology classification, report generation from images, and visual question answering.
The release has three configurations:
Config
Modality
Subjects
Pairing
cect_3phase
3-phase contrast-enhanced abdominal CT (DICOM)
5
per-subject text report +… See the full description on the dataset page: https://huggingface.co/datasets/Perle-ai/multimodal-ct-radiology-reports.financial-reports-secThe dataset contains the annual report of US public firms filing with the SEC EDGAR system.
Each annual report (10K filing) is broken into 20 sections. Each section is split into individual sentences.
Sentiment labels are provided on a per filing basis from the market reaction around the filing data.
Additional metadata for each filing is included in the dataset.esg_reports_v2
Vidore Benchmark 2 - ESG Restaurant Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "french" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_v2.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.economics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
lk-tourism-weekly-reports-docsgov_reportGovReport long document summarization dataset.
There are three configs:
- plain_text: plain text document-to-summary pairs
- plain_text_with_recommendations: plain text doucment-summary pairs, with "What GAO recommends" included in the summary
- structure: data with section structurelk-tourism-weekly-reports-chunkscbsl-annual-reports-chunkscbsl-annual-reports-docsESG_Report
ESG Report PDF Dataset
Download Instructions
To download the dataset, follow these steps:
Navigate to the data directory in the GitHub repository:
cd data
Install Git LFS (if not already installed):
git lfs install
Clone the dataset from Hugging Face Hub:
git clone https://huggingface.co/datasets/WHATX/ESG_Report
Dataset Description
This dataset contains three main components:
raw_pdf:
A collection of 195 PDFs scraped from TCFD Hub.
The PDFs are… See the full description on the dataset page: https://huggingface.co/datasets/WHATX/ESG_Report.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
megawika-report-generation
Dataset Card for MegaWika for Report Generation
Dataset Summary
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span
50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a
non-English language, an automated English translation is provided.
This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.lk-tourism-monthly-reports-chunkslk-tourism-monthly-reports-docsoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.lk-dmc-situation-reports-chunkscorral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.lk-dmc-situation-reports-docssyntheticDocQA_government_reports_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Government Reports that allow ViDoRe to benchmark administrative/legal documents.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('government reports'). From these documents, we randomly sampled 1000 pages.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_government_reports_test.popcorn-reportsmultilingual-EMIR-reporting-parquet-mixedtwice_kr_market_report_retrieval
FinMarketReport-Retrieval-ko
Constructed a Retrieval dataset related to the stock market, based on Korean Financial Reports.
asrs-aviation-reports
Dataset Card for ASRS Aviation Incident Reports
Dataset Summary
This dataset collects 47,723 aviation incident reports published in the Aviation Safety Reporting System (ASRS) database maintained by NASA.
Supported Tasks and Leaderboards
'summarization': Dataset can be used to train a model for abstractive and extractive summarization. The model performance is measured by how high the output summary's ROUGE score for a given narrative account of an aviation… See the full description on the dataset page: https://huggingface.co/datasets/elihoole/asrs-aviation-reports.mlb-daily-reportannual_reportsThe dataset is a comprehensive collection of financial documents from corporate annual reports.
Reference: https://www.annualreports.com/
syntheticDocQA_government_reports_test_tesseractdocqa_gov_report_beirThis is a copy of https://huggingface.co/datasets/jinaai/docqa_gov_report reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/docqa_gov_report_beir.economics_reports_eng_v2
Vidore Benchmark 2 - World Economics report Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to World economic reports. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_eng_v2.
