datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esg_reports_v2
Vidore Benchmark 2 - ESG Restaurant Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "french" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_v2.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.economics_reports_v2
Vidore Benchmark 2 - World Economics report Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "english" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_v2.syntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
lk-tourism-weekly-reports-chunkslk-tourism-weekly-reports-docscbsl-annual-reports-chunkscbsl-annual-reports-docssyntheticDocQA_government_reports_test_beirBEIR version of vidore/syntheticDocQA_government_reports_test.
lk-tourism-monthly-reports-chunkslk-tourism-monthly-reports-docsoag-nepal-audit-reports
OAG Nepal Audit Reports — Nepali transcripts and ruled tables
Machine-readable transcripts of 6,234 publications of the Office of the Auditor General of Nepal (महालेखा परीक्षकको कार्यालय, OAG) — the annual audit reports of local governments, provinces and central bodies, plus the OAG's own bulletins, journals and financial statements.
The OAG publishes these as PDFs whose text layer is, for most documents, legacy pre-Unicode Devanagari: fonts like Preeti and Fontasy Himali that… See the full description on the dataset page: https://huggingface.co/datasets/damo-da/oag-nepal-audit-reports.lk-dmc-situation-reports-chunkscorral_runs_reports
Corral – Evaluation Score Reports
Reports from Corral evaluation runs across models, scaffolds, scopes, and task granularities in all 8 environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the Reports produced during the evaluation runs of models across all 8 Corral environments.
The dataset is organized into 24 configurations… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral_runs_reports.lk-dmc-situation-reports-docssyntheticDocQA_government_reports_test
Dataset Description
This dataset is part of a topic-specific retrieval benchmark spanning multiple domains, which evaluates retrieval in more realistic industrial applications.
It includes documents about the Government Reports that allow ViDoRe to benchmark administrative/legal documents.
Data Collection
Thanks to a crawler (see below), we collected 1,000 PDFs from the Internet with the query ('government reports'). From these documents, we randomly sampled 1000 pages.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/syntheticDocQA_government_reports_test.popcorn-reportssyntheticDocQA_government_reports_test_tesseractIMF-Reports
IMF Technical Assistance Reports — Recommendation Process Corpus
A page-grounded research corpus of 780 IMF technical-assistance report
records. It contains source PDFs, layout-aware Markdown, page-level text, extracted
visuals, metadata, observations, recommendations, and labeled links between observations
and recommendations.
Required acknowledgement
All research, publications, datasets, models, applications, or other work derived from
this corpus should… See the full description on the dataset page: https://huggingface.co/datasets/FrenchCastle/IMF-Reports.economics_reports_eng_v2
Vidore Benchmark 2 - World Economics report Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of World economic reports from 2024.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to World economic reports. It includes a curated set of documents, queries, relevance judgments (qrels), and page images.… See the full description on the dataset page: https://huggingface.co/datasets/vidore/economics_reports_eng_v2.esg_reports_eng_v2
Vidore Benchmark 2 - ESG Restaurant Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
Each query is in french.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports of fast food companies. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_eng_v2.financial-reports
Description
Topic: Financial Reports
Domains: Finance, Accounting, Economics
Focus: Synthetic raw financial reports for analysis and training
Number of Entries: 1000
Dataset Type: Raw Dataset
Model Used: bedrock/us.amazon.nova-pro-v1:0
Language: English
Generated by: SynthGenAI Package
corral-QAs-reports
Corral – QA Reports
Model completions for question-answer evaluations probing factual knowledge and reasoning across Corral environments
📋 Dataset Summary
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the model completions and reports for the question-answer evaluations used to test the factual knowledge and reasoning ability of models across Corral… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/corral-QAs-reports.osha-severe-injury-reports-enriched
OSHA Severe Injury Reports Enriched Dataset
This dataset packages OSHA Severe Injury Report records into a single analysis-ready dataframe covering January 1, 2015 through August 31, 2025.
Each row represents one OSHA severe injury report submitted under federal OSHA jurisdiction. The dataset preserves the underlying establishment, geography, NAICS, incident classification, and narrative fields while adding normalized dates, NAICS sector mapping, severity profile fields… See the full description on the dataset page: https://huggingface.co/datasets/masonmarker/osha-severe-injury-reports-enriched.esg-reports
ConTEB - ESG Reports
This dataset is part of ConTEB (Context-aware Text Embedding Benchmark), designed for evaluating contextual embedding model capabilities. It focuses on the theme of Industrial ESG Reports, particularly stemming from the fast-food industry.
Dataset Summary
This dataset was designed to elicit contextual information. It is built upon a subset of the ViDoRe Benchmark. To build the corpus, we start from the pre-existing collection of ESG Reports, extract… See the full description on the dataset page: https://huggingface.co/datasets/illuin-conteb/esg-reports.climateqa-ipcc-ipbes-reports-1.0Dataset card WIP
This dataset is storing all the chunks used in the ClimateQ&A assistant https://huggingface.co/spaces/Ekimetrics/climate-question-answering when using RAG to answer any questions about the environment.
It constitutes all of the recent IPCC and IPBES reports listed here https://climateqa.com/docs/sources parsed in a machine-readable format.
Dataset construction
The dataset was constructed with a layout segmentation algorithm + an OCR algorithm and then… See the full description on the dataset page: https://huggingface.co/datasets/Ekimetrics/climateqa-ipcc-ipbes-reports-1.0.uav-fault-symptom-reports
UAV Fault Symptom Reports
A synthetic dataset of UAV flight telemetry paired with operator-style symptom reports written by a
language model. Each row is one five-second window of a flight: 20 telemetry channels, the fault
class, a severity derived from simulated consequences, and a one-sentence report.
split
rows
flights
model-written reports
unique reports
benchmark
10,500
2,100
82.0%
79.7%
challenge
3,500
700
88.1%
91.0%
benchmark is balanced across seven… See the full description on the dataset page: https://huggingface.co/datasets/Bashifu/uav-fault-symptom-reports.SnP500-annual-and-sustainability-reportssection1_annual_reports_tokenized_llama3_8bcorporate-emission-reports
Dataset Card for Dataset Name
A dataset of 100 corporate sustainability reports with manually extracted scope 1, 2 and 3 greenhouse gas emission values.
Dataset Details
Dataset Description
Data about corporate greenhouse gas emissions is usually published only as part of sustainability report PDF's, which is not a machine-readable format. Interested actors have to manually extract emission data from these reports, which is a tedious and time-consuming process.… See the full description on the dataset page: https://huggingface.co/datasets/nopperl/corporate-emission-reports.
