datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
esg_reports_v2
Vidore Benchmark 2 - ESG Restaurant Dataset (Multilingual)
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
The dataset contain queries in the following languages : ["english", "french", "german", "spanish"]. Each query was originaly in "french" (see… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_v2.esg_reports_human_labeled_v2
Vidore Benchmark 2 - ESG Human Labeled
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports from the fast food industry.
Dataset Summary
Each query is in english.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports for the fast food industry. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_human_labeled_v2.esg_cid_retrieval
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Usage
from datasets import load_dataset
# document chunks: train/dev/test_gri/test_esrs
documents = load_dataset("esgllm/esg_cid_retrieval", "documents")
# queries (disclosure text): train/dev/test_gri/test_esrs
queries = load_dataset("esgllm/esg_cid_retrieval", "queries")
# training triplets: train/dev
triplets = load_dataset("esgllm/esg_cid_retrieval"… See the full description on the dataset page: https://huggingface.co/datasets/airefinery/esg_cid_retrieval.esg_reports_eng_v2
Vidore Benchmark 2 - ESG Restaurant Dataset
This dataset is part of the "Vidore Benchmark 2" collection, designed for evaluating visual retrieval applications. It focuses on the theme of ESG reports in the fast food industry.
Dataset Summary
Each query is in french.
This dataset provides a focused benchmark for visual retrieval tasks related to ESG reports of fast food companies. It includes a curated set of documents, queries, relevance judgments (qrels), and page… See the full description on the dataset page: https://huggingface.co/datasets/vidore/esg_reports_eng_v2.environmental_2kgovernance_2ksocial_2kalistairking_public-company-esg-ratings-dataset
Public Company ESG Ratings Dataset
ESG ratings for over 700 mid / large-cap companies across various industries
Dataset Info
Source: Kaggle
Original Size: 0.04 MB
Kaggle Downloads: 5,844
Files: 1
Files
data.csv
Mirrored from Kaggle
ESGenius
ESGenius
ESGenius is an EMNLP 2025 Main Conference Oral benchmark for evaluating large language models on Environmental, Social, and Governance (ESG) and sustainability knowledge. The paper was nominated for the EMNLP 2025 Resource and Theme Paper Awards, Top 1%.
Paper: https://aclanthology.org/2025.emnlp-main.739/
Project site: https://angel-ntu.github.io/ESGenius/
GitHub repository: https://github.com/ANGEL-NTU/ESGenius
Interactive heatmap:… See the full description on the dataset page: https://huggingface.co/datasets/cy0307/ESGenius.spx-sustainalytics-esg-scoresagua_e_esgoto_nordeste_brasileiro
Brazilian Northeast Water and Sanitation Crisis Dataset (BNWSC)
Overview
This dataset provides multidisciplinary data on water access, sanitation, public health, and socioeconomic disparities in Brazil's Northeast region. It integrates official sources (2014–2022) and includes projections up to 2030, supporting research in public policy, collective health, and sustainability.
Key Features
Applications
Correlation analysis between sanitation, income… See the full description on the dataset page: https://huggingface.co/datasets/carpenterbb/agua_e_esgoto_nordeste_brasileiro.esg-spanish-events
ESG Spanish Events (2014–2024)
Dataset summary
Multi-label ESG classification dataset for Spanish equity market news
(2014–2024). Contains 1,688 human-annotated canonical news events
(gold set) and 64,488 LLaMA-3.1-8B SFT silver-annotated events
(silver set). Labels cover three ESG pillars (Environmental, Social,
Governance) as independent binary signals plus a 4-class sentiment
dimension (Positive / Negative / Neutral / NA).
Companion model: DReggio/mrbert-es-esg… See the full description on the dataset page: https://huggingface.co/datasets/DReggio/esg-spanish-events.esg-bank-dataset-v4
Vietnamese Banking ESG Disclosure Dataset
Description
Multi-task annotated dataset for ESG (Environmental, Social, Governance)
disclosure analysis in the Vietnamese banking sector. Each sample is a
text chunk from bank sustainability/annual reports, annotated with:
Greenwashing classification (Legitimate / Greenwashing / Uncertain)
ESG pillar (Environmental / Social / Governance / General)
Content quality score (0-100)
ESG score (0-100)
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/hiennthp/esg-bank-dataset-v4.action_500WaterForestBiodiversityNature_2200esg_reports_human_labeled_v2_reasoningESGdatasetsESGDatasetesg_reports_eng_v2_reasoningesg-sentiment
Dataset Card for "esg-sentiment"
More Information needed
ESG_ver1_rationaleesg_reports_eng_v2ViDoRe_esg_reports_v2_multilingualsp500_ESG_Datasetmining-esg-scoringViDoRe_esg_reports_v2ViDoRe_esg_reports_human_labeled_v2corporate_esg_risk_analyticssp500_ESG_Dataset
