datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ClimateFEVER_test_top_250_only_w_correct-v2
ClimateFEVERHardNegatives
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change. The hard negative version has been created by pooling the 250 top documents per query from BM25, e5-multilingual-large and e5-mistral-instruct.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/ClimateFEVER_test_top_250_only_w_correct-v2.jax-gcm-data
jax-gcm boundary conditions and emissions
Input data for jax-gcm
(jcm), a fully differentiable atmospheric GCM in JAX. Two tiers:
products/ — grid-independent source products
store
contents
source
ceds_anthro.zarr
anthropogenic SO2/BC/OC/NH3 flux, sector-summed, 0.5°, monthly 1850–2023 + PI (1850–59) / PD (2005–14) climatologies
CEDS-CMIP-2025-04-18 (input4MIPs CMIP7)
bb4cmip7.zarr
open-burning SO2/BC/OC/NH3 flux, 0.25°, monthly 1850–2023 + PI/PD… See the full description on the dataset page: https://huggingface.co/datasets/climate-analytics-lab/jax-gcm-data.ClimateBench-M-IMGClimateSuite
ClimateSuite
If you use ClimateSuite, please ❤️ like this dataset and ⭐ star the SPF GitHub repository for updates.
ClimateSuite is a large collection of Earth system model simulations for climate emulation. It was curated for Spatiotemporal Pyramid Flow Matching for Climate Emulation, which introduces Spatiotemporal Pyramid Flows (SPF) for efficient probabilistic climate emulation across temporal and spatial scales.
This repository stores ClimateSuite as archive-sharded Zarr… See the full description on the dataset page: https://huggingface.co/datasets/jirvin16/ClimateSuite.ClimateNetclimate-traceclimate-fever
ClimateFEVER
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims (queries) regarding climate-change. The underlying corpus is the same as FVER.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/climate-fever.climate_fever
Dataset Card for ClimateFever
Dataset Summary
A dataset adopting the FEVER methodology that consists of 1,535 real-world claims regarding climate-change collected on the internet. Each claim is accompanied by five manually annotated evidence sentences retrieved from the English Wikipedia that support, refute or do not give enough information to validate the claim totalling in 7,675 claim-evidence pairs. The dataset features challenging claims that relate multiple facets… See the full description on the dataset page: https://huggingface.co/datasets/tdiggelm/climate_fever.climate-learn
Dataset Card for Dataset Name
Dataset Summary
Data used for ClimateLearn's benchmark experiments.
Supported Tasks
Weather forecasting
Statistical downscaling
Climate projection
Additional Information
Dataset Curators
Maintained by the Machine Intelligence Group at UCLA, headed by Professor Aditya Grover. Please contact Jason Jewik at jason.jewik@ucla.edu for any questions, or open an issue on our GitHub/HuggingFace page.… See the full description on the dataset page: https://huggingface.co/datasets/jasonjewik/climate-learn.africa-world-bank-climate-change-indicators-for-cameroon
Cameroon - Climate Change | Africa (original)
Size category: 1K<n<10K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-world-bank-climate-change-indicators-for-cameroon.climate-fever-decontaminated
climate-fever (Decontaminated)
A decontaminated version of the climate-fever dataset from the BEIR benchmark, with samples found in the mgte-en pre-training dataset removed.
Decontamination methodology
Contamination was detected using a two-pass approach against the full mgte-en dataset (484 GB, 1,235 parquet files):
Pass 1: Exact hash matching
All texts (queries and corpus documents) were normalized (lowercased, unicode NFKD, whitespace collapsed) and hashed… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/climate-fever-decontaminated.africa-world-bank-climate-environment-time-series
Africa World Bank Climate and Environment Labeled Time Series Data
This repository is part of the Africa Temporal Intelligence Corpus (ATIC). It contains sector-specific temporal corpus packages for African countries.
ATIC sector repositories are designed for machine consumption first: Parquet tables, stable IDs, reproducible metadata, explicit provenance, review status, and separable semantic layers.
Sector Scope
Temporal climate and environmental indicators for… See the full description on the dataset page: https://huggingface.co/datasets/africatic/africa-world-bank-climate-environment-time-series.climate-fever
Dataset Card for BEIR Benchmark
climate-fever is one of the datasets from the Fact Checking task within BEIR, measuring Wikipedia abstract retrieval for a given claim about climate.
Dataset Summary
BEIR is a heterogeneous benchmark built from 18 diverse datasets representing 9 information retrieval tasks.
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever.climate_sentiment
Dataset Card for climate_sentiment
Dataset Summary
We introduce an expert-annotated dataset for classifying climate-related sentiment of climate-related paragraphs in corporate disclosures.
Supported Tasks and Leaderboards
The dataset supports a ternary sentiment classification task of whether a given climate-related paragraph has sentiment opportunity, neutral, or risk.
Languages
The text in the dataset is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/climate_sentiment.climatecasechart
Data
All cases and included documents from climatecasechart.com, that
are Non-US jurisdiction
filed in a WHO53 country
or Arbitral Tribunal, European Committee on Social Rights, European Union, International Courts & Tribunals, World Trade Organization United Nations OECD
have an (estimated) filing year, estimated if not specified by earliest document filing date, and
that filing year is from 2011 up to and including 2024.
For these cases all documents are collected where… See the full description on the dataset page: https://huggingface.co/datasets/ug-dsc/climatecasechart.climate-evaluationDatasets for Climate Evaluation.france-climate-risk-drias-tracc-2023
France Climate Risk DRIAS TRACC-2023
Ce dataset regroupe des fichiers NetCDF issus de DRIAS TRACC-2023 pour la France métropolitaine, organisés pour un usage plus simple.
Le dépôt contient des indicateurs climatiques :
pour les niveaux de réchauffement +2.0 °C, +2.7 °C et +4.0 °C ;
en produits multi-modèles ;
en produits par modèle ;
pour les valeurs absolues ;
ainsi que pour la période de référence historique.
Structure du dépôt
Le dépôt est organisé par niveau de… See the full description on the dataset page: https://huggingface.co/datasets/saadtaleb/france-climate-risk-drias-tracc-2023.climatecheck
The ClimateCheck Dataset
This dataset is used for the ClimateCheck: Scientific Fact-checking of Social Media Posts on Climate Change Shared Task.
The 2025 iteration was hosted at the Scholarly Document Processing workshop at ACL 2025, and a new 2026 iteration will be hosted at the Natural Scientific Language Processing workshop at LREC 2026.
2026 Update
For running the next iteration of the task, we added manually labelled training data, resulting in 3023… See the full description on the dataset page: https://huggingface.co/datasets/rabuahmad/climatecheck.climate-trace
Climate TRACE Spain Emissions
This dataset mirrors the Climate TRACE country packages for Spain (ISO alpha-3 ESP). It contains Parquet (zstd) extracts converted from the CSVs published by Climate TRACE. All columns are stored as text to avoid type inference failures. Cast as needed.
Contents
raw/<gas>/ABOUT_THE_DATA/: original documentation supplied by Climate TRACE (PDF methodology summary and CSV data dictionary).
raw/<gas>/DATA/: annual emission series (2015 onward)… See the full description on the dataset page: https://huggingface.co/datasets/datania/climate-trace.fineweb-edu-climatevindhya-climateenvironmental_claims
Dataset Card for environmental_claims
Dataset Summary
We introduce an expert-annotated dataset for detecting real-world environmental claims made by listed companies.
Supported Tasks and Leaderboards
The dataset supports a binary classification task of whether a given sentence is an environmental claim or not.
Languages
The text in the dataset is in English.
Dataset Structure
Data Instances
{
"text": "It will enable E.ON to… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/environmental_claims.climate_detection
Dataset Card for climate_detection
Dataset Summary
We introduce an expert-annotated dataset for detecting climate-related paragraphs in corporate disclosures.
Supported Tasks and Leaderboards
The dataset supports a binary classification task of whether a given paragraph is climate-related or not.
Languages
The text in the dataset is in English.
Dataset Structure
Data Instances
{
'text': '− Scope 3: Optional scope that includes… See the full description on the dataset page: https://huggingface.co/datasets/climatebert/climate_detection.tl-climate-projectionsclimatebot-dataclimate-fever-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-generated-queries.ClimateBench-M-TSclimate-fever-qrels
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-qrels.finewebedu-climate-v2ClimateBench-M-TS-processed
