datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
climate-fever
ClimateFEVER
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset adopting the FEVER methodology that consists of 1,535 real-world claims (queries) regarding climate-change. The underlying corpus is the same as FVER.
Task category
t2t
Domains
Encyclopaedic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html
How to evaluate on this task
You can evaluate an embedding model on this dataset using… See the full description on the dataset page: https://huggingface.co/datasets/mteb/climate-fever.fineweb-edu-climateclimatebot-dataclimate-fever-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-generated-queries.ClimateFund
ClimateFund: An Annotated Dataset of Climate Mitigation Projects for Supporting Question Answering
This repository contains a dataset based on funding proposals of 21 climate mitigation projects, submitted to the Green Climate Fund (GCF).
Climate mitigation documentation is challenging to parse and understand, due to the length of this documents, their multi-modality (commonly comprising tables, figures and
free text), and their highly technical and domain-specific content.… See the full description on the dataset page: https://huggingface.co/datasets/JavierSanzCruza/ClimateFund.PCFBench
PCFBench
Paper: arXiv:2608.27716 · Code: watershed-climate/pcfbench
Process-based Product Carbon Footprint benchmark for evaluating LLMs and
agents on the operational steps of life-cycle assessment (LCA): bill-of-
materials decomposition, mapping triage, ecoinvent process matching,
literature extraction of physical input rates, and total kgCO₂e
prediction against expert-grounded EPDs.
Tasks
ID
Task
Items
GT claims
Headline metric
1
Product decomposition… See the full description on the dataset page: https://huggingface.co/datasets/Watershed-Climate/PCFBench.climate-fever-v2
ClimateFEVER.v2
An MTEB dataset
Massive Text Embedding Benchmark
CLIMATE-FEVER is a dataset following the FEVER methodology, containing 1,535 real-world climate change claims. This updated version addresses corpus mismatches and qrel inconsistencies in MTEB, restoring labels while refining corpus-query alignment for better accuracy.
Task category
t2t
Domains
Academic, Written
Reference
https://www.sustainablefinance.uzh.ch/en/research/climate-fever.html… See the full description on the dataset page: https://huggingface.co/datasets/mteb/climate-fever-v2.LLM-Misalign-Climate-ChangeThis dataset accompanies the paper:
LLM Benchmark-User Need Misalignment for Climate Change
📄 Paper link: https://arxiv.org/abs/2603.26106;
🛠️ Github link: https://github.com/OuchengLiu/LLM-Misalign-Climate-Change
📌 Dataset Summary
This dataset provides a unified collection of climate change–related knowledge interactions for conducting the analyses presented in the paper. It integrates data from multiple sources, covering both human–AI and human–human interactions, and… See the full description on the dataset page: https://huggingface.co/datasets/Westing/LLM-Misalign-Climate-Change.corect-climate-feverHindi_Climate_Feverclimate-change-MRCThe Climate Change MRC dataset, also known as CCMRC, is a part of the work "Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering", accepted at IJCAI-ECAI 2022. The paper was accepted in the special system demo track "AI for Good".
If you use the dataset, cite the following paper:
@inproceedings{rony2022climatemrc,
title={Climate Bot: A Machine Reading Comprehension System for Climate Change Question Answering.},
author={Rony, Md Rashad Al Hasan and Zuo… See the full description on the dataset page: https://huggingface.co/datasets/rony/climate-change-MRC.iecc-climate-zone-by-county
IECC/Building America climate zone by U.S. county, 2021 code cycle
Canonical, always-current version: https://referencesource.org/iecc-climate-zone-by-county/
Machine-readable: https://referencesource.org/iecc-climate-zone-by-county/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2028-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 3134
Which IECC climate zone (1-8, with… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/iecc-climate-zone-by-county.hvac-climate-data
HVAC Climate Telemetry Dataset
⚠️ Synthetic Data Disclaimer: This dataset contains synthetically generated data for demonstration and testing purposes. It does not represent real building measurements or actual climate conditions.
Overview
Real-time HVAC and climate telemetry data for building energy management systems. This dataset provides indoor/outdoor environmental conditions, HVAC status, and 24-hour trends.
Schema
{
"pipeline": "hvac_climate_data"… See the full description on the dataset page: https://huggingface.co/datasets/shahabsalehi/hvac-climate-data.ai-climate-compute-2026climate-fever-fa
Dataset Summary
ClimateFEVER-Fa is a Persian (Farsi) dataset tailored for the Retrieval task, specifically focused on fact-checking climate-related claims. It is the translated version of the English ClimateFEVER dataset and is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) under the BEIR-Fa collection.
Language(s): Persian (Farsi)
Task(s): Retrieval (Fact Checking, Evidence Retrieval for Climate Claims)
Source: Translated from the English ClimateFEVER dataset… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/climate-fever-fa.Indian_Climate_Disaster_data
BharatCRIC: Indian Climate Disaster Data (Heatwave Advisories + Scam Pairs)
Generated by scripts/grade_a_rebuild.py with blueprint grade_a_2026_04_30.
File
Use
instruction_dataset_main.jsonl
Full structured instruction upload (1680 rows)
instruction_dataset_smoke.jsonl
50-row smoke test, 5 languages x 5 formats x 2 rows
instruction_dataset_smoke_v2.jsonl
Same smoke set for the previous filename expected by notes
preference_pairs_scams.jsonl
Mirrored genuine-vs-scam… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Disaster_data.Indian_Climate_Resilience_Instruction_Corpus_
IndianCRIC — Indian Climate Resilience Instruction Corpus
5 languages · 5 formats · genuine ↔ scam pairs
Built for the Adaption Labs Uncharted Data Challenge 2026
Why this dataset exists
The Vulnerable people of Bihar, Uttar Pradesh, and Jharkhand sit at the intersection of high heat vulnerability and low AI coverage.
During extreme weather events, official advisories compete with misinformation — fake helplines, fraudulent relief schemes, and OTP scams disguised… See the full description on the dataset page: https://huggingface.co/datasets/sahilmaniyar888/Indian_Climate_Resilience_Instruction_Corpus_.fineweb-climate-filteredearth-love-united-climate-knowledge
🌍 Earth Love United Climate Knowledge Dataset
The most comprehensive open climate science knowledge dataset.
10,128 text chunks + 124 structured facts + 4.54B year geological memory + 10 tipping points.
Built to power GAIA — an AI that embodies the living consciousness of Earth.
Dataset Overview
This dataset gives an AI system authoritative, sourced knowledge about climate change,
carbon, Earth science, and solutions. It has four layers:
Layer 1: Text Knowledge… See the full description on the dataset page: https://huggingface.co/datasets/ego0op/earth-love-united-climate-knowledge.ClimSight-GPT4.1-Mini-Climate-Assesments
Dataset Card for ClimSight Dataset
This dataset, generated using ClimSight, provides a curated collection of localized climate insights. ClimSight is an advanced tool
that integrates Large Language Models (LLMs) with comprehensive climate data to deliver actionable insights for critical
decision-making across various sectors, including agriculture, urban planning, disaster management, and policy development. It
transforms complex climate data into easily digestible and relevant… See the full description on the dataset page: https://huggingface.co/datasets/CliDyn/ClimSight-GPT4.1-Mini-Climate-Assesments.beir-nl-climate-fever
Dataset Card for BEIR-NL Benchmark
Dataset Summary
BEIR-NL is a Dutch-translated version of the BEIR benchmark, a diverse and heterogeneous collection of datasets covering various domains from biomedical and financial texts to general web content. Our benchmark is integrated into the Massive Multilingual Text Embedding Benchmark (MMTEB).
BEIR-NL contains the following tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018… See the full description on the dataset page: https://huggingface.co/datasets/clips/beir-nl-climate-fever.benchmark-climatefeverclimate-fever-top-20-gen-queries
NFCorpus: 20 generated queries (BEIR Benchmark)
This HF dataset contains the top-20 synthetic queries generated for each passage in the above BEIR benchmark dataset.
DocT5query model used: BeIR/query-gen-msmarco-t5-base-v1
id (str): unique document id in NFCorpus in the BEIR benchmark (corpus.jsonl).
Questions generated: 20
Code used for generation: evaluate_anserini_docT5query_parallel.py
Below contains the old dataset card for the BEIR benchmark.
Dataset Card for BEIR… See the full description on the dataset page: https://huggingface.co/datasets/income/climate-fever-top-20-gen-queries.adaption-bsad-agri-climate-resilience
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_agri_climate_resilience
This dataset contains synthetic entries for agricultural biotechnology studies focused on climate resilience, with a specific emphasis on the Global South. Each record includes simulated study details, organism information, genetic engineering methods, and biosafety levels formatted as CSV-like text. The data is designed to represent low-risk research scenarios under… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-agri-climate-resilience.adaption-bsad-agri-climate-resilience-v1
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_agri_climate_resilience
This dataset contains synthetic entries for agricultural biotechnology studies focused on climate resilience, with a specific emphasis on the Global South. Each record includes simulated study details, organism information, genetic engineering methods, and biosafety levels formatted as CSV-like text. The data is designed to represent low-risk research scenarios under… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-agri-climate-resilience-v1.ClimateChangeQA_hardadaption-bsad-climate-agri-synthetic
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
bsad_climate_agri_synthetic
This dataset contains synthetic entries for climate-smart agriculture and food systems research, specifically focusing on genetic engineering applications. Each record includes simulated study details, organism identifiers, biosafety levels (BSL-1), and global scope indicators formatted as CSV-like text. The data is artificially generated to mimic structured biological… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-bsad-climate-agri-synthetic.ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3 30B A3B FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3-30B-A3B-Instruct-2507-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 2, non-eager mode
Max tokens: 4096
Generation config: temperature 0.7, top_p 0.8, top_k 20, min_p 0.0… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen3-30b-a3b-fp8-10k-seed42.africa-open-climate-data
Open Government Climate & Agricultural Data for Africa
Catalog of openly available US government and international datasets for Africa —
NASA, NOAA, USDA, FEWS NET. All public domain or open access.
These are the raw material for AI coordination tools doing drought monitoring,
agricultural planning, and food security analysis.
Datasets Cataloged
Source
Dataset
Resolution
Period
NASA TRMM
Monthly precipitation
0.25°
1998-2015
NASA MODIS
Land surface… See the full description on the dataset page: https://huggingface.co/datasets/gmahia/africa-open-climate-data.ClimateMBERT-syn-qwen35-122b-fp8-10k-seed42
ClimateMBERT Synthetic Qwen3.5 FP8 10K Seed42
Synthetic continuation dataset generated from WxChat/ClimateMBERT_syn train split.
Source dataset: WxChat/ClimateMBERT_syn
Source split: train
Sampling: shuffled with random seed 42, ranks 0..9999
Rows: 10,000
Generator: Qwen/Qwen3.5-122B-A10B-FP8
Inference: vLLM on Clariden GH200 GPUs, tensor parallel size 4, non-eager mode
Max tokens: 4096
No-thinking mode: chat_template_kwargs={"enable_thinking": false}
Generation config: temperature… See the full description on the dataset page: https://huggingface.co/datasets/JingweiNi/ClimateMBERT-syn-qwen35-122b-fp8-10k-seed42.
