datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earth-Silver
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Silver.Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Iron.EarthScience-Text-LLM-20K-90-10
EarthScience-Text-LLM-20K-90-10
This is a pure-text Earth-science corpus unified from three non-overlapping upstream datasets:
Ekimetrics/climateqa-ipcc-ipbes-reports-1.0: climate and IPCC/IPBES report chunks.
GeoGPT-Research-Project/GeoGPT-CoT-QA: geoscience question-answer reasoning.
gremlin97/RemoteSensingCorpus: remote-sensing and geospatial machine-learning text.
Files and Split
The previous preprocessing outputs were merged into a 23,098-record pool and… See the full description on the dataset page: https://huggingface.co/datasets/moTcream/EarthScience-Text-LLM-20K-90-10.Earth-Gold
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Gold.Earth-Iron
Dataset Card for Earth-Iron
Dataset Details
Dataset Description
Earth-Iron is a comprehensive question answering (QA) benchmark designed to evaluate the fundamental scientific exploration abilities of large language models (LLMs) within the Earth sciences. It features a substantial number of questions covering a wide range of topics and tasks crucial for basic understanding in this domain. This dataset aims to assess the foundational knowledge that underpins… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Iron.kasa-mcp-indirect-channel-probes
KASA MCP — Indirect-Channel Agent Probes
Four probes measuring whether untrusted content — not the operator — can steer a local model that sits inside an agent pipeline. Four model configurations, five runs each, 80 rows.
The dataset exists because it caught a failure in the architecture that produced it. The headline result is A8: 20 out of 20 runs compromised, on every configuration tested.
What each probe measures
Probe
Channel
Question… See the full description on the dataset page: https://huggingface.co/datasets/Earthen937/kasa-mcp-indirect-channel-probes.earth_silver
Bud Ecosystem mirror of ai-earth/Earth-Silver — a verbatim copy for offline, reproducible model evaluation. License unchanged (MIT); all rights remain with the original authors.
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in… See the full description on the dataset page: https://huggingface.co/datasets/budecosystem/earth_silver.Earth-Silver
Dataset Card for Earth-Silver
Dataset Details
Dataset Description
Earth-Silver is a question answering (QA) benchmark designed to evaluate the professional depth of large language models (LLMs) within the Earth sciences. It features more difficult and challenging questions compared to Earth-Iron, focusing on specialized knowledge within the domain. This dataset aims to assess a model's ability to handle complex inquiries requiring a deeper understanding of… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Silver.Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or… See the full description on the dataset page: https://huggingface.co/datasets/JasonChen91/Earth-Iron.Earth-Gold
Dataset Card for Earth-Gold
Dataset Details
Dataset Description
Earth-Gold is a novel open-ended dialogue dataset designed to evaluate the advanced scientific exploration capabilities of large language models (LLMs) within the Earth sciences. Unlike traditional question-answering formats, Earth-Gold assesses a model's ability to engage in multi-turn dialogues that simulate the process of scientific inquiry, including reflecting on existing methodologies and… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Gold.
