datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or cover isolated… See the full description on the dataset page: https://huggingface.co/datasets/ai-earth/Earth-Iron.Earth-Iron
Dataset Card for Earth-Iron
Dataset Details
Dataset Description
Earth-Iron is a comprehensive question answering (QA) benchmark designed to evaluate the fundamental scientific exploration abilities of large language models (LLMs) within the Earth sciences. It features a substantial number of questions covering a wide range of topics and tasks crucial for basic understanding in this domain. This dataset aims to assess the foundational knowledge that underpins… See the full description on the dataset page: https://huggingface.co/datasets/PrismaX/Earth-Iron.Earth-Iron
(ICLR'26) EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
Updates/News 🆕
🚩 News (2026-01-26) EarthSE has been accepted by ICLR 2026 🎉.
Abstract
Advancements in Large Language Models (LLMs) drive interest in scientific applications, necessitating specialized benchmarks such as Earth science. Existing benchmarks either present a general science focus devoid of Earth science specificity or… See the full description on the dataset page: https://huggingface.co/datasets/JasonChen91/Earth-Iron.ptbr-irony-idioms-regionalism
Sotaques Digitais — Benchmark LLM Português Brasileiro
Dataset Summary
Sotaques Digitais é um benchmark de avaliação de competência pragmática e cultural para LLMs em português brasileiro. O dataset contém 90 cenários de teste distribuídos em três categorias linguísticas, construídos a partir de contextos reais (redes sociais, WhatsApp, atendimento ao cliente, avaliações de produto, ambiente de trabalho).
O benchmark foi desenvolvido para a pesquisa "Sotaques… See the full description on the dataset page: https://huggingface.co/datasets/ramondomiingos/ptbr-irony-idioms-regionalism.
