datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.GeoGrid_Bench
GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
We present GeoGrid-Bench, a benchmark designed to evaluate the ability of foundation models to understand geo-spatial data in the grid structure. Geo-spatial datasets pose distinct challenges due to their dense numerical values, strong spatial and temporal dependencies, and unique multimodal representations including tabular data, heatmaps, and geographic visualizations. To assess how foundation… See the full description on the dataset page: https://huggingface.co/datasets/bowen-upenn/GeoGrid_Bench.geoparser-benchmark-results
Geoparser benchmark results
Geoparsing pipelines from geoparser scored on English and
multilingual corpora, run on Grid'5000 (one Tesla T4).
Benchmarks
Benchmark
Languages
Docs
Toponyms
Source
GeoVirus
en
229
2167
WikiNews articles on epidemics (Gritta et al., 2018).
HIPE-2020
de, en, fr
129
1516
Historical Swiss, Luxembourgish and American newspapers, OCR.
NewsEye
de, fi, fr, sv
77
1772
Historical European newspapers, OCR (HIPE-2022).
TopRes19th… See the full description on the dataset page: https://huggingface.co/datasets/NoeFlandre/geoparser-benchmark-results.georgia-layoffs-warn-act-notices-daily
Georgia WARN Act layoff notices — every filing we hold since 2023, one CSV, rebuilt daily
284 Georgia WARN notices — every one this dataset holds, back to 2023 — free to download in full: no paywalled years, no login, no account · most recent notice filed 2026-09-14
· state source last checked 2026-09-22T12:26Z · official source: Technical College System of Georgia (WorkSource Georgia) — WARN notices.
Georgia employers must file a WARN Act notice with the state before a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/georgia-layoffs-warn-act-notices-daily.geom_drugs
GEOM: Molecular Conformations (Drugs Subset)
Note: This is a mirrored and specifically preprocessed version of the GEOM dataset (Drugs subset), originally created by Simon Axelrod and Rafael Gómez-Bombarelli. All credit for the original conformational sampling and DFT calculations goes to the original authors. This repository exists to guarantee availability and exact reproducibility for downstream machine learning projects.
Dataset Description
The Geometric Ensemble… See the full description on the dataset page: https://huggingface.co/datasets/raulsofia/geom_drugs.GeoGPT-CoT-QA
GeoGPT-CoT-QA Dataset: A Large-scale Geoscience Chain-of-Thought QA Dataset for Supervised Fine-Tuning of LLMs
1. Dataset Description
We introduce GeoGPT-CoT-QA Dataset, a large-scale synthetic question–answer (QA) corpus enriched with chain-of-thought (CoT) reasoning traces, developed to support supervised fine-tuning (SFT) of geoscience reasoning models. The GeoGPT-R1-Preview is specifically fine-tuned using this dataset to enhance its geoscience reasoning capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-CoT-QA.GEOMeta
GEOMeta: Large-Scale Human Bulk RNA-seq Dataset with Curated Metadata
Overview
GEOMeta is a large-scale human bulk RNA-seq dataset derived from the ARCHS4 resource, containing transcript abundance profiles for ~474,000 samples across training, test, and held-out test splits. Each sample is annotated with standardized metadata including sex, organ system, disease category, age group, and experimental setting.
Dataset Construction
Transcript abundance profiles… See the full description on the dataset page: https://huggingface.co/datasets/binchenlab/GEOMeta.GeoGPT-QA
GeoGPT-QA Dataset: A Large-scale Geoscience QA Dataset for Supervised Fine-tuning of LLMs
1. Dataset Description
We introduce GeoGPT-QA Dataset, a large-scale synthetic question–answer (QA) corpus developed to support supervised fine-tuning (SFT) of geoscience foundation models.
The dataset is derived from open-access geoscience publications distributed under the CC BY license. Using an automated data synthesis pipeline, we generated professional QA pairs from article… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA.turkish-disaster-news-geonlp
Turkish Disaster News GeoNLP Dataset
Dataset Summary
This dataset was created and submitted as part of the Uncharted Data Challenge by Adaption. The LLM-enhanced instruction pairs (turkish_earthquake_news.csv) were generated using Adaptive Data by Adaption — an AI-powered data adaptation platform.
The first open-source Turkish-language disaster news dataset with district-level geocoding, humanitarian category labels, and multi-dimensional damage classification.… See the full description on the dataset page: https://huggingface.co/datasets/FatmaElik/turkish-disaster-news-geonlp.geotree
Tree Monitoring Dataset - Bangladesh
This dataset is prepared for training object detection and segmentation models to monitor tree canopies in Bangladesh.
Dataset Structure
sentinel2/: Raw Sentinel-2 Level-2A imagery for Bandarban, Rangamati, Sylhet, and Gazipur districts.
deepforest/: DeepForest annotations and tree crown samples.
zenodo/: Reference training datasets.
selvabox/: Tree canopy labels and annotations.
global_forest_change/: Hansen Global Forest… See the full description on the dataset page: https://huggingface.co/datasets/the-shoaib2/geotree.geolayers
Geolayers-Data
-->
This dataset card contains usage instructions and metadata for all data-products released with our paper:Using Multiple Input Modalities can Improve Data-Efficiency and O.O.D. Generalization for ML with Satellite Imagery. We release 3 modified versions of 3 benchmark datasets spanning land-cover segmentation, tree-cover regression, and multi-label land-cover classification tasks. These datasets are augmented with auxiliary, geographic inputs. A full list of… See the full description on the dataset page: https://huggingface.co/datasets/arjunrao2000/geolayers.geometric_optics_physical_consistency_eval
Geometric Optics – Physical Consistency Evaluation Dataset
This dataset contains visual failure cases in geometric optics for multimodal image generation models.
Scope
The dataset focuses on physical and geometric inconsistencies related to:
mirror reflections (law of reflection)
refraction and dispersion in prisms
light ray direction consistency
shadow direction vs light source
camera–object–light spatial coherence
Motivation
Current image generation models… See the full description on the dataset page: https://huggingface.co/datasets/8Planetterraforming/geometric_optics_physical_consistency_eval.geoagent-results
GeoAgent — raw experiment results
Raw per-run output from GeoAgent, a benchmark that has vision language models
play a GeoGuessr-style game on Google Street View: the model looks at a
panorama, may rotate and move around, and finally guesses where it is.
Code: https://github.com/sohambuilds/geoagent
This repository holds the trajectories the paper's numbers were computed from.
The location manifests, the aggregated metric tables and the figures stay in
the code repo, under… See the full description on the dataset page: https://huggingface.co/datasets/Kartikeyatrivedi/geoagent-results.GeoNuclearData
Predicting Whether a Nuclear Reactor Is PWR or BWR
Dataset Overview
This project is based on the Geo Nuclear Data dataset from Kaggle.The dataset contains information about nuclear reactors around the world, including geographic and technical features such as:
Country
Plant name
Latitude
Longitude
Reactor type
Capacity
Status
Construction and operational dates
The original dataset was cleaned and filtered in order to focus only on two reactor types:
PWR
BWR… See the full description on the dataset page: https://huggingface.co/datasets/jonblustein/GeoNuclearData.sec-edgar-geographic-revenue-breakdowns
US S&P 500 Companies Geographic Revenue Exposure (SEC EDGAR)
This dataset contains a comprehensive, reconciled, and audited map of the geographic and regional revenue breakdowns for major US-listed corporations (including S&P 500 companies). The data was extracted directly from corporate 10-K filings submitted to the US Securities and Exchange Commission (SEC) EDGAR system.
By reconciling structured SEC XBRL segment dimensions with unstructured HTML R-file disclosures (using the… See the full description on the dataset page: https://huggingface.co/datasets/Metricshour/sec-edgar-geographic-revenue-breakdowns.visible-pack-dcfc32
visible-pack-dcfc32
Synthetic products test data: 51 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/GeorgeGarcia/visible-pack-dcfc32.state-of-geo-2026
The 2026 State of Generative Engine Optimization
860 AI answers scored across 85 B2B software companies in 61 categories, with every cited source traced: 5,160 citations. All citation URLs were string-matched against reddit.com, quora.com and stackoverflow.com: zero hits. Vendor-authored pages were 77.4% of citations. Volume II, The Absence Ladder, classifies the 616 answers where a company was absent and finds the shape of absence shifts with existing visibility (chi-square… See the full description on the dataset page: https://huggingface.co/datasets/Broadcastwell/state-of-geo-2026.ozone_training_data
Ozone Training Data
Dataset Summary
The Ozone training dataset contains information about ozone levels, temperature, wind speed, pressure, and other related atmospheric variables across various geographic locations and time periods. It includes detailed daily observations from multiple data sources for comprehensive environmental and air quality analysis. Geographic coordinates (latitude and longitude) and timestamps (month, day, and hour) provide spatial and temporal… See the full description on the dataset page: https://huggingface.co/datasets/Geoweaver/ozone_training_data.NCERT_Geography_12thpositivequotation-public-domain-quotes
PositiveQuotation Source-Verified Public Domain Quotes
This small dataset contains exactly 30 English proverbs matched to numbered entries in a public-domain U.S. source. It is designed for examples, prototypes, educational projects, and applications that need compact quotation records with auditable provenance.
Homepage: https://positivequotation.com/public-domain-quotes
API documentation: https://positivequotation.com/developers/public-domain-quotes-api
Live JSON API:… See the full description on the dataset page: https://huggingface.co/datasets/geosfero/positivequotation-public-domain-quotes.GeoGallery
GeoGallery: 3D geological fields
Dataset description
This dataset consists of geological grids of four different geological sedimentation environments:
Barrier Island
Tidal
Shelf
Wave Delta
Each grid has 4 different types of fields:
facies
net-to-gross (NTG)
porosity
permeability
The dataset was created using the framework proposed by K.J. Webber and L.C. van Geuns and then elaborated by N. Tyler and R. J. Finley
There are 20004 total geological grids, 5001… See the full description on the dataset page: https://huggingface.co/datasets/Klimkou/GeoGallery.WESAD_raw_dataWESAD_raw_data.csv contains 62 features (and four additional ones: Time, subject_id, SSSQ, condition).
Data is split into 4 classes: baseline, stress, amusement, meditation.
Data was recorded from 15 different subjects.
geomlamaData from the paper GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models
(The 125 raw instances are combined with the 500 augmented instances (without "Corresponding Questions"))
GitHub
georges-1913-normalization
Normalized Georges 1913
Description
This dataset was created as part of the Burchard's Dekret Digital project (www.burchards-dekret-digital.de),
funded by the Academy of Sciences and Literature | Mainz.
It is based on 55,000 lemmata from Karl Georges, Ausführliches lateinisch-deutsches Handwörterbuch, Hannover 1913 (Georges 1913)
and was developed to train models for normalization tasks in the context of medieval Latin.
The dataset consists of approximately 5 million… See the full description on the dataset page: https://huggingface.co/datasets/mschonhardt/georges-1913-normalization.NCERT_Geography_11thsentiment.csvdumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.stability-recovery-geometry-v0.1
What this dataset does
This dataset tests whether a model can evaluate recovery geometry.
The task is simple:
Given a scenario and a recovery-geometry claim, predict whether the claim is supported.
Core stability idea
Recovery is not merely the return of performance.
Recovery geometry evaluates whether the system is moving toward a healthier basin of operation.
Favorable recovery geometry typically includes:
restoration of function
restoration of margin
reduction of… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/stability-recovery-geometry-v0.1.each-expression-f760b3
each-expression-f760b3
Synthetic products test data: 32 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/GeorgeMartinez/each-expression-f760b3.EmeraldData
Dataset Card for EmeraldData
EmeraldData is a large-scale, semi-synthetic benchmark dataset containing 620 instances designed for greenwashing detection. It was created to address the absence of large-scale, annotated, real-world benchmarks with verified instances of greenwashing.
Dataset Details
Dataset Description
Existing research on greenwashing is limited by the lack of large-scale annotated real-world benchmarks. This scarcity is due to vague greenwashing… See the full description on the dataset page: https://huggingface.co/datasets/geoka/EmeraldData.
