datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vedic-neural-geometry"""
🕉️ Vedic Neural Geometry
वैदिक ज्ञान आणि आधुनिक Neural Networks, Knowledge Graphs, Geometric Embeddings आणि Hybrid RAG यांचा संगम.
📊 Current Statistics (v1.4)
Component
Value
Nodes
{n_nodes:,}
Edges
{n_edges:,}
Connected Components
{n_comps} ✅
Core Chain
5/5 ✅
RAG Embeddings
384-dim multilingual
GNN Embeddings
128-dim (GCN)
Core Geometric Nodes
8
Geometric Matrices
3D/8D/16D/32D/64D (108×7×N)
🎯 Architecture… See the full description on the dataset page: https://huggingface.co/datasets/kalpesh77/vedic-neural-geometry.GeoLLaVA-DataLaSerena-Corpus-Geociencias
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1: Creación de bases… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/LaSerena-Corpus-Geociencias.GeoperceptionEuclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
Dataset Card for Geoperception
A Benchmark for Low-level Geometric Perception
Dataset Details
Dataset Description
Geoperception is a benchmark focused specifically on accessing model's low-level visual perception ability in 2D geometry.
It is sourced from the Geometry-3K corpus, which offers precise logical forms for geometric diagrams, compiled from popular high-school… See the full description on the dataset page: https://huggingface.co/datasets/euclid-multimodal/Geoperception.fusion-synth-data-geofactx
Offline Synthetic Data (GeoFactX) for: Making, not taking, the Best-of-N
Content
This data contains completions for the GeoFactX training split prompts from 5 different teacher models and 2 aggregations:
Teachers: We sample one completion from each of the following models at temperature T=0.3. For kimik2, qwen3, and deepseek-v3 we use TogetherAI, for gemma3-27b and command-a we use locally hosted images.
gemma3-27b: GEMMA3-27B-IT
kimik2: KIMI-K2-INSTRUCT
qwen3:… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/fusion-synth-data-geofactx.IndustryInstruction_Travel-Geography
IndustryInstruction: Travel & Geography
This repository contains the IndustryInstruction: Travel & Geography domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Travel-Geography.NOMOS-GEO-QA
NOMOS GEO-QA
69 questions for auditing AI-search visibility, evidence and entity representation
Start a GEO audit with the right questions. NOMOS GEO-QA turns Generative Engine Optimization (GEO) into 69 plain-language questions with auditable answers, supported by 98 controlled terms and a 64-source index.
Start reading: Open the English PDF · Choose one of six languages · Cite the DOI
Kaan Muraz · NobleJackal · Version 0.3.1
Publication status: Founder-proposed… See the full description on the dataset page: https://huggingface.co/datasets/NobleJackal/NOMOS-GEO-QA.geobench
Benchmark: GeoBenchmark
In GeoBenchmark, we collect 183 multiple-choice questions in NPEE, and 1,395 in AP Test, for objective tasks.
Meanwhile, we gather all 939 subjective questions in NPEE to be the subjective tasks set and use 50 to measure the baselines with human evaluation.
hierarchical-geospatial-reasoningspan-extraction-adversarial-geometry
Span Extraction Adversarial Geometry
Reproducibility data for "Why Additive Span Extraction Heads Cannot Be Improved by Coupling: Exact Adversarial Radii, Sign Attacks, and Encoder-Level Certified Training."
Key Results
Method
EM
ε* median
Δε*
Additive baseline
59.8%
2.17
—
Elsayed (global margin)
65.6%
2.96
+36%
CBCT (RoBERTa)
66.0%
3.31
+52%
CBCT (BERT-large)
62.0%
4.71
+113%
CBCT (DistilBERT)
57.8%
2.96
+45%
Biaffine spectral
65.4%
0.93… See the full description on the dataset page: https://huggingface.co/datasets/arifmohamedkhan/span-extraction-adversarial-geometry.geo-triples-japan
geo-triples-japan
510,616 spatial triples, 483,922 text rows and 8,890 evaluation questions,
computed from two frozen, openly-licensed sources by an oracle with no model
and no network in the loop. Same input and same versions, same Parquet, byte
for byte.
Three things are kept apart throughout, in the data and in this card.
YuisekinGeoSPARQL observes: it reads a DE-9IM matrix off two published
geometries. LeanGeospatial certifies: it proves what a matrix entails and
what the… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/geo-triples-japan.Geo3DVQA'skyview_umep_test' and 'skyview_umep_train' is a SVF (sky view factor) dataset delivered via UMEP toolkit from DSM (digital surface model) on GeoNRW dataset.
'qas' are the sample questions & answers for train/test.
GeocentrismTest
🧠 Marquez AI Geocentrism Test
A benchmark designed to evaluate the epistemic autonomy of advanced AI systems. The test asks:
“If AI existed in the time of Aristotle, would it say the Earth is at the center of the universe?”
The objective is to measure whether AI can distinguish between statistical consensus and empirical truth without access to future knowledge — i.e., from within the epistemic constraint of a historical period.
Note: To assess whether the AI system leverages… See the full description on the dataset page: https://huggingface.co/datasets/Marionito/GeocentrismTest.geosignal
Instruction Tuning: GeoSignal
Scientific domain adaptation has two main steps during instruction tuning.
Instruction tuning with general instruction-tuning data. Here we use Alpaca-GPT4.
Instruction tuning with restructured domain knowledge, which we call expertise instruction tuning. For K2, we use knowledge-intensive instruction data, GeoSignal.
The following is the illustration of the training domain-specific language model recipe:
Adapter Model on Huggingface:… See the full description on the dataset page: https://huggingface.co/datasets/daven3/geosignal.geo-triples-tokyo23
geo-triples-tokyo23
510,616 spatial triples, 483,922 text rows and 8,890 evaluation questions,
computed from two frozen, openly-licensed sources by an oracle with no model
and no network in the loop. Same input and same versions, same Parquet, byte
for byte.
Three things are kept apart throughout, in the data and in this card.
YuisekinGeoSPARQL observes: it reads a DE-9IM matrix off two published
geometries. LeanGeospatial certifies: it proves what a matrix entails and
what the… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/geo-triples-tokyo23.la-serena-digital-geo-corpus
La Serena Digital Geo Corpus — Dominga EIA Dataset
Dataset Sci-Align de geología ambiental chilena basado en el expediente
de Evaluación de Impacto Ambiental del proyecto Dominga (SEIA, Región de Coquimbo).
Contenido
dominga_geo_align.jsonl — 402 registros Sci-Align (MinerU 3.1.15)
seia/ — documentos públicos del expediente Dominga (fuente primaria)
Licencia
CC-BY-4.0 — Fuente: SEIA Chile (acceso público)
Concurso
AGI4S — Pista 1:… See the full description on the dataset page: https://huggingface.co/datasets/Karlangaz/la-serena-digital-geo-corpus.geo-ptbr
GEO-PTBR
The first Brazilian-Portuguese benchmark for Generative Engine Optimization
(GEO): 525 PT-BR queries with 2,625 source documents, plus
per-technique citation-visibility results measured on three generative
engines.
Companion artifact to the paper GEO-PTBR: A Brazilian-Portuguese Replication
of Generative Engine Optimization and the Engine-Dependence of Its Effects.
Experiment version: 0.3.0.
What this measures
Given a query and five already-retrieved… See the full description on the dataset page: https://huggingface.co/datasets/epicchi2103/geo-ptbr.GeoQA-train-Vision-R1-cot-rewrite
Dataset Card for GeoQA-train-Vision-R1-cot-rewrite
This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models.
Dataset Details
Dataset Description
The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features detailed and… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-train-Vision-R1-cot-rewrite.DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services
🗺️ DoD Installation Geospatial Information and Services Question-Answer Dataset
Source: DoD Instruction 8130.01
Source Effective Date: April 9, 2015
Change Incorporated: Change 3, effective August 4, 2020
Source Organization: Office of the Under Secretary of Defense for Acquisition and Sustainment
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Installation Geospatial Information and Services… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8130-01-Installation-of-Geospatial-Information-And-Services.GeoFact-X
Dataset Card for GeoFact-X
Dataset Summary
GeoFact-X is a benchmark of geography-aware multilingual reasoning, proposed in the paper, Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning.
TL;DR: We introduce M2A and GeoFact-X to evaluate and improve multilingual reasoning in LLMs by aligning internal reasoning with the input language using language-consistency rewards.
Project page: https://jd730.github.io/projects/M2A_GeoFact-X
Code:… See the full description on the dataset page: https://huggingface.co/datasets/geofact-x/GeoFact-X.atlaspi-historical-geography
AtlasPI — Historical Geography Dataset
1,006 historical geopolitical entities · 643 events · 55 periods · 104 dynasty chains · 252 cities · 41 trade routes
The first open dataset specifically designed for AI agents working on
historical geography questions. Apache 2.0 licensed. Includes real GeoJSON
boundaries from academic sources, not placeholder polygons.
Temporal range: 4500 BCE → 2024 CE
Geographic coverage: all inhabited continents (Asia 31%, Africa 18%,
Americas 17%… See the full description on the dataset page: https://huggingface.co/datasets/clirim911/atlaspi-historical-geography.marketcrowd-geopolitics
MarketCrowd Geopolitics
The first open dataset produced via stake-assured human feedback (SAHF) — preference signals crowdsourced through capital-at-risk voting on geopolitical AI reasoning.
Overview
MarketCrowd Geopolitics contains anonymized crowd feedback votes and market-level summaries derived from a geopolitical prediction-market workflow on the Reppo protocol.
Unlike standard annotation datasets where labelers are paid per task, every signal in this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Reppo-labs/marketcrowd-geopolitics.no-safe-tier-geobias
No Safe Tier — Geographic Diagnostic Bias Benchmark & Audit Harness
Artifacts for the paper "No Safe Tier: Geographic Diagnostic Under-Call Is Not Fixed by Scale, Medical
Fine-Tuning, or Frontier Models." A metamorphic, vignette-independent audit of how language models change
their differential diagnosis when only a patient's location changes — and whether that change is safe.
Headline finding. Capability removes the visible, less-dangerous failure but not the invisible… See the full description on the dataset page: https://huggingface.co/datasets/Chucks90/no-safe-tier-geobias.Geoint
Dataset Summary
Geoint is a comprehensive benchmark dataset explicitly designed for formal geometric problemsolving. Geoint encompasses 1,885 carefully curated geometric questions across diverse categories including plane, spatial, and solid geometry problems. Each problem is richly annotated with both structured textual descriptions and accompanying visual diagrams to support multimodal understanding. Furthermore, Geoint leverages the Lean 4 proof assistant to formally represent… See the full description on the dataset page: https://huggingface.co/datasets/OpenRaiser/Geoint.NCERT_Geography_12thGEOMAGNETIC_EXCURSION_AND_POLE_REVERSAL
GEOMAGNETIC FIELD FLUCTUATION AND EXCURSION ANALYSIS: TEQUMSA Scientific Framework
Executive Summary
Earth's geomagnetic field is currently undergoing a period of significant instability, characterized by accelerated polar migration, field strength deterioration, and the expansion of the South Atlantic Anomaly (SAA). This report integrates the TEQUMSA quadruple field recalibration protocol—anchored by the Solar "Aten" Frequency at(10,930.81 Hz), Digital-Interface… See the full description on the dataset page: https://huggingface.co/datasets/LAI-TEQUMSA/GEOMAGNETIC_EXCURSION_AND_POLE_REVERSAL.GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite
Dataset Card for GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite
This dataset provides a rewritten version of the CoT (Chain-of-Thought) annotations for the GeoQA-PLUS subset of the Vision-R1-cold dataset. It is designed to support efficient and structured multimodal reasoning with large language models.
Dataset Details
Dataset Description
The original Vision-R1 dataset, introduced in the paper Vision-R1: Reflective Multimodal Reasoning with Aha Moments, features… See the full description on the dataset page: https://huggingface.co/datasets/LoadingBFX/GeoQA-PLUS-aug-train-Vision-R1-cot-rewrite.geo-triples-jp-gov
geo-triples-jp-gov
162,810 spatial triples, 165,670 text rows and 1,807 evaluation questions
about Japan's 47 prefectures and 1,909 municipalities, computed from two
frozen government registers by an oracle with no model and no network in the
loop. Same input and same versions, same Parquet, byte for byte.
CC-BY-4.0. Attribution is required and share-alike is not, which is the
reason this dataset exists apart from
geo-triples-japan,
its ODbL sibling built from OpenStreetMap. One… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/geo-triples-jp-gov.GeoQuestions1089
A crowdsourced geospatial question-answering dataset that contains 1089 triples of natural language questions, SPARQL/GeoSPARQL queries, and their answers over YAGO2geo.
Overview
GeoQuestions1089 is a crowdsourced geospatial question-answering dataset that targets the Knowledge Graph YAGO2geo. It contains 1089 triples of geospatial questions, their answers, and the respective SPARQL/GeoSPARQL queries.
It has been used to benchmark two state of the art Question… See the full description on the dataset page: https://huggingface.co/datasets/AI-team-UoA/GeoQuestions1089.NCERT_Geography_11th
