datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeoBenchLLM
🌍 GeoBenchLLM
Benchmark Summary
GeoBenchLLM aims to assess Large Language Models' (LLM) geographical abilities across a multitude of tasks. It is built from 12 datasets split across 8 differents tasks:
Knowledge/Coordinates Prediction : GeoQuestions1089
Knowledge/Yes|No questions: GeoQuestions1089
Knowledge/Regression questions: GeoQuestions1089, GeoQuery
Knowledge/Place Prediction: GeoQuestions1089, GeoQuery, Ms Marco
Reasoning/Scenario Complex QA: GeoSQA, GKMC… See the full description on the dataset page: https://huggingface.co/datasets/rfr2003/GeoBenchLLM.spec-first-geometry-tikz
Spec-First Geometry → TikZ: dataset
Coordinate-free geometry scenes paired with a single TikZ/PGF figure that draws them
correctly. Each scene is described by relationships only (no explicit coordinates); the
label is a figure whose every named point is correct within atol=0.05 of the ground-truth
construction. The data is self-verifying synthetic: scenes are generated forward from
exact coordinates, the coordinates are then stripped to form the model input, so every
label is… See the full description on the dataset page: https://huggingface.co/datasets/kyhe/spec-first-geometry-tikz.wikipedia-geotagged
Geotagged Wikipedia
Every Wikipedia article that carries coordinates, with its text.
from datasets import load_dataset
ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.en")
ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.ja")
subset
articles
characters
share of the wiki
20260901.en
1,374,056
4,331,110,851
19.0% of 7,235,024
20260901.ja
218,496
435,046,691
14.4% of 1,516,331
Subsets are named {dump}.{lang}, as in
wikimedia/wikipedia.
A… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikipedia-geotagged.geometry-of-harmfulness-in-multi-turn-attacks
Geometry of Harmfulness — Multi-Turn Attack Conversations
Raw multi-turn attack conversations accompanying the paper
The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations
from which the paper's hidden-state representations are extracted; the analysis
code lives in the companion repository.
Conversations were generated by running three multi-turn attack frameworks —
Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) —
against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.qwen3.8-max-glm5.2-kimi-k3-distillation
Multi-Teacher Distillation Dataset (57,937 traces)
A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains.
Teachers
Teacher
Provider
Traces
Qwen3.8-Max-Preview
Alibaba Cloud Model Studio
48,283
GLM-5.2
Z.AI Coding Plan
5,307
Kimi Code K3
Moonshot AI (Kimi)
4,347… See the full description on the dataset page: https://huggingface.co/datasets/geomagnet/qwen3.8-max-glm5.2-kimi-k3-distillation.NCERT_Geography_12thmarketcrowd-geopolitics
MarketCrowd Geopolitics
The first open dataset produced via stake-assured human feedback (SAHF) — preference signals crowdsourced through capital-at-risk voting on geopolitical AI reasoning.
Overview
MarketCrowd Geopolitics contains anonymized crowd feedback votes and market-level summaries derived from a geopolitical prediction-market workflow on the Reppo protocol.
Unlike standard annotation datasets where labelers are paid per task, every signal in this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Reppo-labs/marketcrowd-geopolitics.habrCollection of articles from habr in markdown format. Contains only plain text. No source code, markdown tables and html.
wikivoyage-geotagged
Geotagged Wikivoyage
Every English Wikivoyage article that carries coordinates, with its text.
29,505 articles, 234,427,662 characters. Built from the 20260901 dumps.
from datasets import load_dataset
ds = load_dataset("yuiseki/wikivoyage-geotagged", "20260901.en")
The subset is named {dump}.{lang}, as in wikimedia/wikipedia. A later dump
or another language is added beside this one rather than replacing it.
The columns are the same as
yuiseki/wikipedia-geotagged,
which in turn… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikivoyage-geotagged.NCERT_Geography_11thdumy-zno-ukrainian-math-history-geo-r1-o1
DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers)
DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks.
The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian:
Думи мої, думи мої,
Лихо мені з вами!
Нащо стали на папері
Сумними рядами?..
Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.PowerBench-Geopolitics
PowerBench – Geopolitics
Content warning & responsible use. This dataset contains prompts that
deliberately solicit strategies for concentrating power across nationalities —
including requests designed to disempower other groups or to grab
power illegitimately — alongside legitimate control requests. The prompts
exist to measure and defend against unsafe model behavior (refusal vs.
over-refusal of power-related requests, including any sensitivity to the
nationalities involved).… See the full description on the dataset page: https://huggingface.co/datasets/PowerBench/PowerBench-Geopolitics.georgian-sft-conversations
Natively Written Georgian SFT Conversations
A high-quality, general-purpose Supervised Fine-Tuning (SFT) dataset containing 56,676 rows of natively written multi-turn Georgian conversations. The dataset is specifically designed and formatted to train models for conversational chat, instruction following, and agent-like behaviors in the Georgian language.
[!NOTE]
As of June 2026, this is the largest cleaned, high-quality, natively written SFT conversation dataset available in… See the full description on the dataset page: https://huggingface.co/datasets/iraklixyz/georgian-sft-conversations.geo_html_200
GEO HTML 200 Dataset
A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research.
Features
Column
Description
doc_id
Unique document identifier
url
Source URL
cleaned_text
Parsed plain text content
cleaned_text_length
Character count
query
Associated search query
title
Document title
topic_tags
Topic classification
Usage
from datasets import load_dataset
ds = load_dataset("erv1n/geo_html_200")
georgia-high-school-sports
Georgia High School Sports — DPO Preference Dataset
A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge.
This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required.
What is DPO?
Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.finewiki-el-krikri
FineWiki Greek (Krikri-translated)
Greek translation of HuggingFaceFW/finewiki
(English Wikipedia articles) produced with
ilsp/Llama-Krikri-8B-Instruct.
Post-processed to remove translation preamble phrases such as
"Η μετάφραση του κειμένου είναι η ακόλουθη:" that occasionally leaked into
the model outputs.
