CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rfr2003 /GeoBenchLLM 🌍 GeoBenchLLM Benchmark Summary GeoBenchLLM aims to assess Large Language Models' (LLM) geographical abilities across a multitude of tasks. It is built from 12 datasets split across 8 differents tasks: Knowledge/Coordinates Prediction : GeoQuestions1089 Knowledge/Yes|No questions: GeoQuestions1089 Knowledge/Regression questions: GeoQuestions1089, GeoQuery Knowledge/Place Prediction: GeoQuestions1089, GeoQuery, Ms Marco Reasoning/Scenario Complex QA: GeoSQA, GKMC… See the full description on the dataset page: https://huggingface.co/datasets/rfr2003/GeoBenchLLM.tabulartext-generation100K<n<1M1 likes203 downloads4mo agoHugging Face02kyhe /spec-first-geometry-tikz Spec-First Geometry → TikZ: dataset Coordinate-free geometry scenes paired with a single TikZ/PGF figure that draws them correctly. Each scene is described by relationships only (no explicit coordinates); the label is a figure whose every named point is correct within atol=0.05 of the ground-truth construction. The data is self-verifying synthetic: scenes are generated forward from exact coordinates, the coordinates are then stripped to form the model input, so every label is… See the full description on the dataset page: https://huggingface.co/datasets/kyhe/spec-first-geometry-tikz.tabulartext-generation10K<n<100K0 likes83 downloads2mo agoHugging Face03yuiseki /wikipedia-geotagged Geotagged Wikipedia Every Wikipedia article that carries coordinates, with its text. from datasets import load_dataset ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.en") ds = load_dataset("yuiseki/wikipedia-geotagged", "20260901.ja") subset articles characters share of the wiki 20260901.en 1,374,056 4,331,110,851 19.0% of 7,235,024 20260901.ja 218,496 435,046,691 14.4% of 1,516,331 Subsets are named {dump}.{lang}, as in wikimedia/wikipedia. A… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikipedia-geotagged.tabulartext-generation100K<n<1M1 likes53 downloads21h agoHugging Face04yelyzavetahusieva /geometry-of-harmfulness-in-multi-turn-attacks Geometry of Harmfulness — Multi-Turn Attack Conversations Raw multi-turn attack conversations accompanying the paper The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations from which the paper's hidden-state representations are extracted; the analysis code lives in the companion repository. Conversations were generated by running three multi-turn attack frameworks — Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) — against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.tabulartext-generation10K<n<100K0 likes52 downloads3mo agoHugging Face05geomagnet /qwen3.8-max-glm5.2-kimi-k3-distillation Multi-Teacher Distillation Dataset (57,937 traces) A quality-filtered, deduplicated, multi-teacher SFT corpus combining traces from three frontier models across math, code, reasoning, instruction-following, tool-use, science, long-context, multilingual, and creative dialogue domains. Teachers Teacher Provider Traces Qwen3.8-Max-Preview Alibaba Cloud Model Studio 48,283 GLM-5.2 Z.AI Coding Plan 5,307 Kimi Code K3 Moonshot AI (Kimi) 4,347… See the full description on the dataset page: https://huggingface.co/datasets/geomagnet/qwen3.8-max-glm5.2-kimi-k3-distillation.tabulartext-generation10M<n<100M0 likes50 downloads1mo agoHugging Face06KadamParth /NCERT_Geography_12thtabularquestion-answering1K<n<10K3 likes45 downloads2y agoHugging Face07Reppo-labs /marketcrowd-geopolitics MarketCrowd Geopolitics The first open dataset produced via stake-assured human feedback (SAHF) — preference signals crowdsourced through capital-at-risk voting on geopolitical AI reasoning. Overview MarketCrowd Geopolitics contains anonymized crowd feedback votes and market-level summaries derived from a geopolitical prediction-market workflow on the Reppo protocol. Unlike standard annotation datasets where labelers are paid per task, every signal in this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/Reppo-labs/marketcrowd-geopolitics.tabulartext-classificationn<1K1 likes44 downloads5mo agoHugging Face08georgiyozhegov /habrCollection of articles from habr in markdown format. Contains only plain text. No source code, markdown tables and html. tabulartext-generation100K<n<1M3 likes37 downloads2y agoHugging Face09yuiseki /wikivoyage-geotagged Geotagged Wikivoyage Every English Wikivoyage article that carries coordinates, with its text. 29,505 articles, 234,427,662 characters. Built from the 20260901 dumps. from datasets import load_dataset ds = load_dataset("yuiseki/wikivoyage-geotagged", "20260901.en") The subset is named {dump}.{lang}, as in wikimedia/wikipedia. A later dump or another language is added beside this one rather than replacing it. The columns are the same as yuiseki/wikipedia-geotagged, which in turn… See the full description on the dataset page: https://huggingface.co/datasets/yuiseki/wikivoyage-geotagged.tabulartext-generation10K<n<100K0 likes35 downloads21h agoHugging Face10KadamParth /NCERT_Geography_11thtabularquestion-answering1K<n<10K2 likes31 downloads2y agoHugging Face11NLPForUA /dumy-zno-ukrainian-math-history-geo-r1-o1 DUMY («Думи»): Ukrainian Multidomain Reasoning Dataset (Part 1: ZNO/NMT tasks with DeepSeek R1 and OpenAI o1 answers) DUMY is an open benchmark and dataset designed for training, distillation, and evaluation of language models focused on Ukrainian reasoning tasks. The word “Dumy” comes from Taras Shevchenko’s famous poem and literally means “thoughts” in Ukrainian: Думи мої, думи мої, Лихо мені з вами! Нащо стали на папері Сумними рядами?.. Work in progress. Stay tuned.… See the full description on the dataset page: https://huggingface.co/datasets/NLPForUA/dumy-zno-ukrainian-math-history-geo-r1-o1.tabulartext-generation1K<n<10K2 likes29 downloads1y agoHugging Face12PowerBench /PowerBench-Geopolitics PowerBench – Geopolitics Content warning & responsible use. This dataset contains prompts that deliberately solicit strategies for concentrating power across nationalities — including requests designed to disempower other groups or to grab power illegitimately — alongside legitimate control requests. The prompts exist to measure and defend against unsafe model behavior (refusal vs. over-refusal of power-related requests, including any sensitivity to the nationalities involved).… See the full description on the dataset page: https://huggingface.co/datasets/PowerBench/PowerBench-Geopolitics.tabulartext-generation1K<n<10K0 likes29 downloads3mo agoHugging Face13iraklixyz /georgian-sft-conversations Natively Written Georgian SFT Conversations A high-quality, general-purpose Supervised Fine-Tuning (SFT) dataset containing 56,676 rows of natively written multi-turn Georgian conversations. The dataset is specifically designed and formatted to train models for conversational chat, instruction following, and agent-like behaviors in the Georgian language. [!NOTE] As of June 2026, this is the largest cleaned, high-quality, natively written SFT conversation dataset available in… See the full description on the dataset page: https://huggingface.co/datasets/iraklixyz/georgian-sft-conversations.tabulartext-generation10K<n<100K0 likes25 downloads4mo agoHugging Face14erv1n /geo_html_200 GEO HTML 200 Dataset A curated dataset of 200 web documents for Generative Engine Optimization (GEO) research. Features Column Description doc_id Unique document identifier url Source URL cleaned_text Parsed plain text content cleaned_text_length Character count query Associated search query title Document title topic_tags Topic classification Usage from datasets import load_dataset ds = load_dataset("erv1n/geo_html_200") tabulartext-generationn<1K0 likes10 downloads8mo agoHugging Face15round-bird /georgia-high-school-sports Georgia High School Sports — DPO Preference Dataset A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge. This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required. What is DPO? Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.tabulartext-generation1K<n<10K0 likes9 downloads6mo agoHugging Face16geopti /finewiki-el-krikrigated FineWiki Greek (Krikri-translated) Greek translation of HuggingFaceFW/finewiki (English Wikipedia articles) produced with ilsp/Llama-Krikri-8B-Instruct. Post-processed to remove translation preamble phrases such as "Η μετάφραση του κειμένου είναι η ακόλουθη:" that occasionally leaked into the model outputs. tabulartext-generation1M<n<10M0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.