datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
toponym-adoption-data
Toponym adoption data
English-language adoption of Ukrainian vs Russian toponym transliterations
(Kyiv/Kiev, Chornobyl/Chernobyl, ...), 2010-2026.
Layout
Four stages, each a function of the previous. Only raw is expensive; everything
downstream is a recompute.
file
what it is
<source>_raw.parquet
exactly what the provider returned, nothing dropped
<source>_processed.parquet
cleaned and regex-matched; only records containing a spelling… See the full description on the dataset page: https://huggingface.co/datasets/KyivNotKiev/toponym-adoption-data.tatarstan-toponyms
Dataset Card for Toponyms of Tatarstan
Dataset Details
Dataset Description
A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.kabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.Toponymes_NETSCITY
[!NOTE]
Dataset origin: https://loterre-skosmos.loterre.fr/BVM/fr/
Description
Cette ressource contient 3403 entrées terminologiques regroupées en 418 collections (regroupements par agglomération urbaine et par département). Le périmètre géographique de couverture de cette ressource est la France. Le critère utilisé pour hiérarchiser entre les homonymes est celui de l’activité de publication scientifique : cette ressource a été créée en vue de travailler sur la répartition des… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/Toponymes_NETSCITY.pl-de-toponym-translation-outputs
Dataset structure
polish column is the Polish name of a place. It's the input for both the models, which is translated to produce finetuned_output and basemodel_output.
german column is the actual German name of a place.
Outputs of Polish→German Toponym Translation model.
V0 subset, "finetuned_output" column is outputs of this model → https://huggingface.co/DebasishDhal99/polish-to-german-toponym-model-opus-mt-pl-de (Fine-tuned specifically for Toponym translation)
V0 subset… See the full description on the dataset page: https://huggingface.co/datasets/DebasishDhal99/pl-de-toponym-translation-outputs.ce_ru_toponymsVarious toponyms in Chechen and Russian.
ToponymExtractorDatasetskab-en-toponyms-sentences
English-Kabyle Parallel Corpus for Machine Translation
This dataset contains 32,024 grammatically flawless parallel sentence pairs mapping English to literary Kabyle (Taqbaylit kab).
This corpus was synthesized using a linguistically-informed morphosyntactic rule engine paired with clean OpenStreetMap toponym registries from boffire/kabyle-toponyms. It handles complex phonetic mutations natively, making it a state-of-the-art bootstrapping asset for fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kab-en-toponyms-sentences.tatarstan-toponyms-qa
Dataset Card for Tatarstan Toponyms QA Dataset
A question-answering dataset about toponyms (place names) of Tatarstan, containing 38,696 QA pairs in Russian and Tatar languages. The dataset covers various aspects of geographical names including their type, coordinates, sources, etymology, administrative region, location, and physical characteristics.
Dataset Details
Dataset Description
This dataset provides extractive question-answering pairs about toponyms… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms-qa.tajikistan-toponyms-corpus
🇹🇯 Tajikistan Toponyms Corpus
A comprehensive corpus of geographical names (toponyms) of Tajikistan, covering rivers, lakes, mountains, villages, districts, and other geographic features across all regions of the country.
📖 Description
This dataset contains over 5,559 unique toponyms collected from official administrative sources, cartographic materials, and local knowledge. Each entry includes the Tajik name, part-of-speech tag, semantic type, administrative region… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajikistan-toponyms-corpus.
