toponym
Datasets
All datasets matching “toponym”toponym-adoption-data
Toponym adoption data
English-language adoption of Ukrainian vs Russian toponym transliterations
(Kyiv/Kiev, Chornobyl/Chernobyl, ...), 2010-2026.
Layout
Four stages, each a function of the previous. Only raw is expensive; everything
downstream is a recompute.
file
what it is
<source>_raw.parquet
exactly what the provider returned, nothing dropped
<source>_processed.parquet
cleaned and regex-matched; only records containing a spelling… See the full description on the dataset page: https://huggingface.co/datasets/KyivNotKiev/toponym-adoption-data.tatarstan-toponyms
Dataset Card for Toponyms of Tatarstan
Dataset Details
Dataset Description
A comprehensive dataset of 9,688 toponyms (place names) from Tatarstan and Tatar-populated regions, curated by TatarNLPWorld as part of the Tat2Vec project. Each entry provides detailed linguistic, geographical, and etymological information about Tatar and Russian place names. The dataset is specifically designed for linguistic research, onomastic studies, and training NLP… See the full description on the dataset page: https://huggingface.co/datasets/TatarNLPWorld/tatarstan-toponyms.kabyle-toponyms
Algeria French–Kabyle Toponym Corpus
A reproducible, georeferenced parallel corpus of Algerian place names extracted from OpenStreetMap, mapping name:fr to name:kab.
Description
This dataset contains every OpenStreetMap object in Algeria that is simultaneously tagged with both French (name:fr) and Kabyle (name:kab) names. It covers cities, towns, villages, hamlets, roads, administrative boundaries, and localized points of interest (POI).
The corpus is designed… See the full description on the dataset page: https://huggingface.co/datasets/taqbaylit/kabyle-toponyms.Toponymes_NETSCITY
[!NOTE]
Dataset origin: https://loterre-skosmos.loterre.fr/BVM/fr/
Description
Cette ressource contient 3403 entrées terminologiques regroupées en 418 collections (regroupements par agglomération urbaine et par département). Le périmètre géographique de couverture de cette ressource est la France. Le critère utilisé pour hiérarchiser entre les homonymes est celui de l’activité de publication scientifique : cette ressource a été créée en vue de travailler sur la répartition des… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/Toponymes_NETSCITY.pl-de-toponym-translation-outputs
Dataset structure
polish column is the Polish name of a place. It's the input for both the models, which is translated to produce finetuned_output and basemodel_output.
german column is the actual German name of a place.
Outputs of Polish→German Toponym Translation model.
V0 subset, "finetuned_output" column is outputs of this model → https://huggingface.co/DebasishDhal99/polish-to-german-toponym-model-opus-mt-pl-de (Fine-tuned specifically for Toponym translation)
V0 subset… See the full description on the dataset page: https://huggingface.co/datasets/DebasishDhal99/pl-de-toponym-translation-outputs.ce_ru_toponymsVarious toponyms in Chechen and Russian.
