open-travel/japan-travel-mcp-data
Japan Travel MCP — Data The runtime data for the japan-travel-mcp Model Context Protocol server. Comprehensive Japanese travel data for AI agents, built from public official sources, covering all 47 prefectures and 1,938 local government entities. Code lives on GitHub: github.com/ookami0210/japan-travel-mcp Data lives here. The npm package downloads this dataset on first run. Why this dataset exists Japan's tourism information — created to reach the world — is… See the full description on the dataset page: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data.
Japan Travel MCP — Data
The runtime data for the japan-travel-mcp Model Context Protocol server. Comprehensive Japanese travel data for AI agents, built from public official sources, covering all 47 prefectures and 1,938 local government entities.
Code lives on GitHub: github.com/ookami0210/japan-travel-mcp Data lives here. The npm package downloads this dataset on first run.
Why this dataset exists
Japan's tourism information — created to reach the world — is published across thousands of municipal websites. Almost none of it is accessible to AI agents in a structured, multilingual form. This dataset fixes that gap.
— KJ Sunada, founder of KabuK Style
What's inside
translations/
descriptions_complete.jsonl # 13,394 attractions × 18 languages — rich
# 200-300 char tourism descriptions
multilingual_complete.jsonl # 13,961 attractions × 18 languages — names
multilingual_wikipedia.jsonl # 18-language names from Wikipedia sitelinks
jp_en.jsonl # JP → EN canonical name mapping
prefectures/ # 47 prefecture files: municipal-scrape spots
# + Wikidata attractions per prefecture
hotels/
master.json # ~20,000 accommodations (Wikidata + OSM merged)
r3/ # Official designation registries
maff_gi.json # 172 MAFF Geographical Indications (food / agri-products)
meti_densan.json # 231 METI-designated Traditional Crafts (Dentō Kōgeihin)
japan_heritage.json # 104 Japan Heritage stories (Nihon Isan)
bunka_intangible.json # 125 Important Intangible Cultural Properties
unesco_japan.json # 58 UNESCO ICH inscriptions for Japan
translations/
r3_translations.jsonl # 690 designation records × 18 languages
glossary/
seed_canonical.json # House style for translations
mlit_canonical.json # Japan Tourism Agency (MLIT) official terminology
wikipedia_multilingual.json # 18-language Wikipedia sitelinks (build-time)
_state/
wikidata_attractions.json # 41,404 Wikidata attractions, ja-anchored
municipalities.json # 1,938 municipalities + designated-city wards
municipality_centroids.json # JIS-coded centroid coordinates
official_urls.json # Resolved official tourism site URLsSource policy — official build-up only
This dataset only contains records that an authoritative public body has designated, scraped from that body's own publication. No editorial picks, no AI-curated lists, no UGC.
17 supported languages
English (en), Japanese (ja), Chinese Simplified (zh), Korean (ko), French (fr), Spanish (es), German (de), Italian (it), Portuguese (pt), Russian (ru), Thai (th), Vietnamese (vi), Indonesian (id), Malay (ms), Arabic (ar), Hindi (hi), Tagalog (tl).
The 17 were chosen to cover the JNTO inbound-tourism priority languages plus major source markets across Asia-Pacific, Europe, and the Middle East.
Coverage
All 47 prefectures are populated; every entity has descriptions in all 17 languages (no per-language gaps inside the 13,394-entity description set). The chart shows the per-prefecture entity count — the long tail outside Kyoto / Tokyo / Hokkaido is the actual point of this dataset.
13,394 attractions × 18 languages = 241,092 description cells. Traditional Chinese (zh-Hant, Taiwan lexicon) is derived from the quality-controlled Simplified layer via deterministic OpenCC conversion (s2twp); records added after 2026-08 are generated natively in both Chinese scripts. Plus 690 official-designation records (MAFF GI, METI crafts, Japan Heritage, Bunka-cho intangible records, UNESCO ICH) translated to the same 17 languages = 12,420 more cells. Plus 13,961 canonical names × 18 languages = 237,337 more cells.
Refresh cadence
The GitHub Actions cron in the code repo refreshes data on two tracks and re-publishes to this dataset:
Each domain is hit at most once per cycle.
How to use
Via the MCP server (recommended)
npm install -g japan-travel-mcp
japan-travel-mcp # downloads this dataset to ~/.japan-travel-mcp/data/ on first runThen add to your AI agent's MCP config (Claude Desktop, Cursor, etc.).
Direct download
from huggingface_hub import snapshot_download
local_dir = snapshot_download(
repo_id="open-travel/japan-travel-mcp-data",
repo_type="dataset",
)# Or via git-lfs:
git clone https://huggingface.co/datasets/open-travel/japan-travel-mcp-dataCitation
If you use this dataset in research or a product, please cite:
KJ Sunada, "Japan Travel MCP", 2026.
GitHub: https://github.com/ookami0210/japan-travel-mcp
HF dataset: https://huggingface.co/datasets/open-travel/japan-travel-mcp-data
License: CC BY 4.0License
Data: CC BY 4.0 — free to use including commercially, attribution required.
Code (separate repo): MIT.
Underlying source data carries its own licenses (see source-policy table above).
