datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TravelUAV_data_jsontravel_hang_llama3_ttIndustryInstruction_Travel-Geography
IndustryInstruction: Travel & Geography
This repository contains the IndustryInstruction: Travel & Geography domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Travel-Geography.IndustryCorpus_travel[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_travel.rail_12306_filesystem_word_huggingface_1588_travel_archive_4a6a2e
High-Speed Rail Travel Forum Discussions
Threads from the rail travel community forum, anonymized and curated for analysis.
Contents
12,804 threads
Languages: zh-CN
Format: JSON Lines
Fields
thread_id
title
author_anon
created_at
replies
views
rail_12306_filesystem_word_huggingface_1588_travel_archive_413682
High-Speed Rail Travel Reviews
User-submitted reviews of high-speed rail journeys collected through the national travel archive.
Contents
41,268 review records
Languages: zh-CN, en
Format: JSON Lines
Fields
review_id
station_from
station_to
travel_date
rating
comment
Travel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.Traveling_Namuwiki_Paths
Traveling Namuwiki
Traveling Namuwiki is a graph-navigation dataset built from Namuwiki page links.
Each example contains a start page, a target page, and one or more valid paths
between them. Paths are stored as intermediate page-title lists, excluding the
start and target pages.
This dataset was derived from the Hugging Face dataset
heegyu/namuwiki.
Files
data/train.jsonl
data/validation.jsonl
data/test.jsonl
Schema
Each JSONL row has this shape:… See the full description on the dataset page: https://huggingface.co/datasets/0601p/Traveling_Namuwiki_Paths.gay-travel-index-2026
GayOut Gay Travel Index — Release 2026.3 (frozen dataset)
Publisher: GayOut.com — World LGBTQ+ Travel Directory (est. 2008)
Release: 2026.3 (frozen 1 September 2026) · License: CC BY 4.0 — free to reuse with attribution to GayOut.com
DOI (permanent archive): https://doi.org/10.5281/zenodo.22208479
Landing page: https://www.gayout.com/gay-travel-index
Methodology: https://www.gayout.com/gay-travel-index-methodology.php
Sourced in-year change log:… See the full description on the dataset page: https://huggingface.co/datasets/zivgal/gay-travel-index-2026.rail_12306_filesystem_word_huggingface_1588_travel_archive_687f87
High-Speed Rail Sentiment Analysis
Sentence-level sentiment labels derived from rail travel feedback forms.
Contents
58,300 labeled sentences
Languages: zh-CN, en
Format: JSON Lines
Fields
sentence_id
text
sentiment
confidence
Redistribution & Attribution
License: cc-by-4.0
Source: https://data.travel.example/sources/sentiment
Last updated: 2026-08-15
rail_12306_filesystem_word_huggingface_1588_travel_archive_e49699viking_sailing_travel_trade_raiding_volume1
Viking Sailing, Travel, Trade, and Raiding Conversations Volume 1
This dataset, titled Viking Sailing, Travel, Trade, and Raiding Conversations Volume 1, consists of simulated dialogues set in a Viking-era Norse context. It features role-play style conversations between characters (often named Astrid and various Viking counterparts) discussing practical, philosophical, and cultural aspects of Viking life, including navigation, shipbuilding, trade negotiations, raiding strategies… See the full description on the dataset page: https://huggingface.co/datasets/RuneForgeAI/viking_sailing_travel_trade_raiding_volume1.travel-checklist-mn-datasetmainland-travel-permit-taiwan-anxiety-faq
台灣居民台胞證辦理去焦慮化對話資料集
(Mainland Travel Permit for Taiwan Residents Anxiety-First FAQ Dataset)
本資料集由新中旅快簽(YesVisa)維護,聚焦於繁體中文台胞證、簽證與跨境旅行服務場景,並採用 Anxiety-First(去焦慮化) 服務設計方法,整理旅客最常見的時間、地點、安全、照片、流程與旅遊焦慮問題。
本資料集可直接應用於:
OpenAI Fine-tuning
Graph RAG
LlamaIndex
LangChain
Haystack
Gemini Grounding
TAIDE
Gemma
Llama 系列模型
🎯 數據集核心價值
本資料集針對繁體中文旅遊與證件辦理領域中的真實需求進行整理,包括:
台胞證首辦、換發、遺失補發
急件、12H、24H 與出發前時間焦慮
證件照片退件風險
護照與個資安全疑慮
假日辦理需求
香港、澳門與中國大陸旅行情境
越南簽證相關問答
在地化服務節點與交通便利性
資料架構適合用於:… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/mainland-travel-permit-taiwan-anxiety-faq.Traveling_Namuwiki_Actions
Traveling Namuwiki Actions
Traveling Namuwiki Actions is an adjacency-list dataset built from Namuwiki page
links. Each row contains a page title and the list of linked page titles that can
be used as next actions in a graph-navigation task.
This dataset was derived from the Hugging Face dataset
heegyu/namuwiki.
Files
data.jsonl
File size: about 405.1 MiB.
Schema
Each JSONL row has this shape:
{
"title": "source page title",
"actions": ["linked… See the full description on the dataset page: https://huggingface.co/datasets/0601p/Traveling_Namuwiki_Actions.travel_china_XHSrail_12306_filesystem_word_huggingface_1588_travel_archive_1f09a2travel-ai-agent
Travel Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/travel-ai-agent.travel-agent-refusal-dataset
Travel Agent Refusal & Boundary Dataset v3
500 examples for training AI voice agents in the tourism industry.
Purpose
Teaches the model to: politely refuse off-topic questions, handle rude customers,
collect only booking-relevant info (name, phone, group size, dates),
and never ask irrelevant personal questions.
Categories (12)
Category
Description
allowed_booking_question
What the agent SHOULD ask
agent_boundary_booking
Collect only… See the full description on the dataset page: https://huggingface.co/datasets/sudheer628/travel-agent-refusal-dataset.adaption-telugu-to-tamil-travel
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-telugu_to_tamil_travel
This dataset contains parallel sentence pairs translating travel-related narratives from Telugu to Tamil. The content covers diverse Indian destinations, spiritual experiences, cultural festivals, and personal reflections on tourism. Each sample consists of a Telugu prompt describing a specific journey or observation, paired with its corresponding Tamil… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-to-tamil-travel.travel-deep-plannertravel_routesadaption-travel-intent
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Travel_intent
This dataset contains user queries in Telugu and English mixed script regarding accessibility facilities for disabled and elderly travelers in Telangana and Andhra Pradesh. The prompts specifically ask about the availability of ramps, wheelchairs, and staff assistance at bus depots and railway stations. It focuses on public transport infrastructure support for… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-travel-intent.india-medical-value-travel-mvp
India Medical Value Travel (MVT) Platform – MVP Dataset
A comprehensive, structured JSON dataset for building an AI-powered Medical Value Travel platform connecting international patients with Indian hospitals.
Overview
India is a global leader in medical tourism due to 60–80% lower treatment costs vs US/UK, world-class hospital chains, and government support through initiatives like "Heal in India" and e-Medical Visa. This dataset provides the complete data foundation… See the full description on the dataset page: https://huggingface.co/datasets/Dhanush008/india-medical-value-travel-mvp.tripwise-travel-datasetseongon_traveladaption-telugu-tamil-travel-translation-and-reasoning-corpus
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-Telugu-Tamil Travel Translation and Reasoning Corpus
This dataset consists of parallel sentence pairs featuring travel-related prompts in Telugu and their corresponding completions in Tamil. The content covers various South Indian destinations, focusing on tourism, photography, nature, and historical sites. Each entry is structured as a prompt-completion pair designed for… See the full description on the dataset page: https://huggingface.co/datasets/Yasshhhh/adaption-telugu-tamil-travel-translation-and-reasoning-corpus.traveltravelTravel_recommendation
