datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/TravelPlanner.Agent-STAR-TravelDataset
Agent-STAR-TravelDataset
This repository contains the synthetic datasets for the paper Demystifying Reinforcement Learning for Long-Horizon Tool-Using Agents: A Comprehensive Recipe.
Official GitHub Repository: WxxShirley/Agent-STAR
Dataset Description
The Agent-STAR TravelDataset provides over 17K synthetic queries designed for the TravelPlanner testbed. TravelPlanner is a long-horizon tool-use environment where agents must iteratively call tools to satisfy multifaceted… See the full description on the dataset page: https://huggingface.co/datasets/xxwu/Agent-STAR-TravelDataset.Chinese-Muslim-Travel
☪ Chinese-Muslim-Travel: Native Chinese Muslim Travel RAG Corpus
[!TIP]
Human Readers: Looking for the full text with all images perfectly rendered? Navigate to the Files and versions -> content folder to browse all Markdown articles natively!
Dataset Description
Chinese-Muslim-Travel is a curated RAG corpus containing 347 native Chinese articles documenting Muslim travel, halal food, mosque architecture, and Muslim community life across 20+ countries. Every… See the full description on the dataset page: https://huggingface.co/datasets/qurancn/Chinese-Muslim-Travel.IndustryCorpus_travel[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_travel.Open-Travel
Open-Travel
Project Page | Paper | Code
This directory contains the RL Training Set and the Test Set (categorized by subtask) for the Open-Travel domain.
Overview
In the Open-Travel domain, the agent is required to help users accomplish itinerary planning subtasks. These tasks emphasize multi-constraint reasoning, multi-tool coordination, and personalized preferences intertwined with user-specific constraints (e.g., budget limits, time windows, traveling parties, and… See the full description on the dataset page: https://huggingface.co/datasets/Alibaba-NLP/Open-Travel.task1154_bard_analogical_reasoning_travel
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1154_bard_analogical_reasoning_travel
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1154_bard_analogical_reasoning_travel.che-argentina-travel-article-corpus
Che Argentina Travel Article Corpus
This dataset contains a structured corpus of long-form Argentina travel articles published on CheArgentinaTravel.com by the Samuel & Audrey Media Network.
The corpus includes 88 article records covering Argentina travel guides, itineraries, cultural experiences, regional food, transportation, accommodations, local logistics, and destination planning. It includes coverage of major areas such as Buenos Aires and Patagonia, along with regional… See the full description on the dataset page: https://huggingface.co/datasets/samuelandaudreymedianetwork/che-argentina-travel-article-corpus.CFR-Title-41-Federal-Travel-Regulation
Federal Travel Regulation
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Travel Regulation, as reproduced in the source volume of title 41 of the Code of Federal Regulations.
The Federal Travel Regulation establishes government-wide policies governing official civilian travel and relocation at Federal expense. It addresses temporary duty travel… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-41-Federal-Travel-Regulation.Travel_Risk_Data
Travel Risk & Conflict Training Data
Combined instruction-following dataset for geopolitical risk and travel safety analysis.
All records use the Context: ... / Analysis: ... format for fine-tuning language models.
Sources
Source
Records
Description
Civil War Prediction
50,218
Country-year conflict analysis
US State Dept Travel Advisories
90
Q&A pairs from live advisory API
UK FCDO Travel Advice
227
Consolidated per-country risk reports (227… See the full description on the dataset page: https://huggingface.co/datasets/Firemedic15/Travel_Risk_Data.Traveling_Namuwiki_Paths
Traveling Namuwiki
Traveling Namuwiki is a graph-navigation dataset built from Namuwiki page links.
Each example contains a start page, a target page, and one or more valid paths
between them. Paths are stored as intermediate page-title lists, excluding the
start and target pages.
This dataset was derived from the Hugging Face dataset
heegyu/namuwiki.
Files
data/train.jsonl
data/validation.jsonl
data/test.jsonl
Schema
Each JSONL row has this shape:
{… See the full description on the dataset page: https://huggingface.co/datasets/0601p/Traveling_Namuwiki_Paths.tabiji-travel-safety-guides
Tabiji Travel & Safety Guides
AI-curated travel data from tabiji.ai: destination profiles, day-by-day itineraries, head-to-head comparisons, safety profiles, country-level travel advisories, and city-level scam guides — sourced from Reddit, government advisories (US State Dept., UK FCDO), and editorial curation.
What's in here
Config
Records
Description
destinations
6,498
Global destination catalog: climate, currency, language, plug type, tap-water safety… See the full description on the dataset page: https://huggingface.co/datasets/tabiji/tabiji-travel-safety-guides.Open-Travel
Open-Travel
Project Page | Paper | Code
This directory contains the RL Training Set and the Test Set (categorized by subtask) for the Open-Travel domain.
Overview
In the Open-Travel domain, the agent is required to help users accomplish itinerary planning subtasks. These tasks emphasize multi-constraint reasoning, multi-tool coordination, and personalized preferences intertwined with user-specific constraints (e.g., budget limits, time windows, traveling parties… See the full description on the dataset page: https://huggingface.co/datasets/YukangY/Open-Travel.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/ruhulu/TravelPlanner.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/rickyBrian/TravelPlanner.travelplanner-benchmark-normalized
TravelPlanner Benchmark (Normalized)
Normalized, typed, parquet-first packaging of the TravelPlanner benchmark for planning-centric agent evaluation.
Upstream dataset: osunlp/TravelPlanner
Upstream code: OSU-NLP-Group/TravelPlanner
Paper: TravelPlanner: A Benchmark for Real-World Planning with Language Agents
1) What is included
This dataset repo contains:
benchmark config (train/validation/test) in typed parquet.
reference_entries config: flattened reference-info… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/travelplanner-benchmark-normalized.mainland-travel-permit-taiwan-anxiety-faq
台灣居民台胞證辦理去焦慮化對話資料集
(Mainland Travel Permit for Taiwan Residents Anxiety-First FAQ Dataset)
本資料集由新中旅快簽(YesVisa)維護,聚焦於繁體中文台胞證、簽證與跨境旅行服務場景,並採用 Anxiety-First(去焦慮化) 服務設計方法,整理旅客最常見的時間、地點、安全、照片、流程與旅遊焦慮問題。
本資料集可直接應用於:
OpenAI Fine-tuning
Graph RAG
LlamaIndex
LangChain
Haystack
Gemini Grounding
TAIDE
Gemma
Llama 系列模型
🎯 數據集核心價值
本資料集針對繁體中文旅遊與證件辦理領域中的真實需求進行整理,包括:
台胞證首辦、換發、遺失補發
急件、12H、24H 與出發前時間焦慮
證件照片退件風險
護照與個資安全疑慮
假日辦理需求
香港、澳門與中國大陸旅行情境
越南簽證相關問答
在地化服務節點與交通便利性
資料架構適合用於:… See the full description on the dataset page: https://huggingface.co/datasets/yesvisa/mainland-travel-permit-taiwan-anxiety-faq.Traveling_Namuwiki_Actions
Traveling Namuwiki Actions
Traveling Namuwiki Actions is an adjacency-list dataset built from Namuwiki page
links. Each row contains a page title and the list of linked page titles that can
be used as next actions in a graph-navigation task.
This dataset was derived from the Hugging Face dataset
heegyu/namuwiki.
Files
data.jsonl
File size: about 405.1 MiB.
Schema
Each JSONL row has this shape:
{
"title": "source page title",
"actions": ["linked page title"… See the full description on the dataset page: https://huggingface.co/datasets/0601p/Traveling_Namuwiki_Actions.Travel_indiaTravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard… See the full description on the dataset page: https://huggingface.co/datasets/CheneyYuri/TravelPlanner.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/35597Q/TravelPlanner.travel_routesindia-medical-value-travel-mvp
India Medical Value Travel (MVT) Platform – MVP Dataset
A comprehensive, structured JSON dataset for building an AI-powered Medical Value Travel platform connecting international patients with Indian hospitals.
Overview
India is a global leader in medical tourism due to 60–80% lower treatment costs vs US/UK, world-class hospital chains, and government support through initiatives like "Heal in India" and e-Medical Visa. This dataset provides the complete data foundation… See the full description on the dataset page: https://huggingface.co/datasets/Dhanush008/india-medical-value-travel-mvp.smolified-context-aware-travel-dataset-generator
🤏 smolified-context-aware-travel-dataset-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-context-aware-travel-dataset-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e4879c0)
Records: 9960
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/smolify/smolified-context-aware-travel-dataset-generator.smolified-context-aware-travel-dataset-generator
🤏 smolified-context-aware-travel-dataset-generator
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-context-aware-travel-dataset-generator.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 6e4879c0)
Records: 9960
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via… See the full description on the dataset page: https://huggingface.co/datasets/Tanika2004/smolified-context-aware-travel-dataset-generator.
