datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/TravelPlanner.Bitext-travel-llm-chatbot-training-dataset
Bitext - Travel Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the [Travel] sector can be easily achieved using our two-step approach to LLM Fine-Tuning. An overview of… See the full description on the dataset page: https://huggingface.co/datasets/bitext/Bitext-travel-llm-chatbot-training-dataset.vn-provinces-travel-agency-revenue
Vietnam provinces travel agency revenue
Travel agency (tour operator) revenue at current prices (billion VND). Coverage 2010, 2012-2024. Year 2024 is preliminary. Geographic labels are English (UN/GSO style ASCII romanization). Tables cover provinces, regions and national total where present. Province names follow ar_core.vn_geo (historical 63-province system).
Figures
Hero
Comparison
Color key
Files
provinces (854 rows)
data/provinces.csv… See the full description on the dataset page: https://huggingface.co/datasets/letrinhan/vn-provinces-travel-agency-revenue.travel-conversations-finetuning
UltraChat Dataset (HuggingFace)
For prototyping and model training, the project utilized the "UltraChat" dataset available from HuggingFace. This dataset comprises 10 JSONLines files, totaling 1.5 million conversations, each stored as lists of strings. The initial preprocessing involved standardizing the text data by converting it to lowercase, removing punctuation using regular expressions, and applying lemmatization with part-of-speech tagging. These steps ensured uniformity and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/travel-conversations-finetuning.travel-QAreddit-travel-QA-finetuningThis dataset was sourced through a series of daily requests to the Reddit API, aiming to capture diverse and real-time travel-related discussions from multiple travel-related subreddits, sourced from this list: https://www.reddit.com/r/travel/comments/1100hca/the_definitive_list_of_travel_subreddits_to_help/, along with subreddits for common travel destinations. Requested was top 100 of the year, this was executed only one, then hot 50 daily. Data aggregation involved concatenating and… See the full description on the dataset page: https://huggingface.co/datasets/soniawmeyer/reddit-travel-QA-finetuning.green-books-travel-guides
African American Travel Guides: The Green Book & Companion Directories (1930–1966)
A unified, structured dataset of 113,827 business and lodging listings transcribed from 50 volumes of mid-20th-century African American travel guides, spanning 1930–1966. During the Jim Crow era, these guides told Black travelers which hotels, restaurants, tourist homes, service stations, and other businesses would serve them safely. This dataset brings The Negro Motorist Green Book together with… See the full description on the dataset page: https://huggingface.co/datasets/hadro/green-books-travel-guides.CFR-Title-41-Federal-Travel-Regulation
Federal Travel Regulation
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on the Federal Travel Regulation, as reproduced in the source volume of title 41 of the Code of Federal Regulations.
The Federal Travel Regulation establishes government-wide policies governing official civilian travel and relocation at Federal expense. It addresses temporary duty travel… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/CFR-Title-41-Federal-Travel-Regulation.travel-multi-turn-chat-geminiTravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/ruhulu/TravelPlanner.travel-attractions-synthetic
Exploratory Data Analysis (EDA)
An exploratory data analysis was conducted to assess the structure, quality, and distributions within the dataset.
Dataset Overview
The dataset contains 9,998 records and 6 columns. All variables are categorical except for derived text-length metrics used for analysis. The dataset focuses exclusively on travel attractions to maintain consistency with the project objective.
Entity Type Distribution
All records in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/maorsoul/travel-attractions-synthetic.travel-reco-dataset
license: mit
task_categories:
text-classification
tags:
travel
reccomendations
tourism
pretty_name: Travel Recommendation Dataset"
size_categories:
1K<n<10K
---# Travel Recommendation Dataset (Synthetic)
Size: 1,200 rowsModality: Text tabular (CSV)Use case: Recommend 3 travel destinations based on user preferences.
Columns
id: unique identifier
name: destination (City, Country)
city, country: components of name
continent: {Africa, Asia, Europe, North America, South America… See the full description on the dataset page: https://huggingface.co/datasets/isaac2006/travel-reco-dataset.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/rickyBrian/TravelPlanner.south-america-travel-planning-index-2026
South America Travel Planning Index 2026
Version 1.1 is a source-linked travel-planning dataset covering all 12 sovereign South American countries. It combines editorial trip-length ranges, buffer-day guidance, gateways, route intensity, seasonality, signature experiences, official tourism links, and entry-check links with a separate registry of 44 official-source records.
Canonical record
Version DOI: https://doi.org/10.5281/zenodo.21765168
Concept DOI for all… See the full description on the dataset page: https://huggingface.co/datasets/visaadvisor1/south-america-travel-planning-index-2026.reddit-travel-qaTravel_indiaTravelDatatravel_itineraries_dataset_IndiaTravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard… See the full description on the dataset page: https://huggingface.co/datasets/CheneyYuri/TravelPlanner.ESSPL_TravelPolicy_Formattedsabay_sai_travel_data
Sabay Sai Travel Club Dataset
This dataset contains information about tours, stays, guides, and travel tips for Kazakhstan, curated by Sabay Sai Travel Club. It is intended for AI, travel recommendation systems, and general data exploration. Each entry includes details about the experience, location, description, URL,
TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/35597Q/TravelPlanner.travel-recommenderthis is for testing my travel-recommender based on RAG models.
travelSource
travel-multi-turn-chat-geminitravel_insurancetravel-wikipediatravel-prompt-lora-evaluation-datasetTravelDatatravel2
