datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
alia_tourism
📘 ALIA_TOURISM Dataset
The ALIA_TOURISM dataset is a multilingual resource designed for text generation within the tourism domain.
The source field indicates the data origin.
Documents in the train partition are restricted to sources with LLM-permissive licenses or explicit donations to the ALIA project,
whereas the test partition contains no such restrictions.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_tourism.code-tourisme
Code du tourisme, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-tourisme.rosetta-ko-tourism-synth-cpt
rosetta-ko-tourism-synth-cpt
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt (this repo)
continued-pretraining corpus (plain… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-cpt.rosetta-ko-tourism-synth-rlvr
rosetta-ko-tourism-synth-rlvr
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-rlvr.rosetta-ko-tourism-synth-dpo
rosetta-ko-tourism-synth-dpo
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo.rosetta-ko-tourism-synth-dpo-think
rosetta-ko-tourism-synth-dpo-think
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-dpo-think.rosetta-ko-tourism-synth-sft
rosetta-ko-tourism-synth-sft
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft.rosetta-ko-tourism-synth-sft-think
rosetta-ko-tourism-synth-sft-think
Korean tourism-domain data grounded in public tourism records — a cleaned CPT corpus plus source-grounded QA/preference/RLVR sets.
Synthetic data generated with the Qwen3.6-27B teacher model — part of the Rosetta-KO suite for the Rosetta Korean LLM (PoSTMEDIA).
Tourism Suite
Sibling datasets from the same pipeline (each a separate repo):
repo
format
rosetta-ko-tourism-synth-cpt
continued-pretraining corpus (plain text)… See the full description on the dataset page: https://huggingface.co/datasets/PoSTMEDIA/rosetta-ko-tourism-synth-sft-think.Personal-Tourism-Khmer
Personal Tourism Khmer
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) focused on generating realistic personas and profiles for tourism professionals, tour guides, and hospitality workers in Cambodia.
Project Vision
I hope this dataset helps your project succeed! Whether you are building a travel assistant chatbot, a tour recommendation system, or conducting research on… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Tourism-Khmer.indonesian-tourism
Wisata Indonesia 🏝️
Kumpulan 146 deskripsi tempat wisata Indonesia — dari Masjid Raya Baiturrahman (Aceh) sampai Raja Ampat (Papua Barat Daya). Deskripsi hangat ala panduan wisata, lengkap dengan waktu terbaik & estimasi tiket.
Kenapa dataset ini ada?
Dataset wisata Indonesia di HF cuma ada 1 (50 record, 5 downloads) — sangat lemah. Ini yang pertama lengkap: 146 tempat asli dari 30+ provinsi, 7 tipe (alam/pantai/gunung/budaya/kuliner/bahari/sejarah), 3 tingkat… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-tourism.morocco_tourism_darija
🇲🇦 Moroccan Tourism Darija Dialogues (maroc_tourism_darja)
Repository | datasets/maroc_tourism_darjaSize | 1 000 dialogues (≈ 14 k turns)Language | Darija (Moroccan Arabic)Topics | Transport · Accommodation · Culture · GastronomyLicense | CC‑BY‑4.0Format | JSONL (one dialogue per line)
1. Dataset Summary
maroc_tourism_darja is a synthetic, high‑quality collection of tourist‑guide conversations in Moroccan Darija.Each dialogue is a short two‑turn exchange in which a… See the full description on the dataset page: https://huggingface.co/datasets/oabai/morocco_tourism_darija.Tourism_Softpower
Tourism Softpower Thai Only
Thai-only text corpus for tourism and broad soft power experiments.
Columns
text: Thai text chunk
source: source project, mostly thwiki
doc_id: stable document id
title: source page title
page: chunk/page index starting from 1
Counts
Rows: 100,416
Source counts: {'thwiki': 100416}
License
Source text is from Wikimedia projects and is marked CC BY-SA 4.0. Keep attribution to source pages when using or… See the full description on the dataset page: https://huggingface.co/datasets/SPAISS6F1/Tourism_Softpower.
