VLAI-AIVN/vietnamtourism-instruction-dataset
VietnamTourism LLaVA Instruction VietnamTourism LLaVA Instruction is a Vietnamese multimodal instruction-tuning dataset built from public tourism article images and metadata, then converted into LLaVA-style multi-turn conversations. The dataset is intended for research and internal experimentation on Vietnamese visual question answering, image-grounded dialogue, and tourism-domain multimodal assistants. Dataset Summary split samples train 5,978… See the full description on the dataset page: https://huggingface.co/datasets/VLAI-AIVN/vietnamtourism-instruction-dataset.
This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.
