datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UVW-2026
UVW 2026: Underthesea Vietnamese Wikipedia Dataset
Dataset Description
UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining.
Key Features
Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.
