datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UDD-1
UDD-1: Universal Dependency Dataset for Vietnamese
Dataset Description
Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines.
Dataset Summary
Language: Vietnamese (vi)
Version: 1.1
Domains: ⚖️ Legal (Laws) + 📰 News
Total Sentences: 20,000
Total Tokens: ~453,551
Annotation: Machine-generated using Underthesea NLP toolkit
Data Sources
Source
Domain
Sentences… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-1.UTS_WTKUTS_WTKUTS_VLC
Dataset Card for Vietnamese Legal Corpus (UTS_VLC)
A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution,
maintained by Underthesea NLP. The flagship 2026 split is a
verified in-force snapshot — every document is currently in force, de-duplicated, and validated
against Vietnam's official legal database vbpl.vn.
Dataset Details
Dataset Description
UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.UVW-2026
UVW 2026: Underthesea Vietnamese Wikipedia Dataset
Dataset Description
UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining.
Key Features
Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.UTS2017_Bank
UTS2017_Bank Dataset
Dataset Description
Dataset Summary
The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.UTS_TextUTSTextUVN-1
Vietnamese News Dataset
A dataset of Vietnamese news articles collected from 6 major Vietnamese newspapers for NLP research.
Dataset Summary
This dataset contains 3,268 Vietnamese news articles covering various topics including politics, business, sports, entertainment, education, health, and technology. It is designed for Vietnamese NLP research tasks such as:
Text classification (news categorization)
Language modeling
Text generation
Named entity recognition
Keyword… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVN-1.uts2025_vietipa
Vietnamese IPA Dataset
A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications.
Dataset Description
Dataset Summary
This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for:
Text-to-speech systems development
Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.UVB-v0.1
UVB - Underthesea Vietnamese Books Dataset
A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research.
Dataset Summary
UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.UTS_DictionaryUTSTextUDD-v0.1
UDD-v0.1: Universal Dependency Dataset for Vietnamese
Dataset Description
Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines.
Dataset Summary
Language: Vietnamese (vi)
Version: 0.1
Domain: ⚖️ Legal (Laws)
Sentences: 3,000
Tokens: 64,814
Source: Vietnamese Legal Corpus (UTS_VLC)
Annotation: Machine-generated using Underthesea NLP toolkit
Validation: Passes all UD validation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-v0.1.UVD-1UVD-1sentence-segmentation-1
Sentence Segmentation
Test set for evaluating and improving Vietnamese sentence boundary detection (sent_tokenize) in underthesea.
Problem
The current PunktSentenceTokenizer in underthesea fails on several Vietnamese-specific patterns, primarily in legal text where article titles are merged with sentence bodies without punctuation boundaries.
Current Results
Category
Total
Correct
Accuracy
title_content_merge
38
0
0.0%
repeated_title
13
0
0.0%… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/sentence-segmentation-1.
