CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01undertheseanlp /UDD-1 UDD-1: Universal Dependency Dataset for Vietnamese Dataset Description Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines. Dataset Summary Language: Vietnamese (vi) Version: 1.1 Domains: ⚖️ Legal (Laws) + 📰 News Total Sentences: 20,000 Total Tokens: ~453,551 Annotation: Machine-generated using Underthesea NLP toolkit Data Sources Source Domain Sentences… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-1.documenttoken-classification10K<n<100K0 likes242 downloads7mo agoHugging Face02undertheseanlp /UTS_VLC Dataset Card for Vietnamese Legal Corpus (UTS_VLC) A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution, maintained by Underthesea NLP. The flagship 2026 split is a verified in-force snapshot — every document is currently in force, de-duplicated, and validated against Vietnam's official legal database vbpl.vn. Dataset Details Dataset Description UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.texttext-generationn<1K2 likes186 downloads4mo agoHugging Face03undertheseanlp /UVW-2026 UVW 2026: Underthesea Vietnamese Wikipedia Dataset Dataset Description UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining. Key Features Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.tabulartext-generation1M<n<10M1 likes171 downloads8mo agoHugging Face04undertheseanlp /UTS_TextUTSTexttexttext-generation10K<n<100K0 likes91 downloads4y agoHugging Face05undertheseanlp /UTS2017_Bank UTS2017_Bank Dataset Dataset Description Dataset Summary The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.texttext-classification1K<n<10K2 likes89 downloads1y agoHugging Face06undertheseanlp /UVN-1 Vietnamese News Dataset A dataset of Vietnamese news articles collected from 6 major Vietnamese newspapers for NLP research. Dataset Summary This dataset contains 3,268 Vietnamese news articles covering various topics including politics, business, sports, entertainment, education, health, and technology. It is designed for Vietnamese NLP research tasks such as: Text classification (news categorization) Language modeling Text generation Named entity recognition Keyword… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVN-1.texttext-classification1K<n<10K0 likes81 downloads8mo agoHugging Face07undertheseanlp /uts2025_vietipa Vietnamese IPA Dataset A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications. Dataset Description Dataset Summary This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for: Text-to-speech systems development Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.audiotext-to-speechn<1K1 likes73 downloads1y agoHugging Face08undertheseanlp /UVB-v0.1 UVB - Underthesea Vietnamese Books Dataset A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research. Dataset Summary UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.tabulartext-generationn<1K0 likes65 downloads8mo agoHugging Face09undertheseanlp /UDD-v0.1 UDD-v0.1: Universal Dependency Dataset for Vietnamese Dataset Description Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines. Dataset Summary Language: Vietnamese (vi) Version: 0.1 Domain: ⚖️ Legal (Laws) Sentences: 3,000 Tokens: 64,814 Source: Vietnamese Legal Corpus (UTS_VLC) Annotation: Machine-generated using Underthesea NLP toolkit Validation: Passes all UD validation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-v0.1.texttoken-classification1K<n<10K0 likes48 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.