CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01undertheseanlp /UDD-1 UDD-1: Universal Dependency Dataset for Vietnamese Dataset Description Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines. Dataset Summary Language: Vietnamese (vi) Version: 1.1 Domains: ⚖️ Legal (Laws) + 📰 News Total Sentences: 20,000 Total Tokens: ~453,551 Annotation: Machine-generated using Underthesea NLP toolkit Data Sources Source Domain Sentences… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-1.documenttoken-classification10K<n<100K0 likes237 downloads7mo agoHugging Face02undertheseanlp /UTS_WTKUTS_WTKtoken-classification0 likes225 downloads3y agoHugging Face03undertheseanlp /UTS_VLC Dataset Card for Vietnamese Legal Corpus (UTS_VLC) A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution, maintained by Underthesea NLP. The flagship 2026 split is a verified in-force snapshot — every document is currently in force, de-duplicated, and validated against Vietnam's official legal database vbpl.vn. Dataset Details Dataset Description UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.texttext-generationn<1K2 likes197 downloads4mo agoHugging Face04undertheseanlp /UVW-2026 UVW 2026: Underthesea Vietnamese Wikipedia Dataset Dataset Description UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining. Key Features Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.tabulartext-generation1M<n<10M1 likes148 downloads8mo agoHugging Face05undertheseanlp /UTS2017_Bank UTS2017_Bank Dataset Dataset Description Dataset Summary The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.texttext-classification1K<n<10K2 likes90 downloads1y agoHugging Face06undertheseanlp /UTS_TextUTSTexttexttext-generation10K<n<100K0 likes87 downloads4y agoHugging Face07undertheseanlp /UVN-1 Vietnamese News Dataset A dataset of Vietnamese news articles collected from 6 major Vietnamese newspapers for NLP research. Dataset Summary This dataset contains 3,268 Vietnamese news articles covering various topics including politics, business, sports, entertainment, education, health, and technology. It is designed for Vietnamese NLP research tasks such as: Text classification (news categorization) Language modeling Text generation Named entity recognition Keyword… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVN-1.texttext-classification1K<n<10K0 likes86 downloads8mo agoHugging Face08undertheseanlp /uts2025_vietipa Vietnamese IPA Dataset A comprehensive Vietnamese IPA (International Phonetic Alphabet) dataset with word pronunciations and MP3 audio files for text-to-speech and pronunciation learning applications. Dataset Description Dataset Summary This dataset contains 50 common Vietnamese words with their IPA (International Phonetic Alphabet) transcriptions and corresponding audio files. It's designed for: Text-to-speech systems development Vietnamese pronunciation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/uts2025_vietipa.audiotext-to-speechn<1K1 likes65 downloads1y agoHugging Face09undertheseanlp /UVB-v0.1 UVB - Underthesea Vietnamese Books Dataset A collection of 447 Vietnamese books with full text content and Goodreads metadata for NLP research. Dataset Summary UVB (Underthesea Vietnamese Books) is a dataset containing 447 Vietnamese books with full text content, mapped to Goodreads for metadata enrichment including genres, ratings, and publication years. The dataset is designed for Vietnamese language model training, text generation, and other NLP tasks.… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVB-v0.1.tabulartext-generationn<1K0 likes64 downloads8mo agoHugging Face10undertheseanlp /UTS_DictionaryUTSTexttext-generation10K<n<100K2 likes52 downloads7mo agoHugging Face11undertheseanlp /UDD-v0.1 UDD-v0.1: Universal Dependency Dataset for Vietnamese Dataset Description Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines. Dataset Summary Language: Vietnamese (vi) Version: 0.1 Domain: ⚖️ Legal (Laws) Sentences: 3,000 Tokens: 64,814 Source: Vietnamese Legal Corpus (UTS_VLC) Annotation: Machine-generated using Underthesea NLP toolkit Validation: Passes all UD validation… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-v0.1.texttoken-classification1K<n<10K0 likes52 downloads8mo agoHugging Face12undertheseanlp /UVD-1UVD-1text-generation10K<n<100K0 likes24 downloads7mo agoHugging Face13undertheseanlp /sentence-segmentation-1 Sentence Segmentation Test set for evaluating and improving Vietnamese sentence boundary detection (sent_tokenize) in underthesea. Problem The current PunktSentenceTokenizer in underthesea fails on several Vietnamese-specific patterns, primarily in legal text where article titles are merged with sentence bodies without punctuation boundaries. Current Results Category Total Correct Accuracy title_content_merge 38 0 0.0% repeated_title 13 0 0.0%… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/sentence-segmentation-1.0 likes18 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.