CoolFace
14 results

underthesea

undertheseanlp /UDD-1 UDD-1: Universal Dependency Dataset for Vietnamese Dataset Description Vietnamese Universal Dependency dataset created by Underthesea NLP. This dataset follows the Universal Dependencies annotation guidelines. Dataset Summary Language: Vietnamese (vi) Version: 1.1 Domains: ⚖️ Legal (Laws) + 📰 News Total Sentences: 20,000 Total Tokens: ~453,551 Annotation: Machine-generated using Underthesea NLP toolkit Data Sources Source Domain Sentences… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UDD-1.documenttoken-classification10K<n<100K0 likes237 downloads7mo agoHugging Faceundertheseanlp /UTS_WTKUTS_WTKtoken-classification0 likes225 downloads3y agoHugging Faceundertheseanlp /UTS_VLC Dataset Card for Vietnamese Legal Corpus (UTS_VLC) A curated corpus of Vietnamese Laws and Codes (Luật, Bộ luật) and the Constitution, maintained by Underthesea NLP. The flagship 2026 split is a verified in-force snapshot — every document is currently in force, de-duplicated, and validated against Vietnam's official legal database vbpl.vn. Dataset Details Dataset Description UTS_VLC contains the full text of Vietnamese legislation at the top of the… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS_VLC.texttext-generationn<1K2 likes197 downloads4mo agoHugging Faceundertheseanlp /UVW-2026 UVW 2026: Underthesea Vietnamese Wikipedia Dataset Dataset Description UVW 2026 (Underthesea Vietnamese Wikipedia) is a high-quality, cleaned dataset of Vietnamese Wikipedia articles enriched with Wikidata metadata. Designed for Vietnamese NLP research including language modeling, text generation, text classification, named entity recognition, and model pretraining. Key Features Clean text: Wikipedia markup, templates, references, and formatting… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UVW-2026.tabulartext-generation1M<n<10M1 likes148 downloads8mo agoHugging Faceundertheseanlp /UTS2017_Bank UTS2017_Bank Dataset Dataset Description Dataset Summary The UTS2017_Bank dataset is a comprehensive Vietnamese banking domain dataset containing customer feedback and reviews about banking services. It contains 2,471 annotated examples (1,977 train, 494 test) with both aspect labels and sentiment annotations. The dataset supports multiple NLP tasks including aspect classification, sentiment analysis, and aspect-based sentiment analysis in the Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/undertheseanlp/UTS2017_Bank.texttext-classification1K<n<10K2 likes90 downloads1y agoHugging Faceundertheseanlp /UTS_TextUTSTexttexttext-generation10K<n<100K0 likes87 downloads4y agoHugging Face