CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6.3k downloads3y agoHugging Face02hungnm /vietnamese-tokenized0 likes4.2k downloads1y agoHugging Face03minhnguyent546 /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M0 likes2.2k downloads2mo agoHugging Face04thanhnp-uel /vietnam-listed-companies-financial-statements Vietnamese Listed Companies Financial Data Overview This dataset provides standardized financial statement data for Vietnamese listed companies. The original financial information was collected from publicly available financial statements and annual reports published through official stock exchange portals and company websites. The source data has been transformed into a standardized long-format Parquet structure for research, educational, analytical, and… See the full description on the dataset page: https://huggingface.co/datasets/thanhnp-uel/vietnam-listed-companies-financial-statements.1 likes1.8k downloads1mo agoHugging Face05II-Vietnam /Agentic-Multi-SWE-RLtext1K<n<10K0 likes1.3k downloads1y agoHugging Face06VTSNLP /vietnamese_curated_dataset Dataset Description Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator Developed by: Viettel Solutions Language: Vietnamese Details Please visit our Tech Blog post on NVIDIA's plog page for details. Link Data Collection We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include: The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.text10M<n<100M76 likes1.3k downloads2y agoHugging Face07NTQAI /Vietnamese-Traditional-Musicaudioaudio-classification100K<n<1M5 likes1.3k downloads8mo agoHugging Face08th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face09dolly-vn /dolly-audio-1000h-vietnamese Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus Dataset Summary Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.audio100K<n<1M59 likes1k downloads10mo agoHugging Face10tiennv /vietnamese-corpus Dataset Card for "vietnamese-corpus" More Information needed text10M<n<100M1 likes862 downloads3y agoHugging Face11nvidia /Nemotron-Personas-Vietnam Nemotron-Personas-Vietnam Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam A compound AI approach to personas grounded in real-world distributions Tổng quan (Overview) Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.imagetext-generation100K<n<1M62 likes794 downloads4mo agoHugging Face12doof-ferb /Vietnam-Celeb unofficial mirror of Vietnam-Celeb dataset official announcement: https://www.isca-archive.org/interspeech_2023/pham23b_interspeech.html https://github.com/Vietnam-Celeb/Vietnam-Celeb https://huggingface.co/datasets/hustep-lab/Vietnam-Celeb official download: Part 0: https://drive.google.com/file/d/1pMuT3DFzSwib7SVcRS8VkDwPuLTsemSG/view?usp=share_link Part 1: https://drive.google.com/file/d/1xayHt2HRqE1aJ4HvtUT40_9XlgvfDfRY/view?usp=share_linkPart 2:… See the full description on the dataset page: https://huggingface.co/datasets/doof-ferb/Vietnam-Celeb.audio10K<n<100K2 likes693 downloads1y agoHugging Face13BlossomsAI /vietnamese-corpus Vietnamese Combined Corpus Dataset Statistics Total documents: {<15M:,} Wikipedia articles: {>1.3M:,} News articles: {>13M:,} Text documents: {>200K:,} Processing Details Processed using Apache Spark Minimum document length: {10} characters Text cleaning applied: HTML/special character removal Whitespace normalization URL removal Empty document filtering Data Format Each document has: 'text': The document content 'source': Origin of the… See the full description on the dataset page: https://huggingface.co/datasets/BlossomsAI/vietnamese-corpus.text10M<n<100M8 likes636 downloads2y agoHugging Face14PhongGoldFish /youtube-center-vietnamese-asraudio100K<n<1M0 likes591 downloads2mo agoHugging Face15imthanhlv /laion2B-multi-Vietnamese-subset Dataset Card for LAION-2B-multi Vietnamese subset Dataset Summary Filter the Vietnamese subset from Laion2B-multi To get the subset of your language, check out this notebook text-to-image3 likes579 downloads3y agoHugging Face16tinixai /vietnam-real-estates 🏠 Tinix Vietnam Real Estate Listings (2025-2026) Tinix Vietnam Real Estate Listings 2025-2026 là bộ dữ liệu bất động sản Việt Nam quy mô lớn được thu thập và xử lý bởi TiniX AI, bao gồm đúng 3.500.744 tin đăng bán/cho thuê bất động sản từ tháng 6/2025 đến tháng 3/2026 sau khi đã qua bước lọc loại hình nghiêm ngặt (loại bỏ Nhà mặt phố, Nhà trong ngõ). Đây là tài nguyên phục vụ nghiên cứu về thị trường bất động sản, xây dựng mô hình định giá nhà, phân tích xu hướng thị trường tại… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnam-real-estates.tabulartabular-regression1M<n<10M35 likes571 downloads5mo agoHugging Face17minhnguyent546 /mmarco-vietnamese-split Dataset Summary This dataset contains Vietnamese split of the mMARCO dataset. Subset Split # Rows triples train 39,780,811 queries train 808,731 queries dev.full 101,093 queries dev 6980 collection collection 8,841,823 Note: triples contains (query, positive, negative) triples and can be used for training embeddings models. Citing @article{DBLP:journals/corr/abs-2108-13897, author = {Luiz Bonifacio and Israel Campiotti and… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/mmarco-vietnamese-split.texttext-ranking10M<n<100M0 likes562 downloads6mo agoHugging Face185CD-AI /Vietnamese-THUIR-T2Ranking-gg-translated 📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated 📝 Overview Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese. In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.tabulartext-retrieval100M<n<1B22 likes558 downloads1y agoHugging Face19Toan-Minh-Duong-Son /vietnamese-music-dataset Vietnamese Music Dataset A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset. Repository structure Path Contents Count audio/ MP3 audio files, named by YouTube video ID 4,820 images/ PNG cover thumbnails, same IDs as audio/ 4,820 data/ Parquet metadata files, one per collection session 31 Metadata schema Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.audio1K<n<10K1 likes489 downloads1mo agoHugging Face20hungnm /vietnamese-text-pretokenizetext10M<n<100M0 likes466 downloads1y agoHugging Face21fnooub /dolly-audio-1000h-vietnamese Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus Dataset Summary Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/fnooub/dolly-audio-1000h-vietnamese.audio100K<n<1M0 likes461 downloads6mo agoHugging Face22wanderer2k1 /vietnamese_curated_textstext10M<n<100M1 likes418 downloads2y agoHugging Face23tvu-vlinhd11 /vietnamese-packed-204810M<n<100M0 likes407 downloads10mo agoHugging Face24uitnlp /vietnamese_students_feedbackStudents’ feedback is a vital resource for the interdisciplinary research involving the combining of two different research fields between sentiment analysis and education. Vietnamese Students’ Feedback Corpus (UIT-VSFC) is the resource consists of over 16,000 sentences which are human-annotated with two different tasks: sentiment-based and topic-based classifications. To assess the quality of our corpus, we measure the annotator agreements and classification evaluation on the UIT-VSFC corpus. As a result, we obtained the inter-annotator agreement of sentiments and topics with more than over 91% and 71% respectively. In addition, we built the baseline model with the Maximum Entropy classifier and achieved approximately 88% of the sentiment F1-score and over 84% of the topic F1-score.texttext-classification10K<n<100K32 likes398 downloads4y agoHugging Face255CD-AI /Vietnamese-msMARCO-ggtranslatedtabular10M<n<100M4 likes386 downloads2y agoHugging Face26BlossomsAI /merged_vietnamese_instruction_datasettext1M<n<10M0 likes386 downloads1y agoHugging Face27ihbkaiser /dataset_vietnamesetext10M<n<100M0 likes382 downloads9mo agoHugging Face28quanghd96 /dolly-audio-1000h-vietnamese Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus Dataset Summary Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/quanghd96/dolly-audio-1000h-vietnamese.audio100K<n<1M0 likes373 downloads8mo agoHugging Face291TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes340 downloads2y agoHugging Face30EmilyNguyen235 /data-voice-vietnamese-restaurant-quan-oc Vietnamese Restaurant Order Speech This dataset contains Vietnamese spoken restaurant orders paired with text transcripts. Each utterance typically includes a table number, item quantities, dishes, drinks, and add-ons. Dataset Structure Files are split into subdirectories by filename-derived speaker_code to satisfy Hugging Face repository file-count limits: metadata.csv: one row per audio sample. audio/{speaker_code}/*.wav: mono WAV audio files.… See the full description on the dataset page: https://huggingface.co/datasets/EmilyNguyen235/data-voice-vietnamese-restaurant-quan-oc.audioautomatic-speech-recognition1K<n<10K1 likes330 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.