CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face02tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6.3k downloads3y agoHugging Face03thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes3.3k downloads7mo agoHugging Face04VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.1k downloads27d agoHugging Face05supergoose /sentence_transformer_kmeans100_viewertext100K<n<1M0 likes2.6k downloads2y agoHugging Face06minhnguyent546 /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M0 likes2.2k downloads2mo agoHugging Face07VietFinTabGroup /VietFinTab VietFinTab Dataset Overview VietFinTab is a large-scale Vietnamese financial table dataset designed for Table Structure Recognition (TSR), Document AI, and Optical Character Recognition (OCR) research. The dataset is collected from publicly available financial reports of 13 Vietnamese companies spanning the period from 2023 to 2025. The companies included in the dataset are: EVN Finance Joint Stock Company (EVF) FPT Corporation (FPT) Hoang Anh Gia Lai Joint Stock… See the full description on the dataset page: https://huggingface.co/datasets/VietFinTabGroup/VietFinTab.imageimage-feature-extraction10K<n<100K2 likes1.6k downloads2mo agoHugging Face08ntquang0410 /zh-vie_ecom 1688 zh-vi ecom Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới". Xem dữ liệu ở đâu data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây: bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi (bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.texttranslation10K<n<100K0 likes1.4k downloads4d agoHugging Face09II-Vietnam /Agentic-Multi-SWE-RLtext1K<n<10K0 likes1.3k downloads1y agoHugging Face10VTSNLP /vietnamese_curated_dataset Dataset Description Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator Developed by: Viettel Solutions Language: Vietnamese Details Please visit our Tech Blog post on NVIDIA's plog page for details. Link Data Collection We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include: The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.text10M<n<100M76 likes1.3k downloads2y agoHugging Face11NTQAI /Vietnamese-Traditional-Musicaudioaudio-classification100K<n<1M5 likes1.3k downloads8mo agoHugging Face12VietAI /vi_pubmed Dataset Summary 20M Vietnamese PubMed biomedical abstracts translated by the state-of-the-art English-Vietnamese Translation project. The data has been used as unlabeled dataset for pretraining a Vietnamese Biomedical-domain Transformer model. image source: Enriching Biomedical Knowledge for Vietnamese Low-resource Language Through Large-Scale Translation Language English: Original biomedical abstracts from Pubmed Vietnamese: Synthetic abstract translated by a… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_pubmed.texttext-generation10M<n<100M26 likes1.3k downloads3y agoHugging Face13SultanR /VieMix VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/VieMix.texttext-generation100M<n<1B0 likes1.2k downloads1mo agoHugging Face14quocanh34 /viet_vlsp Dataset Card for "viet_vlsp" More Information needed audio100K<n<1M2 likes1.2k downloads3y agoHugging Face15AdaMLLab /VieMix VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/VieMix.texttext-generation100M<n<1B2 likes1.2k downloads5mo agoHugging Face16physicl /multi-view-bathroom-scene-understanding-camera-relocalization Multi-View Bathroom Scene Understanding & Camera Relocalization Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.imagen<1K0 likes1.2k downloads3mo agoHugging Face17lidingm /ViewSpatial-Bench ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models Dataset Description We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.imagevisual-question-answering1K<n<10K23 likes1.2k downloads3mo agoHugging Face18weikaih /ai2thor-random-views-20kimage10K<n<100K0 likes1.2k downloads1y agoHugging Face19th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face20weikaih /habitat-views-20kimage10K<n<100K0 likes1.1k downloads1y agoHugging Face21dolly-vn /dolly-audio-1000h-vietnamese Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus Dataset Summary Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team. Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling. This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.audio100K<n<1M59 likes1k downloads10mo agoHugging Face22truongpdd /vietnews-datasettext10M<n<100M5 likes961 downloads4y agoHugging Face23vietgpt /open-web-math Dataset Card for "open-web-math" More Information needed text1M<n<10M0 likes930 downloads3y agoHugging Face24hustep-lab /VieSpeaker-DatasetThis repo introduces the VieSpeaker dataset for Vietnamese Speaker Verification task. The paper corresponds to the dataset has been accepted at INTERSPEECH 2026: Preprint. Please cite as: @misc{pham2026viespeakerlargescalevietnamesespeaker, title={VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency}, author={Viet Hoang Pham and Tran Trung Nguyen and Bao Thu Ho and Phuong Tuan Dat and Thi Thu Trang Nguyen}, year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/hustep-lab/VieSpeaker-Dataset.text100K<n<1M2 likes909 downloads2mo agoHugging Face25Simuletic /UAV-Aerial-View-Battle-Tank-Detection-Dataset Aerial UAV Perspective: Battle Tank Detection Dataset Overview This is an open-source synthetic dataset specifically engineered to train computer vision models in identifying main battle tanks (MBTs) and armored combat vehicles from tactical overhead and drone perspectives. Obtaining real-world tactical aerial imagery for defense analytics is heavily restricted, operationally dangerous, or classified. This dataset addresses that critical data bottleneck by… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/UAV-Aerial-View-Battle-Tank-Detection-Dataset.imagen<1K1 likes898 downloads3mo agoHugging Face26tiennv /vietnamese-corpus Dataset Card for "vietnamese-corpus" More Information needed text10M<n<100M1 likes862 downloads3y agoHugging Face27vietgpt /the_pile_openwebtext2 Dataset Card for "the_pile_openwebtext2" More Information needed text10M<n<100M5 likes846 downloads3y agoHugging Face28LLDDSS /Awesome_Spatial_VQA_Benchmarks_ViewSpatial-Benchimage1K<n<10K0 likes807 downloads1y agoHugging Face29b00l26 /VietPET-RoI VietPET-RoI VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding boxes. It is intended for medical multimodal research, report generation, visual question answering, and ROI grounding. Research use only. This dataset is not intended for diagnosis, treatment decisions, or direct patient care. Summary Split Patients CT/PET region pairs ROIs Train 160 480 1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.tabularimage-to-text1K<n<10K2 likes794 downloads2mo agoHugging Face30nvidia /Nemotron-Personas-Vietnam Nemotron-Personas-Vietnam Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam A compound AI approach to personas grounded in real-world distributions Tổng quan (Overview) Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.imagetext-generation100K<n<1M62 likes794 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.