CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RongWei-at /3dfront_render_viewsimage1K<n<10K0 likes29k downloads5mo agoHugging Face02Dangindev /viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage. It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English. The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing, landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life, and transportation.visual-question-answering10K<n<100K8 likes16k downloads11mo agoHugging Face03RongWei-at /3dfront-render-viewsimage1K<n<10K0 likes7.6k downloads5mo agoHugging Face04abdoelsayed /dart_laser_vie dart_laser_vie — data Training data + mixes for the DART-LaSER BRIGHT 4B retriever (best model 38.74). Full guide: https://github.com/abdoelsayed2016/dart_laser_vie (DOCUMENTATION.md). Contents reason-embed-data-0928/ — the real per-domain training data (12 <domain>-formatted.jsonl): ReasonEmbed-format {prompt, query, pos, neg, train_group_size, batch_size}. This is what everything is built from. combo_rank_aug/ — the base mix (per-domain + aug.jsonl 81k +… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/dart_laser_vie.0 likes7.2k downloads2mo agoHugging Face05Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face06tmnam20 /Vietnamese-News Dataset Card for "VietnameseNewsparquet" More Information needed text1M<n<10M0 likes6k downloads3y agoHugging Face07hungnm /vietnamese-tokenized0 likes5.6k downloads1y agoHugging Face08QuangDuy /FineWeb2-vie-mds0 likes5.3k downloads11mo agoHugging Face09Youki2026 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples… See the full description on the dataset page: https://huggingface.co/datasets/Youki2026/VietSuperSpeech.0 likes5.2k downloads1mo agoHugging Face10SultanR /VieMix VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication. Subsets Subset Description quality_filtered Quality-filtered data before deduplication minhash_deduped Document-level MinHash deduplication matched Documents appearing in 2+ source datasets The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/VieMix.texttext-generation100M<n<1B0 likes4k downloads1mo agoHugging Face11VietMedTeam /MultiMed-WS Large-scale Joint Weakly-Supervised and Instruction Learning for Medical Speech Translation This repo is under development. Stay tuned! 1 likes3.6k downloads9mo agoHugging Face12thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes3.4k downloads7mo agoHugging Face13VietPhong /kitti-yolo11n-robustness-benchmark KITTI YOLO11n Robustness & Adversarial Benchmark Suite This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels. ?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n) Clean Baseline AP50: 0.3555 Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.tabularobject-detection100K<n<1M0 likes3.3k downloads27d agoHugging Face14diliash /partnetsim-1024-fixed-viewpointsimage1K<n<10K0 likes2.9k downloads2y agoHugging Face15supergoose /sentence_transformer_kmeans100_viewertext100K<n<1M0 likes2.5k downloads2y agoHugging Face16vietanh1999 /vietanh19997 likes2.5k downloads20d agoHugging Face17weikaih /habitat-views-20kimage10K<n<100K0 likes2.4k downloads1y agoHugging Face18minhnguyent546 /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M0 likes2.3k downloads2mo agoHugging Face19weikaih /ai2thor-random-views-20kimage10K<n<100K0 likes2.2k downloads1y agoHugging Face20mlinhbng /viet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage. It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English. The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing, landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life, and transportation.visual-question-answering10K<n<100K0 likes1.9k downloads10mo agoHugging Face21thanhnp-uel /vietnam-listed-companies-financial-statements Vietnamese Listed Companies Financial Data Overview This dataset provides standardized financial statement data for Vietnamese listed companies. The original financial information was collected from publicly available financial statements and annual reports published through official stock exchange portals and company websites. The source data has been transformed into a standardized long-format Parquet structure for research, educational, analytical, and… See the full description on the dataset page: https://huggingface.co/datasets/thanhnp-uel/vietnam-listed-companies-financial-statements.1 likes1.8k downloads1mo agoHugging Face22VietFinTabGroup /VietFinTab VietFinTab Dataset Overview VietFinTab is a large-scale Vietnamese financial table dataset designed for Table Structure Recognition (TSR), Document AI, and Optical Character Recognition (OCR) research. The dataset is collected from publicly available financial reports of 13 Vietnamese companies spanning the period from 2023 to 2025. The companies included in the dataset are: EVN Finance Joint Stock Company (EVF) FPT Corporation (FPT) Hoang Anh Gia Lai Joint Stock… See the full description on the dataset page: https://huggingface.co/datasets/VietFinTabGroup/VietFinTab.imageimage-feature-extraction10K<n<100K2 likes1.7k downloads2mo agoHugging Face23raymondt /viet_co_video_9imagen<1K0 likes1.4k downloads3h agoHugging Face24ntquang0410 /zh-vie_ecom 1688 zh-vi ecom Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới". Xem dữ liệu ở đâu data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây: bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi (bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.texttranslation10K<n<100K0 likes1.4k downloads3d agoHugging Face25XR-Bot0 /xense-lerobot-viewer-workbenck-log0 likes1.3k downloads2d agoHugging Face26II-Vietnam /Agentic-Multi-SWE-RLtext1K<n<10K0 likes1.3k downloads1y agoHugging Face27VTSNLP /vietnamese_curated_dataset Dataset Description Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator Developed by: Viettel Solutions Language: Vietnamese Details Please visit our Tech Blog post on NVIDIA's plog page for details. Link Data Collection We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include: The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.text10M<n<100M76 likes1.3k downloads2y agoHugging Face28VietAI /vi_pubmed Dataset Summary 20M Vietnamese PubMed biomedical abstracts translated by the state-of-the-art English-Vietnamese Translation project. The data has been used as unlabeled dataset for pretraining a Vietnamese Biomedical-domain Transformer model. image source: Enriching Biomedical Knowledge for Vietnamese Low-resource Language Through Large-Scale Translation Language English: Original biomedical abstracts from Pubmed Vietnamese: Synthetic abstract translated by a… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_pubmed.texttext-generation10M<n<100M26 likes1.3k downloads3y agoHugging Face29raymondt /viet_co_video_3imagen<1K0 likes1.3k downloads5h agoHugging Face30ncc2 /ff4d-viewer4d-scenes0 likes1.2k downloads9d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.