datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CulturaY
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.Vietnamese-News
Dataset Card for "VietnameseNewsparquet"
More Information needed
VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.kitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.sentence_transformer_kmeans100_viewerTranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.VietFinTab
VietFinTab
Dataset Overview
VietFinTab is a large-scale Vietnamese financial table dataset designed for Table Structure Recognition (TSR), Document AI, and Optical Character Recognition (OCR) research.
The dataset is collected from publicly available financial reports of 13 Vietnamese companies spanning the period from 2023 to 2025.
The companies included in the dataset are:
EVN Finance Joint Stock Company (EVF)
FPT Corporation (FPT)
Hoang Anh Gia Lai Joint Stock… See the full description on the dataset page: https://huggingface.co/datasets/VietFinTabGroup/VietFinTab.zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.Agentic-Multi-SWE-RLvietnamese_curated_dataset
Dataset Description
Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator
Developed by: Viettel Solutions
Language: Vietnamese
Details
Please visit our Tech Blog post on NVIDIA's plog page for details. Link
Data Collection
We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include:
The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.Vietnamese-Traditional-Musicvi_pubmed
Dataset Summary
20M Vietnamese PubMed biomedical abstracts translated by the state-of-the-art English-Vietnamese Translation project. The data has been used as unlabeled dataset for pretraining a Vietnamese Biomedical-domain Transformer model.
image source: Enriching Biomedical Knowledge for Vietnamese Low-resource Language Through Large-Scale Translation
Language
English: Original biomedical abstracts from Pubmed
Vietnamese: Synthetic abstract translated by a… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_pubmed.VieMix
VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/VieMix.viet_vlsp
Dataset Card for "viet_vlsp"
More Information needed
VieMix
VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/VieMix.multi-view-bathroom-scene-understanding-camera-relocalization
Multi-View Bathroom Scene Understanding & Camera Relocalization
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is… See the full description on the dataset page: https://huggingface.co/datasets/physicl/multi-view-bathroom-scene-understanding-camera-relocalization.ViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.ai2thor-random-views-20kvietnamese-legal-documents
Vietnamese Legal Documents
A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.).
Curated by: Thịnh Ngô
Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.habitat-views-20kdolly-audio-1000h-vietnamese
Dolly-Audio: Vietnamese Multi-Speaker High-Quality Speech Corpus
Dataset Summary
Dolly-Audio is a large-scale, high-quality Vietnamese speech corpus created by the Dolly AI Team.
Inspired by Dolly, the world’s first cloned mammal, the project aims to advance research in Vietnamese speech synthesis, speech recognition, and voice modeling.
This release provides nearly 1,000 hours of professionally cleaned audio, featuring 152 speakers across different Vietnamese regions and… See the full description on the dataset page: https://huggingface.co/datasets/dolly-vn/dolly-audio-1000h-vietnamese.vietnews-datasetopen-web-math
Dataset Card for "open-web-math"
More Information needed
VieSpeaker-DatasetThis repo introduces the VieSpeaker dataset for Vietnamese Speaker Verification task. The paper corresponds to the dataset has been accepted at INTERSPEECH 2026: Preprint.
Please cite as:
@misc{pham2026viespeakerlargescalevietnamesespeaker,
title={VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency},
author={Viet Hoang Pham and Tran Trung Nguyen and Bao Thu Ho and Phuong Tuan Dat and Thi Thu Trang Nguyen},
year={2026}… See the full description on the dataset page: https://huggingface.co/datasets/hustep-lab/VieSpeaker-Dataset.UAV-Aerial-View-Battle-Tank-Detection-Dataset
Aerial UAV Perspective: Battle Tank Detection Dataset
Overview
This is an open-source synthetic dataset specifically engineered to train computer vision models in identifying main battle tanks (MBTs) and armored combat vehicles from tactical overhead and drone perspectives.
Obtaining real-world tactical aerial imagery for defense analytics is heavily restricted, operationally dangerous, or classified. This dataset addresses that critical data bottleneck by… See the full description on the dataset page: https://huggingface.co/datasets/Simuletic/UAV-Aerial-View-Battle-Tank-Detection-Dataset.vietnamese-corpus
Dataset Card for "vietnamese-corpus"
More Information needed
the_pile_openwebtext2
Dataset Card for "the_pile_openwebtext2"
More Information needed
Awesome_Spatial_VQA_Benchmarks_ViewSpatial-BenchVietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.
