datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3dfront_render_viewsviet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.3dfront-render-viewsdart_laser_vie
dart_laser_vie — data
Training data + mixes for the DART-LaSER BRIGHT 4B retriever (best model 38.74).
Full guide: https://github.com/abdoelsayed2016/dart_laser_vie (DOCUMENTATION.md).
Contents
reason-embed-data-0928/ — the real per-domain training data (12 <domain>-formatted.jsonl): ReasonEmbed-format
{prompt, query, pos, neg, train_group_size, batch_size}. This is what everything is built from.
combo_rank_aug/ — the base mix (per-domain + aug.jsonl 81k +… See the full description on the dataset page: https://huggingface.co/datasets/abdoelsayed/dart_laser_vie.CulturaY
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.Vietnamese-News
Dataset Card for "VietnameseNewsparquet"
More Information needed
vietnamese-tokenizedFineWeb2-vie-mdsVietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples… See the full description on the dataset page: https://huggingface.co/datasets/Youki2026/VietSuperSpeech.VieMix
VieMix (https://arxiv.org/abs/2512.18834) is a Vietnamese pretraining corpus built by combining six publicly available Vietnamese datasets, applying Vietnamese-specific quality filtering, and performing cross-dataset deduplication.
Subsets
Subset
Description
quality_filtered
Quality-filtered data before deduplication
minhash_deduped
Document-level MinHash deduplication
matched
Documents appearing in 2+ source datasets
The matched subset uses… See the full description on the dataset page: https://huggingface.co/datasets/SultanR/VieMix.MultiMed-WS
Large-scale Joint Weakly-Supervised and Instruction Learning for Medical Speech Translation
This repo is under development. Stay tuned!
VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.kitti-yolo11n-robustness-benchmark
KITTI YOLO11n Robustness & Adversarial Benchmark Suite
This dataset contains 649,425 benchmark samples evaluating the perception robustness of YOLO11n (Ultralytics YOLOv11 nano in original FP32 precision) on the official KITTI Object Detection train set (3,711 images) under 35 attack & corruption techniques across 5 severity levels.
?? Benchmark Leaderboard (mAP@0.5 Drop on YOLO11n)
Clean Baseline AP50: 0.3555
Evaluation Model: YOLO11n (Original weights:… See the full description on the dataset page: https://huggingface.co/datasets/VietPhong/kitti-yolo11n-robustness-benchmark.partnetsim-1024-fixed-viewpointssentence_transformer_kmeans100_viewervietanh1999habitat-views-20kTranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.ai2thor-random-views-20kviet-cultural-vqaVietnamese Cultural VQA Dataset is a comprehensive multimodal dataset focusing on Vietnamese cultural heritage.
It contains 28,505 images across 12 cultural categories with 119,012 question-answer pairs in Vietnamese and English.
The dataset covers diverse aspects of Vietnamese culture including architecture, cuisine, traditional clothing,
landscapes, festivals, folk culture, traditional games, sports, handicrafts, musical instruments, daily life,
and transportation.vietnam-listed-companies-financial-statements
Vietnamese Listed Companies Financial Data
Overview
This dataset provides standardized financial statement data for Vietnamese listed companies.
The original financial information was collected from publicly available financial statements and annual reports published through official stock exchange portals and company websites.
The source data has been transformed into a standardized long-format Parquet structure for research, educational, analytical, and… See the full description on the dataset page: https://huggingface.co/datasets/thanhnp-uel/vietnam-listed-companies-financial-statements.VietFinTab
VietFinTab
Dataset Overview
VietFinTab is a large-scale Vietnamese financial table dataset designed for Table Structure Recognition (TSR), Document AI, and Optical Character Recognition (OCR) research.
The dataset is collected from publicly available financial reports of 13 Vietnamese companies spanning the period from 2023 to 2025.
The companies included in the dataset are:
EVN Finance Joint Stock Company (EVF)
FPT Corporation (FPT)
Hoang Anh Gia Lai Joint Stock… See the full description on the dataset page: https://huggingface.co/datasets/VietFinTabGroup/VietFinTab.viet_co_video_9zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.xense-lerobot-viewer-workbenck-logAgentic-Multi-SWE-RLvietnamese_curated_dataset
Dataset Description
Vietnamese Curated Text Dataset. This dataset is collected from multiple open Vietnamese datasets, and curated with NeMo Curator
Developed by: Viettel Solutions
Language: Vietnamese
Details
Please visit our Tech Blog post on NVIDIA's plog page for details. Link
Data Collection
We utilize a combination of datasets that contain samples in Vietnamese language, ensuring a robust and representative text corpus. These datasets include:
The… See the full description on the dataset page: https://huggingface.co/datasets/VTSNLP/vietnamese_curated_dataset.vi_pubmed
Dataset Summary
20M Vietnamese PubMed biomedical abstracts translated by the state-of-the-art English-Vietnamese Translation project. The data has been used as unlabeled dataset for pretraining a Vietnamese Biomedical-domain Transformer model.
image source: Enriching Biomedical Knowledge for Vietnamese Low-resource Language Through Large-Scale Translation
Language
English: Original biomedical abstracts from Pubmed
Vietnamese: Synthetic abstract translated by a… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_pubmed.viet_co_video_3ff4d-viewer4d-scenes
