datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.vietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5 over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm
Languages: Vietnamese (vi) answers · English (en) reasoning
Modality: image + text → text
Records: 544,795… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-ImageText-Reasoning.laion-2b-vietnamese-subset
Dataset Card for "laion-2b-vietnamese-subset"
More Information needed
TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.Traffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/star092304/Traffic-sign-detection-VietNam.vietnam_traffic_signTranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.Vietnamese-yfcc15m-OpenAICLIPvietnam_traffic_signTraffic-sign-detection-VietNam
Vietnam Traffic Sign Detection Dataset
This repository contains the dataset for detecting road traffic signs in Vietnam using the state-of-the-art YOLO object detection model.
📂 Repository Structure
The dataset is structured in the standard YOLO format, containing images and corresponding annotations divided into training, validation, and testing sets.
├── classid.xlsx # Excel file mapping class IDs to names
├── dataset/
│ ├── train/ #… See the full description on the dataset page: https://huggingface.co/datasets/Minh124689/Traffic-sign-detection-VietNam.VietnameseTableVQA
Overview
This dataset builds from Vietnamese Wikipedia table and questions-answers generated by Gemini-1.5-Flash model for the task Vietnamese Table Visual Question Answering.
The dataset contains table images from diverse domains of life, including sports, games, technology, movie, music, competition, history, and more.
Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedvietnamese_food_imagesvietnam_traffic_qaVietnamese-Entrance-Exam
Vietnamese Entrance Exam Dataset
The Vietnamese Entrance Exam dataset is a collection of 432 problems derived from Vietnamese University entrance examinations. The dataset aims to provide a novel benchmark for testing reasoning capabilities of language models in several low resource domains specifically designed to minimize potential data contamination from pre-training or post-training exposure.
Domain
Count
Physics
95
Chemistry
94
Math
243
Data… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/Vietnamese-Entrance-Exam.Vietnamese_invoicesvietnamese-vlm
Vietnamese Industries Insights
About Me
I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents.
You can connect with me on LinkedIn: Matteo Khan
Dataset Details
Purpose / Mục Đích
Tiếng Việt:
Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese-vlm.VietnameseOCRdatasetvietnamese-ocr-dataset-aggregatedVietnamese-OpenGVLab-ShareGPT-4o-gg-translatedVietnamese_Handwriting_OCRNemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/Nemotron-Personas-Vietnam.vietnamese-rag-benchmark-1kface-celeb-vietnamese
Dataset Card for "face-celeb-vietnamese"
Dataset Summary
This dataset contains information on over 8,000 samples of well-known Vietnamese individuals, categorized into three professions: singers, actors, and beauty queens. The dataset includes data on more than 100 celebrities in each of the three job categories.
Languages
Vietnamese: The label is used to indicate the name of celebrities in Vietnamese.
Dataset Structure
The image and Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/fptudsc/face-celeb-vietnamese.generated-vietnamese-passeports-datasetData generation in machine learning involves creating or manipulating data to train
and evaluate machine learning models. The purpose of data generation is to provide
diverse and representative examples that cover a wide range of scenarios, ensuring the
model's robustness and generalization.
The dataset contains GENERATED Vietnamese passports, which are replicas of official
passports but with randomly generated details, such as name, date of birth etc.
The primary intention of generating these fake passports is to demonstrate the
structure and content of a typical passport document and to train the neural network to
identify this type of document.
Generated passports can assist in conducting research without accessing or compromising
real user data that is often sensitive and subject to privacy regulations. Synthetic
data generation allows researchers to *develop and refine models using simulated
passport data without risking privacy leaks*.vietnam_celeb_facevietnamese
Vietnamese Industries Insights
About Me
I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents.
You can connect with me on LinkedIn: Matteo Khan
Dataset Details
Purpose / Mục Đích
Tiếng Việt:
Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese.Nemotron-Personas-Vietnam
Nemotron-Personas-Vietnam
Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam
A compound AI approach to personas grounded in real-world distributions
Tổng quan (Overview)
Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/bond2bill/Nemotron-Personas-Vietnam.
