datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.vietnamese-music-dataset
Vietnamese Music Dataset
A collection of 4,820 Vietnamese music tracks with matching cover thumbnails and per-track metadata collected from YouTube, packaged as an audiofolder dataset.
Repository structure
Path
Contents
Count
audio/
MP3 audio files, named by YouTube video ID
4,820
images/
PNG cover thumbnails, same IDs as audio/
4,820
data/
Parquet metadata files, one per collection session
31
Metadata schema
Each Parquet file in… See the full description on the dataset page: https://huggingface.co/datasets/Toan-Minh-Duong-Son/vietnamese-music-dataset.TranNhiem-Vietnamese-ImageText-Reasoning
TranNhiem Vietnamese Image-Text Reasoning (V-LAION)
Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on
natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces
and Answer were synthesized by Qwen3.5 over images from the LAION-derived Vi-Laion-gemini-VQA set.
Curated by: Trần Nhiệm
Languages: Vietnamese (vi) answers · English (en) reasoning
Modality: image + text → text
Records: 544,795… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-ImageText-Reasoning.laion-2b-vietnamese-subset
Dataset Card for "laion-2b-vietnamese-subset"
More Information needed
TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.TranNhiem-Vietnamese-DocumentImage-Reasoning
TranNhiem Vietnamese Document-Image Reasoning (V-Doc)
Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering
grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each
answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5
over the Viet-Doc-VQA-II document collection.
Curated by: Trần Nhiệm..
Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.Vietnamese-yfcc15m-OpenAICLIPvietnamese-music-datasetVietnameseTableVQA
Overview
This dataset builds from Vietnamese Wikipedia table and questions-answers generated by Gemini-1.5-Flash model for the task Vietnamese Table Visual Question Answering.
The dataset contains table images from diverse domains of life, including sports, games, technology, movie, music, competition, history, and more.
Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedvietnamese_food_imagesvietnamese-vlm
Vietnamese Industries Insights
About Me
I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents.
You can connect with me on LinkedIn: Matteo Khan
Dataset Details
Purpose / Mục Đích
Tiếng Việt:
Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese-vlm.Vietnamese_invoicesVietnameseOCRdatasetVietnamese_Handwriting_OCRvietnamese-ocr-dataset-aggregatedvietnamese-rag-benchmark-1kVietnamese-OpenGVLab-ShareGPT-4o-gg-translatedface-celeb-vietnamese
Dataset Card for "face-celeb-vietnamese"
Dataset Summary
This dataset contains information on over 8,000 samples of well-known Vietnamese individuals, categorized into three professions: singers, actors, and beauty queens. The dataset includes data on more than 100 celebrities in each of the three job categories.
Languages
Vietnamese: The label is used to indicate the name of celebrities in Vietnamese.
Dataset Structure
The image and Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/fptudsc/face-celeb-vietnamese.Vietnamese-Entrance-Exam
Vietnamese Entrance Exam Dataset
The Vietnamese Entrance Exam dataset is a collection of 432 problems derived from Vietnamese University entrance examinations. The dataset aims to provide a novel benchmark for testing reasoning capabilities of language models in several low resource domains specifically designed to minimize potential data contamination from pre-training or post-training exposure.
Domain
Count
Physics
95
Chemistry
94
Math
243
Data… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/Vietnamese-Entrance-Exam.vietnamese
Vietnamese Industries Insights
About Me
I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents.
You can connect with me on LinkedIn: Matteo Khan
Dataset Details
Purpose / Mục Đích
Tiếng Việt:
Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese.vietnamese-vqa_merged_2vietnamese_handwriting_line_ocrgenerated-vietnamese-passeports-datasetData generation in machine learning involves creating or manipulating data to train
and evaluate machine learning models. The purpose of data generation is to provide
diverse and representative examples that cover a wide range of scenarios, ensuring the
model's robustness and generalization.
The dataset contains GENERATED Vietnamese passports, which are replicas of official
passports but with randomly generated details, such as name, date of birth etc.
The primary intention of generating these fake passports is to demonstrate the
structure and content of a typical passport document and to train the neural network to
identify this type of document.
Generated passports can assist in conducting research without accessing or compromising
real user data that is often sensitive and subject to privacy regulations. Synthetic
data generation allows researchers to *develop and refine models using simulated
passport data without risking privacy leaks*.vietnamese_character_diacritic_cwlvietnamese_physics_highschool_multimodalvietnamese-traffic-sign-vqa-1k
Vietnamese Traffic Sign VQA — Balanced 1K Test Set
A small, deduplicated and balanced Visual Question Answering (VQA) test set for
Vietnamese traffic signs. Every example is an image + a Vietnamese question + an answer.
The set is intended purely for evaluation (all examples are test).
Quick start
from datasets import load_dataset
ds = load_dataset("tungitachi/vietnamese-traffic-sign-vqa-1k", split="test")
ex = ds[0]
ex["image"] # PIL.Image… See the full description on the dataset page: https://huggingface.co/datasets/tungitachi/vietnamese-traffic-sign-vqa-1k.vietnamese-traffic-sign-vqa-dedup-1k
Vietnamese Traffic-Sign VQA — deduplicated, original-split (1k)
A clean, deduplicated and leakage-free Visual-Question-Answering set for
Vietnamese traffic signs, split train/test by the original detection split.
This is a curated derivative — it is not the same as
Anakonkai/vietnamese-traffic-sign-vqa:
that source is dominated by near-identical video frames and has no clean
train/test boundary. Here those issues are fixed (see How it was built).
What's inside… See the full description on the dataset page: https://huggingface.co/datasets/tungitachi/vietnamese-traffic-sign-vqa-dedup-1k.wit_vietnamese_subset
Dataset Card for "wit_vietnamese_subset"
More Information needed
vqa-vietnamese-charts
