datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CulturaY
CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages
Dataset Summary
From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset.
Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies.
This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.VietSuperSpeech
VietSuperSpeech
Vietnamese Speech Recognition Dataset
Dataset Information
Total samples: 32,267
Train samples: 29,041
Dev samples: 3,226
Total duration: 103.18 hours
Sample rate: 16000 Hz
Average segment length: ~12 seconds
Source Datasets
asr_dataset_nguoivietdailynews
asr_dataset_nguyenkhangofficial
asr_dataset_trinhlieu
Format
The dataset follows Icefall format:
train.json: Training samples
dev.json: Development samples
manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.zh-vie_ecom
1688 zh-vi ecom
Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên
ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới".
Xem dữ liệu ở đâu
data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây:
bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự
dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi
(bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.ViewSpatial-Bench
ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
Dataset Description
We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.VietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/scarlettlin/VietPET-RoI.view2space-v1
VIEW2SPACE v1
VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning.
Associated paper:
VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀
arXiv: 2603.16506
Project Page: Project Page
Related VIEW2SPACE Releases
Training release: Pokerme/view2space-train
4B model checkpoint: Pokerme/view2space_4b
Collection: Pokerme/view2space
The public release is organized into three subsets:
count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.viet-cultural-vqa
🇻🇳 Vietnamese Cultural VQA Dataset
📖 Dataset Description
The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering.
🎯 Dataset Summary
📊 Total Images: 28,505 high-quality cultural images
💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.VietPET-RoI
VietPET-RoI
VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired
cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding
boxes. It is intended for medical multimodal research, report generation,
visual question answering, and ROI grounding.
Research use only. This dataset is not intended for diagnosis, treatment
decisions, or direct patient care.
Summary
Split
Patients
CT/PET region pairs
ROIs
Train
160
480
1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.dataset_vietnameseVietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned
This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 4 hours for 500 examples.
Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated
This translated dataset includes:
LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items.
The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K
Vietnamese-Multi-turn-Chat-Alpacalongbench-view
Introduction
LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.pcl-viewer-kitti-movie
pcl-viewer KITTI movies
Draco-compressed LiDAR frames for the
pcl-viewer demo, in two folders:
geometry/ — sweeps from KITTI raw drive 2011_09_26_drive_0005, positions
plus per-point intensity (Draco color green channel).
seg/ — SemanticKITTI sequence slice with a per-point class id (Draco
color red channel) and intensity (green channel), plus boxes.json (one
axis-aligned 3D box per thing instance per frame).
Attribution & license
Source: KITTI / SemanticKITTI… See the full description on the dataset page: https://huggingface.co/datasets/kolodkin/pcl-viewer-kitti-movie.Vietnamese-yfcc15m-OpenAICLIPCS50-rawpi-sessions-viewer
Coding agent session traces for aaaaliou/pi-sessions-viewer
This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.Vietnamese-Legal-QA
Vietnamese Legal QA — Question Specificity
Phân loại độ cụ thể của câu hỏi pháp luật dân sự Việt Nam: broad (hỏi khái
quát, phải tổng hợp nhiều chế định) hay narrow (hỏi vào một tình huống / một
điều luật xác định). Dùng để định tuyến truy vấn trong hệ RAG pháp luật.
Cấu trúc
Mỗi dòng là một câu hỏi kèm vết gán nhãn. Hai dòng cùng pair_id là một cặp
đối chứng sinh từ cùng một điều luật — một broad, một narrow.
Trường
Ý nghĩa
item_id, pair_id… See the full description on the dataset page: https://huggingface.co/datasets/ThanhVu101/Vietnamese-Legal-QA.Vietnamese-alpaca-gpt4-gg-translatedVietEmbed-RAG-Science
VietEmbed-RAG Science
VietEmbed-RAG Science is a Vietnamese retrieval dataset containing 68,567 query-document examples across seven scientific and technical domains.
Each record consists of:
A Vietnamese query (anchor)
A relevant passage (positive)
A semantically related but non-answering passage (hard_negative)
Topic and domain metadata
The dataset is designed for training and domain adaptation of Vietnamese text embedding, semantic retrieval, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/nhminh107/VietEmbed-RAG-Science.vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary
Research page ·
VietAlpha Lab ·
Source scan
The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines.
What is in the dataset
Direction
Entries
English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.OpenOrca-Viet
🇻🇳 Vietnamese OpenOrca is here 🐋
Dive into the Vietnamese linguistic landscape with OpenOrca, a cutting-edge dataset crafted through a pioneering partnership between Virtual Interactive and Alignment Lab AI. Drawing inspiration and methodology from the renowned Orca paper, we've expanded our horizons to distill knowledge from a more eclectic mix of leading LLMs including GPT-4, PaLM-2, and Claude. Our vision with this dataset is to fuel research and development that will… See the full description on the dataset page: https://huggingface.co/datasets/vilm/OpenOrca-Viet.ovos-wake-word-bench-picovoice-view-glass
OVOS wake_word bench — picovoice-view-glass
Per-clip detection decisions predictions of the registered
OVOS Plugin Arena
wake_word fighters over
Picovoice/wake-word-benchmark.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.Vietnamese_RAG
Dataset Card for Dataset Name
Vi's RAG is an comprehensive Vietnamese dataset optimized for RAG Evaluation, build by ZD AI lab and release under Apache license 2.0.
Dataset Details
There are four datasets in this card :
Vietnamese version of Expert QA that we utilize the strong translation ability of GPT-4 for translation task
RAG ViQuAD which was carefully chosen from UIT-ViQuAD2.0 with additional context column filtered by title
Legal RAG and BKAI_RAG are long form RAG… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/Vietnamese_RAG.vietprofs
VietProfs
VietProfs is a community-maintained directory of Vietnamese and Vietnamese-diaspora academics at universities and eligible public or nonprofit scholarly research institutes worldwide.
This repository publishes the current complete, unmodified roster from VietProfs. It is a JSON array whose records describe a person's current academic appointment and may include research areas, education, honors, public profile links, and portrait source URLs. Portrait image files are… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhvuh/vietprofs.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedvietquill-qcpg
Introduction
This is the primary training data source used to fine-tune VietQuill, the Vietnamese controlled paraphrase generation model presented in our paper.
The dataset is designed to support research on Vietnamese paraphrase generation, particularly controlled paraphrase generation with lexical, syntactic, and semantic transformations.
Citation
If you use VietQuill-QCPG in your research, please cite the corresponding VietQuill-QCPG paper or repository.… See the full description on the dataset page: https://huggingface.co/datasets/ngwgsang/vietquill-qcpg.Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.Vietnamese-nampdn-ai-tiny-webtext-gg-translated
