datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned
This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 4 hours for 500 examples.
Vietnamese-Multi-turn-Chat-AlpacaKTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k
This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 9 hours for 2k examples.
Usage
from datasets import load_dataset
kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1
### Dataset Summary
`magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`.
The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.Vietnamese-alpaca-gpt4-gg-translatedVietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-nampdn-ai-tiny-webtext-gg-translatedVietnamese-395k-meta-math-MetaMathQA-gg-translatedVietnamese-ShareGPT4Vision-gg-translatedVietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedvietnamese-sft-10k
Vietnamese Instruction-Following Dataset (10K)
This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts.
Format: JSONL (one object per line)
Fields: "prompt" (instruction or user message), "response" (assistant reply)
Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.vietnamese-dpo-10k
Vietnamese DPO Dataset (10K)
This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity.
Format: JSONL (one object per line)
Fields:
"prompt":… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-dpo-10k.Vietnamese-cosmos-qa-gg-translatedVietnamese-Openorca-Multiplechoice-gg-translatedVietnamese-microsoft-orca-math-word-problems-200k-gg-translatedVietnamese-LLaVA-Instruct-150K-gg-translatedVietnamese-beyond-rlhf-reward-single-round-gg-translatedVietnamese-meta-math-MetaMathQA-40K-gg-translatedVietnamese-o1-journeyOriginal dataset: https://huggingface.co/datasets/GAIR/o1-journey
This dataset is a Vietnamese translated version of GAIR/o1-journey. Please cite the original dataset if you find it useful in your work.
Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo:
https://github.com/vTuanpham/Large_dataset_translator.
Roughly 2 hours for 649 examples.
vietnamese-sft-10k
Vietnamese Instruction-Following Dataset (10K)
This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts.
Format: JSONL (one object per line)
Fields: "prompt" (instruction or user message), "response" (assistant reply)
Language:… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-sft-10k.Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedvietnamese-medical-dataset
Vietnamese Medical Dataset
A lightweight dataset of (anchor, positive, negative) triplets for training Vietnamese medical text embeddings.
Anchors are short section headers, positives are answer snippets from the same article, and negatives are semantically related snippets from other articles in the same category (semi-hard negatives).
TL;DR
Language: Vietnamese
Domain: Healthcare / Patient education
Format: JSON
Use cases: Contrastive learning (TripletLoss /… See the full description on the dataset page: https://huggingface.co/datasets/mtue29/vietnamese-medical-dataset.Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedVietnamese-NaturalQA-gg-translated-unrefinedVietnamese-mabryCodes-tiny-cot-alpaca-gg-translatedmedical-vietnamese-qa
Medical Vietnamese QA
A dataset of Vietnamese medical question-answer (QA) pairs collected from trusted healthcare websites Vinmec and Long Châu, using the crawl4AI tool.This dataset is intended for research on Question Answering (QA) systems, chatbots, and language model fine-tuning in the medical domain.
Dataset Summary
Language: Vietnamese (vi)
Domain: Healthcare, Medicine, Pharmacy
Task: Question Answering
Source: Public content from Vinmec and Long Châu Pharmacy… See the full description on the dataset page: https://huggingface.co/datasets/Dqdung205/medical-vietnamese-qa.vietnamese-dpo-10k
Vietnamese DPO Dataset (10K)
This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity.
Format: JSONL (one object per line)
Fields:
"prompt":… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-dpo-10k.Vietnamese-395k-meta-math-MetaMathQA-gg-translatedvietnamese-evidence-corpus-chunked
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
47,679 chunks from 13,572 source documents
38,603 Vietnamese chunks and 9,076 English chunks
Maximum chunk length: 512 BGE-M3 tokenizer tokens
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.vietnamese-evidence-corpus-chunked-e5-v3
Vietnamese Evidence Corpus - Chunked
Chunked evidence corpus prepared for multilingual information retrieval,
retrieval-augmented generation, and fact-checking experiments.
Statistics
Chunked with multilingual-E5 token budget
Prefix-aware chunking using `passage: {title}
`
Sentence-aware overlap to preserve local context
Main fields
chunk_id, doc_id, chunk_index
token_start, token_end, token_count
title, text, summary
source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.
