CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face02nvidia /Nemotron-Personas-Vietnam Nemotron-Personas-Vietnam Hệ thống AI kết hợp để tạo personas tổng hợp dựa trên phân bố thực tế của Việt Nam A compound AI approach to personas grounded in real-world distributions Tổng quan (Overview) Nemotron-Personas-Vietnam là tập dữ liệu personas được cung cấp dưới dạng mã nguồn mở (CC BY 4.0) dựa trên phân bố nhân khẩu học, địa lý và đặc điểm tính cách của người Việt Nam. Tập dữ liệu phản ánh một cách toàn diện sự phong phú và đặc trưng… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Personas-Vietnam.imagetext-generation100K<n<1M62 likes794 downloads4mo agoHugging Face031TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes340 downloads2y agoHugging Face04vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes325 downloads6mo agoHugging Face05duyet /vietnamese-legal-instruct Vietnamese Legal Instruction Dataset Dataset: huggingface.co/datasets/duyet/vietnamese-legal-instruct | Source code: github.com/duyet/vietnamese-legal-documents-dataset Instruction-following dataset built from th1nhng0/vietnamese-legal-documents — 127K Vietnamese legal documents from vbpl.vn (Government Legal Document Portal, Ministry of Justice). 467,732 training pairs across 14 QA types with deep Vietnamese legal hierarchy knowledge. Every document has a full_text pair for content… See the full description on the dataset page: https://huggingface.co/datasets/duyet/vietnamese-legal-instruct.texttext-generation100K<n<1M4 likes228 downloads6mo agoHugging Face06daitavan /Vietnam-Law-Raw-Datatext-generation1 likes217 downloads2y agoHugging Face07trungbb8 /vietnamese-news-copus-segmented Dataset Card for vietnamese-news-copus-segmented Dataset Summary This dataset is a refined collection of Vietnamese news articles, originally sourced from ademax/binhvq-news-corpus. It has been processed through a specialized pipeline for cleaning, normalization, and word segmentation. It is ideal for training Vietnamese Language Models (LLMs), word embeddings, or text classification tasks. Original Source: ademax/binhvq-news-corpus Language: Vietnamese (vi)… See the full description on the dataset page: https://huggingface.co/datasets/trungbb8/vietnamese-news-copus-segmented.texttext-generation1M<n<10M0 likes207 downloads4mo agoHugging Face08YuITC /Vietnamese-Legal-Documents Vietnamese Legal Documents Dataset 1. Dataset Summary Raw data: tmnam20/BKAI-Legal-Retrieval The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of: A corpus of legal documents. Train/test splits containing natural language queries and their corresponding relevant documents. This dataset is intended to support research and development in: Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.texttext-retrieval100K<n<1M6 likes191 downloads6mo agoHugging Face091TuanPham /KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 9 hours for 2k examples. Usage from datasets import load_dataset kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.textquestion-answering10K<n<100K1 likes151 downloads2y agoHugging Face101TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes128 downloads2y agoHugging Face11nguyenthanhasia /vsec-vietnamese-spell-correction VSEC: Vietnamese Spell Correction Dataset Dataset Description VSEC (Vietnamese Spell Correction) is a comprehensive dataset for Vietnamese spelling error detection and correction, containing 9,341 sentences with 11,202 human-made misspellings across 5,211 unique error types. This dataset represents the largest publicly available collection of Vietnamese spelling errors with syllable-level annotations, making it an invaluable resource for developing and evaluating… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhasia/vsec-vietnamese-spell-correction.texttext-generation1K<n<10K5 likes116 downloads1y agoHugging Face125CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face135CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face14minhnguyent546 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.texttext-classification1M<n<10M2 likes96 downloads6mo agoHugging Face15ai-enthusiasm-community /vietnamese_health_dataset Team and Homepage Official Website: https://aienthusiasm.vn Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community Contact If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com Dataset Structure The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/vietnamese_health_dataset.texttranslation100K<n<1M0 likes94 downloads4mo agoHugging Face16minhxthanh /Vietnam-History-1M-ViLearn more on GitHub: https://github.com/MinhxThanh/Vietnam-History-Chat-Datasets Thông số chính Tổng số mẫu: 1,000,000 Tỷ lệ có reasoning (analysis): ~77.99% Tỷ lệ chỉ trả lời (final-only): ~22.01% Định dạng: messages theo ShareGPT/ChatML Mẫu có reasoning: system → user → assistant (analysis) → assistant (final) Mẫu final-only: system → user → assistant (final) textquestion-answering1M<n<10M29 likes93 downloads1y agoHugging Face17phamson02 /vietnamese-poetry-corpustexttext-generation100K<n<1M9 likes89 downloads3y agoHugging Face18anhquan12 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes87 downloads5mo agoHugging Face19Phuc-HugigFace /Vietnamese-SFT-Corpus-V2 🇻🇳 Vietnamese SFT Corpus V2.1 (Balanced Safety & Anti-Over-Refusal) Vietnamese SFT Corpus V2.1 là tập dữ liệu Tinh chỉnh có Giám sát (Supervised Fine-Tuning - SFT) chuẩn công nghiệp dành cho mô hình ngôn ngữ lớn (LLM) tiếng Việt. Tập dữ liệu được thiết kế nhằm phục vụ huấn luyện trợ lý ảo thông minh, hội thoại tự nhiên, suy luận logic, đồng thời đặc trị triệt để hiện tượng "từ chối lười biếng / từ chối nhầm" (Lazy Refusal / Over-Refusal) vốn xuất hiện phổ biến ở các mô… See the full description on the dataset page: https://huggingface.co/datasets/Phuc-HugigFace/Vietnamese-SFT-Corpus-V2.texttext-generation10K<n<100K0 likes79 downloads15h agoHugging Face20thangvip /vietnamese-legal-qa thangvip/vietnamese-legal-qa Dataset Description This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application). Dataset Structure Data Fields doc_name: Name of the legal document doc_type_name: Type of document (e.g., "Luật" for Law) article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.textquestion-answering1K<n<10K3 likes64 downloads1y agoHugging Face21vlinhd11 /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes64 downloads22d agoHugging Face22vanhthefirst /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes58 downloads6mo agoHugging Face23bkai-foundation-models /vietnamese-roleplay-realm 🇻🇳 Vietnamese Role-play Realm Dataset This is a dataset of GPT-generated Vietnamese characters made to increase the ability of open-source language models to role-play. It contains 446 characters generated by GPT-3.5 Each character will have 20 topics generated by ChatGPT. And each topic will have a conversation corresponding with it In 446 characters, there are 400 general characters and 46 Vietnamese characters. To construct this dataset, we follow a four-step process:… See the full description on the dataset page: https://huggingface.co/datasets/bkai-foundation-models/vietnamese-roleplay-realm.text-generation3 likes56 downloads3y agoHugging Face24NamSyntax /Vietnamese-Legal-QA-RAG Vietnamese Legal QA Dataset for RAG Evaluation (420 rows) Dataset Summary Vietnamese-Legal-QA-RAG is a specialized dataset designed specifically for evaluating Retrieval-Augmented Generation (RAG) systems in the Vietnamese language. Containing 420 meticulously curated rows, this dataset focuses on the Vietnamese legal domain. It is built not just to test simple fact retrieval, but to rigorously evaluate an LLM's ability to perform multi-hop reasoning and, crucially, to… See the full description on the dataset page: https://huggingface.co/datasets/NamSyntax/Vietnamese-Legal-QA-RAG.textquestion-answeringn<1K3 likes53 downloads6mo agoHugging Face25minhxthanh /Vietnam-History-200K-ViBộ dữ liệu 200,000 mẫu (tiếng Việt) về lịch sử Việt Nam 905–2025, đúng định dạng messages cho mô hình tư duy (analysis) và mô hình thường (final-only). Thông tin chính Ngôn ngữ: Tiếng Việt Phạm vi: 905–2025 (các mốc, nhân vật, trận đánh, văn bản, cải cách, thời kỳ…) Định dạng: messages theo ShareGPT/ChatML Mẫu có reasoning (~78%): system → user → assistant (analysis) → assistant (final) Mẫu final-only (~22%): system → user → assistant (final) Tính đa dạng câu hỏi: theo năm, theo sự… See the full description on the dataset page: https://huggingface.co/datasets/minhxthanh/Vietnam-History-200K-Vi.textquestion-answering100K<n<1M0 likes52 downloads1y agoHugging Face26thanhkt /vietnam-normalize-24ktexttext-generation10K<n<100K3 likes51 downloads2y agoHugging Face27trieunh /Vietnamese_literature_VuTrongPhung Vu Trong Phung Literature Chunks This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature. The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens. Chunking Strategy We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens. Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.tabulartext-generationn<1K0 likes51 downloads1y agoHugging Face28thangvip /combined-vietnamese-legal-text thangvip/combined-vietnamese-legal-text Dataset Description This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets: thangvip/vietnamese-legal-qa (9,715 examples) thangvip/law-reading-comprehension-qa-filtered (205,369 examples) Dataset Structure Data Fields text: Combined text containing legal content followed by question-answer pairs in XML-like format Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.textquestion-answering100K<n<1M1 likes51 downloads1y agoHugging Face29pdt590 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 1,335 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes51 downloads6mo agoHugging Face30vlinhd11 /vietnamese-dpo-10k Vietnamese DPO Dataset (10K) This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity. Format: JSONL (one object per line) Fields: "prompt":… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-dpo-10k.textquestion-answering10K<n<100K0 likes48 downloads22d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.