CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01th1nhng0 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn… See the full description on the dataset page: https://huggingface.co/datasets/th1nhng0/vietnamese-legal-documents.texttext-classification1M<n<10M45 likes1.1k downloads2mo agoHugging Face021TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes340 downloads2y agoHugging Face03vohuutridung /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/vohuutridung/vietnamese-legal-documents.texttext-classification1M<n<10M3 likes325 downloads6mo agoHugging Face04duyet /vietnamese-legal-instruct Vietnamese Legal Instruction Dataset Dataset: huggingface.co/datasets/duyet/vietnamese-legal-instruct | Source code: github.com/duyet/vietnamese-legal-documents-dataset Instruction-following dataset built from th1nhng0/vietnamese-legal-documents — 127K Vietnamese legal documents from vbpl.vn (Government Legal Document Portal, Ministry of Justice). 467,732 training pairs across 14 QA types with deep Vietnamese legal hierarchy knowledge. Every document has a full_text pair for content… See the full description on the dataset page: https://huggingface.co/datasets/duyet/vietnamese-legal-instruct.texttext-generation100K<n<1M4 likes228 downloads6mo agoHugging Face055CD-AI /Vietnamese-Multi-turn-Chat-Alpacatextquestion-answering10K<n<100K29 likes218 downloads2y agoHugging Face06YuITC /Vietnamese-Legal-Documents Vietnamese Legal Documents Dataset 1. Dataset Summary Raw data: tmnam20/BKAI-Legal-Retrieval The Vietnamese Legal Documents Dataset is a benchmark dataset designed for legal information retrieval in the Vietnamese language. It consists of: A corpus of legal documents. Train/test splits containing natural language queries and their corresponding relevant documents. This dataset is intended to support research and development in: Information Retrieval (IR)… See the full description on the dataset page: https://huggingface.co/datasets/YuITC/Vietnamese-Legal-Documents.texttext-retrieval100K<n<1M6 likes191 downloads6mo agoHugging Face07hungnm /vietnamese-medical-qa Dataset Summary Vietnamese-Medical-QA is a question-answering dataset in the healthcare domain, collected from edoctor and vinmec. Size: After merging data from these two sources, obtained 9335 QA pairs. Language: Vietnamese Load with Datasets from datasets import load_dataset # Load dataset from huggingface qa_dataset = load_dataset("hungnm/vietnamese-medical-qa") # print a QA example print(qa_dataset['train'][0]) { "question": "Chào bác sĩ,\nRăng cháu hiện tại… See the full description on the dataset page: https://huggingface.co/datasets/hungnm/vietnamese-medical-qa.textquestion-answering1K<n<10K5 likes178 downloads3y agoHugging Face081TuanPham /KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 9 hours for 2k examples. Usage from datasets import load_dataset kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.textquestion-answering10K<n<100K1 likes151 downloads2y agoHugging Face095CD-AI /Vietnamese-Locutusque-function-calling-chatml-gg-translatedtextquestion-answering100K<n<1M27 likes139 downloads2y agoHugging Face101TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes128 downloads2y agoHugging Face115CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes111 downloads3y agoHugging Face125CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face135CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face14minhnguyent546 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 2,393 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/vietnamese-legal-documents.texttext-classification1M<n<10M2 likes96 downloads6mo agoHugging Face15ai-enthusiasm-community /vietnamese_health_dataset Team and Homepage Official Website: https://aienthusiasm.vn Hugging Face Organization: https://huggingface.co/ai-enthusiasm-community Contact If you encounter any issues with the dataset or have any inquiries, please feel free to reach out to us via email at: aienthusiasm.team@gmail.com Dataset Structure The dataset is provided in a flattened tabular format, optimized for the Hugging Face Dataset Viewer and high-speed Parquet processing.… See the full description on the dataset page: https://huggingface.co/datasets/ai-enthusiasm-community/vietnamese_health_dataset.texttranslation100K<n<1M0 likes94 downloads4mo agoHugging Face16anhquan12 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/anhquan12/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes87 downloads5mo agoHugging Face175CD-AI /Vietnamese-395k-meta-math-MetaMathQA-gg-translatedtextquestion-answering100K<n<1M61 likes85 downloads3y agoHugging Face185CD-AI /Vietnamese-ShareGPT4Video-ShareGPT4Video-gg-translatedtextvisual-question-answering10K<n<100K0 likes80 downloads2y agoHugging Face195CD-AI /Vietnamese-ShareGPT4Vision-gg-translatedtextvisual-question-answering100K<n<1M3 likes78 downloads2y agoHugging Face205CD-AI /Vietnamese-TriviaQA-RC-gg-translatedtextquestion-answering3 likes76 downloads3y agoHugging Face215CD-AI /Vietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedtextvisual-question-answering100K<n<1M0 likes73 downloads2y agoHugging Face225CD-AI /Vietnamese-ComplexWebQuestions-gg-translatedquestion-answering10K<n<100K4 likes65 downloads3y agoHugging Face23thangvip /vietnamese-legal-qa thangvip/vietnamese-legal-qa Dataset Description This dataset contains Vietnamese legal documents with automatically generated question-answer pairs. Each document includes comprehension questions of varying difficulty levels (easy, medium, hard) and types (factual, interpretation, analytical, application). Dataset Structure Data Fields doc_name: Name of the legal document doc_type_name: Type of document (e.g., "Luật" for Law) article_content:… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/vietnamese-legal-qa.textquestion-answering1K<n<10K3 likes64 downloads1y agoHugging Face24vlinhd11 /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes64 downloads22d agoHugging Face25MatteoKhan /vietnamese-vlm Vietnamese Industries Insights About Me I'm Matteo Khan, a computer science apprentice at TW3 Partners, specializing in Generative AI and NLP. My focus is on creating datasets that improve AI's ability to process complex technical documents. You can connect with me on LinkedIn: Matteo Khan Dataset Details Purpose / Mục Đích Tiếng Việt: Bộ dữ liệu này được tạo ra nhằm cung cấp cái nhìn tổng quan về các ngành công nghiệp chủ chốt… See the full description on the dataset page: https://huggingface.co/datasets/MatteoKhan/vietnamese-vlm.imagequestion-answering1K<n<10K0 likes59 downloads2y agoHugging Face26vanhthefirst /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive collection of Vietnamese legal documents — laws, decrees, circulars, decisions, and other normative acts — sourced from vbpl.vn, the official Government Legal Document Portal operated by the Ministry of Justice. The dataset includes structured metadata for every document, raw HTML full-text content, and a rich graph of cross-document legal relationships (amendments, citations, repeals, etc.). Curated by: Thịnh Ngô Source: vbpl.vn Language:… See the full description on the dataset page: https://huggingface.co/datasets/vanhthefirst/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes58 downloads6mo agoHugging Face27NamSyntax /Vietnamese-Legal-QA-RAG Vietnamese Legal QA Dataset for RAG Evaluation (420 rows) Dataset Summary Vietnamese-Legal-QA-RAG is a specialized dataset designed specifically for evaluating Retrieval-Augmented Generation (RAG) systems in the Vietnamese language. Containing 420 meticulously curated rows, this dataset focuses on the Vietnamese legal domain. It is built not just to test simple fact retrieval, but to rigorously evaluate an LLM's ability to perform multi-hop reasoning and, crucially, to… See the full description on the dataset page: https://huggingface.co/datasets/NamSyntax/Vietnamese-Legal-QA-RAG.textquestion-answeringn<1K3 likes53 downloads6mo agoHugging Face28thangvip /combined-vietnamese-legal-text thangvip/combined-vietnamese-legal-text Dataset Description This is a combined Vietnamese legal dataset with question-answer pairs formatted in a single text column. It combines two datasets: thangvip/vietnamese-legal-qa (9,715 examples) thangvip/law-reading-comprehension-qa-filtered (205,369 examples) Dataset Structure Data Fields text: Combined text containing legal content followed by question-answer pairs in XML-like format Format… See the full description on the dataset page: https://huggingface.co/datasets/thangvip/combined-vietnamese-legal-text.textquestion-answering100K<n<1M1 likes51 downloads1y agoHugging Face29pdt590 /vietnamese-legal-documents Vietnamese Legal Documents A comprehensive dataset of 518,255 Vietnamese legal documents sourced from thuvienphapluat.vn — the largest Vietnamese legal document repository. The dataset covers laws, decrees, circulars, decisions, and other official documents issued by Vietnamese government bodies, spanning from 1924 to 2026. At a Glance 🗂️ Total documents 518,255 📅 Date range 1924 – 2026 🏛️ Issuing authorities 1,335 unique bodies 📋 Document types 36… See the full description on the dataset page: https://huggingface.co/datasets/pdt590/vietnamese-legal-documents.texttext-classification1M<n<10M0 likes51 downloads6mo agoHugging Face305CD-AI /Vietnamese-OpenGVLab-ShareGPT-4o-gg-translatedtextvisual-question-answering10K<n<100K0 likes50 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.