CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
011TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes340 downloads2y agoHugging Face021TuanPham /KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 9 hours for 2k examples. Usage from datasets import load_dataset kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.textquestion-answering10K<n<100K1 likes151 downloads2y agoHugging Face031TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes128 downloads2y agoHugging Face045CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes108 downloads2y agoHugging Face055CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes98 downloads3y agoHugging Face06Phuc-HugigFace /Vietnamese-SFT-Corpus-V2 🇻🇳 Vietnamese SFT Corpus V2.1 (Balanced Safety & Anti-Over-Refusal) Vietnamese SFT Corpus V2.1 là tập dữ liệu Tinh chỉnh có Giám sát (Supervised Fine-Tuning - SFT) chuẩn công nghiệp dành cho mô hình ngôn ngữ lớn (LLM) tiếng Việt. Tập dữ liệu được thiết kế nhằm phục vụ huấn luyện trợ lý ảo thông minh, hội thoại tự nhiên, suy luận logic, đồng thời đặc trị triệt để hiện tượng "từ chối lười biếng / từ chối nhầm" (Lazy Refusal / Over-Refusal) vốn xuất hiện phổ biến ở các mô… See the full description on the dataset page: https://huggingface.co/datasets/Phuc-HugigFace/Vietnamese-SFT-Corpus-V2.texttext-generation10K<n<100K0 likes79 downloads11h agoHugging Face07vlinhd11 /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes64 downloads21d agoHugging Face08vlinhd11 /vietnamese-dpo-10k Vietnamese DPO Dataset (10K) This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity. Format: JSONL (one object per line) Fields: "prompt":… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-dpo-10k.textquestion-answering10K<n<100K0 likes48 downloads21d agoHugging Face091TuanPham /Vietnamese-o1-journeyOriginal dataset: https://huggingface.co/datasets/GAIR/o1-journey This dataset is a Vietnamese translated version of GAIR/o1-journey. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 2 hours for 649 examples. textquestion-answeringn<1K0 likes46 downloads2y agoHugging Face105CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face11522H0134-NguyenNhatHuy /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes43 downloads1y agoHugging Face125CD-AI /Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedtexttext-generation10K<n<100K7 likes39 downloads3y agoHugging Face13ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes34 downloads9mo agoHugging Face14522H0134-NguyenNhatHuy /vietnamese-dpo-10k Vietnamese DPO Dataset (10K) This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity. Format: JSONL (one object per line) Fields: "prompt":… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-dpo-10k.textquestion-answering10K<n<100K1 likes29 downloads1y agoHugging Face15anhnon /vietnamese-corporate-legal-articles-fsm Lexora Knowledge - Vietnamese Legal Documents Dataset Summary A structured Vietnamese legal knowledge base crawled from vbpl.vn (CSDL Quốc gia về Pháp luật - Vietnam's National Legal Database), published as 4 linked subsets: full documents, individual articles (Điều), the citation graph between documents/articles, and domain-concept tags. Load a specific subset with datasets.load_dataset("anhnon/vietnamese-corporate-legal-articles-fsm", "articles") etc. Intended… See the full description on the dataset page: https://huggingface.co/datasets/anhnon/vietnamese-corporate-legal-articles-fsm.texttext-generation100K<n<1M1 likes28 downloads2mo agoHugging Face165CD-AI /Vietnamese-argilla-OpenHermesPreferences-66k-gg-translatedtexttext-generation10K<n<100K6 likes18 downloads3y agoHugging Face17Dang-DN-VN /vietnamese-legal-qa-mini-300 Vietnamese Legal Q&A — SFT Dataset A domain-specific supervised fine-tuning dataset for Vietnamese legal question answering, built for LLM fine-tuning and instruction tuning. Dataset Summary This dataset contains 300 Vietnamese legal Q&A samples covering common areas of Vietnamese civil, criminal, labor, and administrative law. All samples are in Vietnamese and follow the Alpaca format. Split Samples Train 250 Validation 50 Total 300… See the full description on the dataset page: https://huggingface.co/datasets/Dang-DN-VN/vietnamese-legal-qa-mini-300.textquestion-answeringn<1K0 likes17 downloads4mo agoHugging Face18nguyenphuthien /vietnamese_no_robotsgated Vietnamese-translated version of HuggingFaceH4/no_robots dataset Dataset Card for No Robots 🙅‍♂️🤖 Look Ma, an instruction dataset that wasn't generated by GPTs! Dataset Summary No Robots is a high-quality dataset of 10,000 instructions and demonstrations created by skilled human annotators. This data can be used for supervised fine-tuning (SFT) to make language models follow instructions better. No Robots was modelled after the instruction dataset described… See the full description on the dataset page: https://huggingface.co/datasets/nguyenphuthien/vietnamese_no_robots.texttext-generation10K<n<100K2 likes14 downloads3y agoHugging Face19nguyenphuthien /vietnamese_ultrachat_200kgated Dataset Card for Vietnamese UltraChat 200k Dataset Description This is a heavily filtered version of the UltraChat dataset and was used to train Zephyr-7B-β, a state of the art 7b chat model. The original datasets consists of 1.4M dialogues generated by ChatGPT and spanning a wide range of topics. To create UltraChat 200k, we applied the following logic: Selection of a subset of data for faster supervised fine tuning. Truecasing of the dataset, as we observed around 5%… See the full description on the dataset page: https://huggingface.co/datasets/nguyenphuthien/vietnamese_ultrachat_200k.texttext-generation100K<n<1M12 likes11 downloads2y agoHugging Face20PhatNguyen41 /vietnamese-instruction-dataset Vietnamese Instruction Dataset This dataset contains instruction-response pairs for training instruction-following models. Format JSON with two fields: instruction response Use Cases Instruction tuning Chatbot training texttext-generationn<1K0 likes10 downloads9mo agoHugging Face21nguyenphuthien /vietnamese_ultrafeedback_binarizedgatedtabulartext-generation10K<n<100K2 likes2 downloads2y agoHugging Face221TuanPham /KTO-mix-14k-vietnamesegatedCompatible with KTO Trainer of trl library. Data was filtered to excluded coding examples, so there is no worry of translation errors. Leave a heart and gud luck, Vietnamese tuners 🤗. texttext-generation10K<n<100K0 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.