CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
011TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes330 downloads2y agoHugging Face025CD-AI /Vietnamese-Multi-turn-Chat-Alpacatextquestion-answering10K<n<100K29 likes224 downloads2y agoHugging Face031TuanPham /KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 9 hours for 2k examples. Usage from datasets import load_dataset kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.textquestion-answering10K<n<100K1 likes145 downloads2y agoHugging Face041TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes119 downloads2y agoHugging Face055CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes112 downloads3y agoHugging Face065CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes107 downloads2y agoHugging Face075CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes96 downloads3y agoHugging Face085CD-AI /Vietnamese-395k-meta-math-MetaMathQA-gg-translatedtextquestion-answering100K<n<1M61 likes81 downloads3y agoHugging Face095CD-AI /Vietnamese-ShareGPT4Vision-gg-translatedtextvisual-question-answering100K<n<1M3 likes77 downloads2y agoHugging Face105CD-AI /Vietnamese-liuhaotian-llava_v1_5_mix665k-gg-translatedtextvisual-question-answering100K<n<1M0 likes73 downloads2y agoHugging Face11vlinhd11 /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes65 downloads22d agoHugging Face12vlinhd11 /vietnamese-dpo-10k Vietnamese DPO Dataset (10K) This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity. Format: JSONL (one object per line) Fields: "prompt":… See the full description on the dataset page: https://huggingface.co/datasets/vlinhd11/vietnamese-dpo-10k.textquestion-answering10K<n<100K0 likes49 downloads22d agoHugging Face135CD-AI /Vietnamese-cosmos-qa-gg-translatedtextquestion-answering10K<n<100K6 likes46 downloads3y agoHugging Face145CD-AI /Vietnamese-Openorca-Multiplechoice-gg-translatedtabularquestion-answering10K<n<100K2 likes45 downloads2y agoHugging Face155CD-AI /Vietnamese-microsoft-orca-math-word-problems-200k-gg-translatedtexttext-generation100K<n<1M4 likes44 downloads3y agoHugging Face165CD-AI /Vietnamese-LLaVA-Instruct-150K-gg-translatedtextvisual-question-answering100K<n<1M27 likes43 downloads3y agoHugging Face175CD-AI /Vietnamese-beyond-rlhf-reward-single-round-gg-translatedtextquestion-answering10K<n<100K6 likes43 downloads3y agoHugging Face185CD-AI /Vietnamese-meta-math-MetaMathQA-40K-gg-translatedtextquestion-answering10K<n<100K16 likes43 downloads3y agoHugging Face191TuanPham /Vietnamese-o1-journeyOriginal dataset: https://huggingface.co/datasets/GAIR/o1-journey This dataset is a Vietnamese translated version of GAIR/o1-journey. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 2 hours for 649 examples. textquestion-answeringn<1K0 likes40 downloads2y agoHugging Face20522H0134-NguyenNhatHuy /vietnamese-sft-10k Vietnamese Instruction-Following Dataset (10K) This dataset comprises 10,000 Vietnamese instruction-style prompt-response pairs curated for supervised fine-tuning (SFT) of language models. It aims to improve conversational and instruction-following abilities in the Vietnamese language, with coverage across diverse social, cultural, and emotional contexts. Format: JSONL (one object per line) Fields: "prompt" (instruction or user message), "response" (assistant reply) Language:… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-sft-10k.texttext-classification10K<n<100K0 likes38 downloads1y agoHugging Face215CD-AI /Vietnamese-nvidia-OpenMathInstruct-1-50k-gg-translatedtexttext-generation10K<n<100K7 likes36 downloads3y agoHugging Face22mtue29 /vietnamese-medical-dataset Vietnamese Medical Dataset A lightweight dataset of (anchor, positive, negative) triplets for training Vietnamese medical text embeddings. Anchors are short section headers, positives are answer snippets from the same article, and negatives are semantically related snippets from other articles in the same category (semi-hard negatives). TL;DR Language: Vietnamese Domain: Healthcare / Patient education Format: JSON Use cases: Contrastive learning (TripletLoss /… See the full description on the dataset page: https://huggingface.co/datasets/mtue29/vietnamese-medical-dataset.textquestion-answering100K<n<1M4 likes34 downloads1y agoHugging Face23ChaosAIVision /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K0 likes33 downloads9mo agoHugging Face245CD-AI /Vietnamese-NaturalQA-gg-translated-unrefinedtextquestion-answering10K<n<100K19 likes31 downloads3y agoHugging Face255CD-AI /Vietnamese-mabryCodes-tiny-cot-alpaca-gg-translatedtextquestion-answering100K<n<1M23 likes31 downloads3y agoHugging Face26Dqdung205 /medical-vietnamese-qa Medical Vietnamese QA A dataset of Vietnamese medical question-answer (QA) pairs collected from trusted healthcare websites Vinmec and Long Châu, using the crawl4AI tool.This dataset is intended for research on Question Answering (QA) systems, chatbots, and language model fine-tuning in the medical domain. Dataset Summary Language: Vietnamese (vi) Domain: Healthcare, Medicine, Pharmacy Task: Question Answering Source: Public content from Vinmec and Long Châu Pharmacy… See the full description on the dataset page: https://huggingface.co/datasets/Dqdung205/medical-vietnamese-qa.textquestion-answering10K<n<100K0 likes29 downloads1y agoHugging Face27522H0134-NguyenNhatHuy /vietnamese-dpo-10k Vietnamese DPO Dataset (10K) This dataset contains 10,000 Vietnamese prompt-response pairs in the Direct Preference Optimization (DPO) format, including a "prompt", a "chosen" response (preferred), and a "rejected" response (less preferred or misaligned). It is intended for training language models to better align with human-preferred responses, particularly in edge cases involving social sensitivity, rudeness, or toxicity. Format: JSONL (one object per line) Fields: "prompt":… See the full description on the dataset page: https://huggingface.co/datasets/522H0134-NguyenNhatHuy/vietnamese-dpo-10k.textquestion-answering10K<n<100K1 likes26 downloads1y agoHugging Face28ChaosAIVision /Vietnamese-395k-meta-math-MetaMathQA-gg-translatedtextquestion-answering100K<n<1M0 likes23 downloads9mo agoHugging Face29Loctran123 /vietnamese-evidence-corpus-chunked Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.tabulartext-retrieval10K<n<100K0 likes23 downloads2mo agoHugging Face30Loctran123 /vietnamese-evidence-corpus-chunked-e5-v3 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics Chunked with multilingual-E5 token budget Prefix-aware chunking using `passage: {title} ` Sentence-aware overlap to preserve local context Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.tabulartext-retrieval10K<n<100K0 likes22 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.