CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01minhnguyent546 /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5-397B-A17B over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Mình rất welcome cho các hợp tác liên quan tới building Data Engine và Model Training at Scale. Contact… See the full description on the dataset page: https://huggingface.co/datasets/minhnguyent546/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M0 likes2.3k downloads2mo agoHugging Face025CD-AI /Vietnamese-THUIR-T2Ranking-gg-translated 📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated 📝 Overview Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese. In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.tabulartext-retrieval100M<n<1B22 likes547 downloads1y agoHugging Face035CD-AI /Vietnamese-msMARCO-ggtranslatedtabular10M<n<100M4 likes384 downloads2y agoHugging Face04trannhiem /TranNhiem-Vietnamese-ImageText-Reasoning TranNhiem Vietnamese Image-Text Reasoning (V-LAION) Large-scale Vietnamese multimodal reasoning: multi-turn visual question–answering grounded on natural images, where every answer ships with an explicit chain-of-thought. Reasoning traces and Answer were synthesized by Qwen3.5 over images from the LAION-derived Vi-Laion-gemini-VQA set. Curated by: Trần Nhiệm Languages: Vietnamese (vi) answers · English (en) reasoning Modality: image + text → text Records: 544,795… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-ImageText-Reasoning.imagevisual-question-answering100K<n<1M7 likes308 downloads2mo agoHugging Face05truongpdd /laion-2b-vietnamese-subset Dataset Card for "laion-2b-vietnamese-subset" More Information needed image10M<n<100M7 likes230 downloads4y agoHugging Face06JBrightmanAI /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5-397B-A17B over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Mình rất welcome cho các hợp tác liên quan tới building… See the full description on the dataset page: https://huggingface.co/datasets/JBrightmanAI/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K0 likes206 downloads2mo agoHugging Face07trannhiem /TranNhiem-Vietnamese-DocumentImage-Reasoning TranNhiem Vietnamese Document-Image Reasoning (V-Doc) Vietnamese document-image understanding with explicit reasoning: multi-turn question–answering grounded on scanned/rendered Vietnamese document pages (textbooks, articles, worksheets). Each answer includes a step-by-step chain-of-thought. Reasoning and Answer was synthesized by Qwen3.5 over the Viet-Doc-VQA-II document collection. Curated by: Trần Nhiệm.. Languages: Vietnamese (vi) answers · English (en) reasoning… See the full description on the dataset page: https://huggingface.co/datasets/trannhiem/TranNhiem-Vietnamese-DocumentImage-Reasoning.imagevisual-question-answering10K<n<100K5 likes169 downloads2mo agoHugging Face08tinixai /vietnamese-job-descriptions 💼 Tinix Vietnam Job Description 1. 📌 Giới Thiệu Tinix Vietnam Job Description Tinix Vietnam Job Description là bộ dữ liệu tuyển dụng tiếng Việt ở định dạng CSV, gồm các tin tuyển dụng có cấu trúc về chức danh, công ty, mức lương, địa điểm, loại hợp đồng, ngành nghề, yêu cầu kinh nghiệm, trình độ học vấn, mô tả công việc, phúc lợi, yêu cầu ứng viên và năm đăng tin. Bộ dữ liệu được thiết kế cho các bài toán NLP và phân tích thị trường lao động tại Việt Nam, đặc biệt trong… See the full description on the dataset page: https://huggingface.co/datasets/tinixai/vietnamese-job-descriptions.tabulartext-classification100K<n<1M3 likes151 downloads5mo agoHugging Face09Loctran123 /vietnamese-evidence-retrieval-indexes-v2-r1 Vietnamese Evidence Retrieval Indexes Prebuilt exact dense and sparse indexes for Loctran123/vietnamese-evidence-corpus-embeddings-e5-large-v2-r1 at revision 2a18d35b6ea2e078db95c1aacdc2a28947268b4e. Rows: 63,699 Source embedding shards: 13 Dense: FAISS IndexFlatIP, 1024 dimensions Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75) BM25 content: title repeated 2 times + chunk text Dense input: title + text Dense rows: deduplicated by content hash Vietnamese tokenization: Unicode word… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes-v2-r1.tabularn<1K0 likes58 downloads1mo agoHugging Face10adachankawai /vietnamese-healthcare-dataset Vietnamese Healthcare Synthetic Patient Records This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator. All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments. Included Files Only the following CSV files are included in this upload: File Rows Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.tabulartabular-classification10K<n<100K0 likes55 downloads4mo agoHugging Face11Loctran123 /vietnamese-evidence-retrieval-indexes Vietnamese Evidence Retrieval Indexes Prebuilt exact dense and sparse indexes for Loctran123/vietnamese-evidence-corpus-embeddings-e5-large at revision e928944361ca7d4c80f80d19bec52ebad55a4f7f. Rows: 52,605 Source embedding shards: 11 Dense: FAISS IndexFlatIP, 1024 dimensions Sparse: BM25S Lucene BM25 (k1=1.5, b=0.75) BM25 content: title repeated 2 times + chunk text Vietnamese tokenization: Unicode word tokens, no stemming and no stopword removal row_id in metadata.parquet is… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-retrieval-indexes.tabularn<1K0 likes53 downloads1mo agoHugging Face12trieunh /Vietnamese_literature_VuTrongPhung Vu Trong Phung Literature Chunks This dataset consists of Vietnamese literary texts written by author Vũ Trọng Phụng, one of the most influential figures of 20th-century Vietnamese literature. The dataset includes both short stories and novels, and has been split into smaller chunks based on the number of tokens. Chunking Strategy We used the tokenizer from vinai/PhoGPT-4B to split the original text into chunks of less than 512 tokens. Each row in the dataset… See the full description on the dataset page: https://huggingface.co/datasets/trieunh/Vietnamese_literature_VuTrongPhung.tabulartext-generationn<1K0 likes51 downloads1y agoHugging Face13kaihanzi /hanzi-sino-vietnamese HSK × Sino-Vietnamese (Hán-Việt) character dataset 768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese. Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.tabular1K<n<10K1 likes50 downloads29d agoHugging Face14thanh29nt /vietnamese-toxic-commenttabular100K<n<1M0 likes48 downloads9mo agoHugging Face15vudang449 /aio2025-vietnamese-exam AIO2025 Vietnamese AI Exam Dataset (v3 - Normalized) Bộ dữ liệu câu hỏi trắc nghiệm AI tiếng Việt từ Kỳ thi AI Việt Nam 2025 (AIO2025). Dataset Summary Split Số câu Train 109 Test 34 Tổng 143 Ngôn ngữ: Tiếng Việt Chủ đề: AI, Machine Learning, Deep Learning, Computer Vision, NLP Nguồn: AIO2025 Vietnam AI Exam Version: v3 (normalized, verified by eval_dataset.py) Schema (13 trường) Trường Mô tả id Mã định danh duy nhất… See the full description on the dataset page: https://huggingface.co/datasets/vudang449/aio2025-vietnamese-exam.tabularmultiple-choicen<1K0 likes48 downloads3mo agoHugging Face16iaiuet /banking_sentiment_vietnamesetabular10K<n<100K0 likes45 downloads1y agoHugging Face175CD-AI /Vietnamese-Openorca-Multiplechoice-gg-translatedtabularquestion-answering10K<n<100K2 likes41 downloads2y agoHugging Face18Loctran123 /vietnamese-evidence-corpus-chunked-e5-v3 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics Chunked with multilingual-E5 token budget Prefix-aware chunking using `passage: {title} ` Sentence-aware overlap to preserve local context Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v3.tabulartext-retrieval10K<n<100K0 likes39 downloads1mo agoHugging Face19justtuananh /Eval-RAG-Vietnamesetabular1K<n<10K1 likes38 downloads2y agoHugging Face2052100303-TranPhuocSang /vietnamese-legal-corpus-20k-rawtabular10K<n<100K1 likes35 downloads2y agoHugging Face21ynguyen1010 /medical_vietnamese_datasetstabular100K<n<1M1 likes35 downloads5mo agoHugging Face22dat201204 /vietnamese-caucu-comments Vietnamese Cau Cuu Facebook Comments Dataset Summary This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection. The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu). This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.tabulartext-classification1K<n<10K0 likes32 downloads6mo agoHugging Face23truongpdd /wit_vietnamese_subset Dataset Card for "wit_vietnamese_subset" More Information needed image100K<n<1M3 likes27 downloads4y agoHugging Face24pqthinh232 /vietnamese-restaurant-review-sentiment-datasettabular10K<n<100K0 likes27 downloads6mo agoHugging Face25truongpdd /vietnamese_10classes_train_test_splittabular10K<n<100K0 likes23 downloads4y agoHugging Face26Loctran123 /vietnamese-evidence-corpus-chunked Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked.tabulartext-retrieval10K<n<100K0 likes23 downloads1mo agoHugging Face27Loctran123 /vietnamese-evidence-corpus-embeddings-e5-large-v2tabularn<1K0 likes23 downloads1mo agoHugging Face28zerostratos /vietnamese_toxic_coretabular10K<n<100K2 likes22 downloads1y agoHugging Face29Loctran123 /vietnamese-evidence-corpus-chunked-e5-v2 Vietnamese Evidence Corpus - Chunked Chunked evidence corpus prepared for multilingual information retrieval, retrieval-augmented generation, and fact-checking experiments. Statistics 47,679 chunks from 13,572 source documents 38,603 Vietnamese chunks and 9,076 English chunks Maximum chunk length: 512 BGE-M3 tokenizer tokens Main fields chunk_id, doc_id, chunk_index token_start, token_end, token_count title, text, summary source, source_type… See the full description on the dataset page: https://huggingface.co/datasets/Loctran123/vietnamese-evidence-corpus-chunked-e5-v2.tabulartext-retrieval10K<n<100K0 likes20 downloads1mo agoHugging Face30bihungba1101 /slam-en-es-vietnamese-prompts SLAM English-Spanish Knowledge Tracing with Vietnamese Prompts This dataset is a derived version of bihungba1101/slam-en-es-knowledge-tracing. It preserves every source row and field and adds prompt_vi, a machine-generated Vietnamese translation of the Spanish prompt. The source track contains English learners who already speak Spanish. Consequently: prompt is the original Spanish text shown to the learner when text is available. prompt_vi is a Vietnamese translation of prompt.… See the full description on the dataset page: https://huggingface.co/datasets/bihungba1101/slam-en-es-vietnamese-prompts.tabulartext-classification1M<n<10M0 likes15 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.