CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Viet-Mistral /CulturaY CulturaY: A Large Cleaned Multilingual Dataset of 75 Languages Dataset Summary From the team that brought you CulturaX, we present CulturaY, another substantial multilingual dataset of 15TB (uncompressed)/3TB (zstd-compressed) that applies the same dataset cleaning methodology to the HPLT v1.1 dataset. Please note that HPLT v1.2 has also been released and is an alternative verison with different cleaning methodolgies. This data was used in part to train our SOTA… See the full description on the dataset page: https://huggingface.co/datasets/Viet-Mistral/CulturaY.texttext-generation1B<n<10B39 likes6.5k downloads2y agoHugging Face02thanhnew2001 /VietSuperSpeech VietSuperSpeech Vietnamese Speech Recognition Dataset Dataset Information Total samples: 32,267 Train samples: 29,041 Dev samples: 3,226 Total duration: 103.18 hours Sample rate: 16000 Hz Average segment length: ~12 seconds Source Datasets asr_dataset_nguoivietdailynews asr_dataset_nguyenkhangofficial asr_dataset_trinhlieu Format The dataset follows Icefall format: train.json: Training samples dev.json: Development samples manifest.json:… See the full description on the dataset page: https://huggingface.co/datasets/thanhnew2001/VietSuperSpeech.audio10K<n<100K6 likes3.4k downloads7mo agoHugging Face03ntquang0410 /zh-vie_ecom 1688 zh-vi ecom Dữ liệu sản phẩm song ngữ Trung–Việt từ 1688.com, phục vụ tiểu luận chuyên ngành "Tối ưu hoá mô hình dịch máy Việt–Trung cho TMĐT xuyên biên giới". Xem dữ liệu ở đâu data/snapshot/ là bảng sạch, cập nhật định kỳ — xem ở đây: bilingual_zh_vi.parquet: sản phẩm có cả tiếng Trung và tiếng Việt (1688 tự dịch máy). Mỗi dòng có title_zh/title_vi, description_zh/description_vi (bảng thuộc tính + SKU, đã ghép theo fid nên hai cột song song từng cặp).… See the full description on the dataset page: https://huggingface.co/datasets/ntquang0410/zh-vie_ecom.texttranslation10K<n<100K0 likes1.4k downloads3d agoHugging Face04lidingm /ViewSpatial-Bench ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models Dataset Description We introduce ViewSpatial-Bench, a comprehensive benchmark with over 5,700 question-answer pairs across 1,000+ 3D scenes from ScanNet and MS-COCO validation sets. This benchmark evaluates VLMs' spatial localization capabilities from multiple perspectives, specifically testing both egocentric (camera) and allocentric (human subject) viewpoints across… See the full description on the dataset page: https://huggingface.co/datasets/lidingm/ViewSpatial-Bench.imagevisual-question-answering1K<n<10K23 likes1.2k downloads3mo agoHugging Face05scarlettlin /VietPET-RoI VietPET-RoI VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding boxes. It is intended for medical multimodal research, report generation, visual question answering, and ROI grounding. Research use only. This dataset is not intended for diagnosis, treatment decisions, or direct patient care. Summary Split Patients CT/PET region pairs ROIs Train 160 480 1,544… See the full description on the dataset page: https://huggingface.co/datasets/scarlettlin/VietPET-RoI.tabularimage-to-text1K<n<10K0 likes596 downloads2mo agoHugging Face06Pokerme /view2space-v1 VIEW2SPACE v1 VIEW2SPACE v1 is a multi-view vision-language evaluation dataset for spatial reasoning. Associated paper: VIEW2SPACE: Studying Multi-View Visual Reasoning from Sparse Observations - ECCV 2026 🚀 arXiv: 2603.16506 Project Page: Project Page Related VIEW2SPACE Releases Training release: Pokerme/view2space-train 4B model checkpoint: Pokerme/view2space_4b Collection: Pokerme/view2space The public release is organized into three subsets: count… See the full description on the dataset page: https://huggingface.co/datasets/Pokerme/view2space-v1.imagevisual-question-answering1K<n<10K6 likes565 downloads3mo agoHugging Face07IAmFuch /viet-cultural-vqa 🇻🇳 Vietnamese Cultural VQA Dataset 📖 Dataset Description The Vietnamese Cultural VQA Dataset is a comprehensive multimodal dataset designed for Visual Question Answering (VQA) tasks focused on Vietnamese cultural heritage. This dataset aims to bridge the gap in understanding and preserving Vietnamese culture through AI-powered visual understanding and question answering. 🎯 Dataset Summary 📊 Total Images: 28,505 high-quality cultural images 💬 Total… See the full description on the dataset page: https://huggingface.co/datasets/IAmFuch/viet-cultural-vqa.imagevisual-question-answering10K<n<100K0 likes534 downloads5mo agoHugging Face08b00l26 /VietPET-RoI VietPET-RoI VietPET-RoI is a Vietnamese whole-body PET/CT dataset containing paired cropped 3D volumes, regional reports, and modality-specific 3D ROI bounding boxes. It is intended for medical multimodal research, report generation, visual question answering, and ROI grounding. Research use only. This dataset is not intended for diagnosis, treatment decisions, or direct patient care. Summary Split Patients CT/PET region pairs ROIs Train 160 480 1,544… See the full description on the dataset page: https://huggingface.co/datasets/b00l26/VietPET-RoI.tabularimage-to-text1K<n<10K2 likes396 downloads2mo agoHugging Face09ihbkaiser /dataset_vietnamesetext10M<n<100M0 likes379 downloads9mo agoHugging Face101TuanPham /Vietnamese-OpenO1-SFTOriginal dataset: https://huggingface.co/datasets/qingy2024/OpenO1-SFT-Cleaned This dataset is a Vietnamese translated version of qingy2024/OpenO1-SFT-Cleaned. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using gemini-flash-2.0-exp via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 4 hours for 500 examples. textquestion-answering10K<n<100K6 likes332 downloads2y agoHugging Face115CD-AI /Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated Dataset Card for 5CD-AI/Vietnamese-lmms-lab-LLaVA-Video-178K-gg-translated This translated dataset includes: LLaVA-Video-178K: 178,509 caption entries, 960,791 open-ended QA (question and answer) items, and 196,198 multiple-choice QA items. The video source of the original dataset is in this repo: lmms-lab/LLaVA-Video-178K textvisual-question-answering1M<n<10M1 likes264 downloads2y agoHugging Face125CD-AI /Vietnamese-Multi-turn-Chat-Alpacatextquestion-answering10K<n<100K29 likes221 downloads2y agoHugging Face13leideng /longbench-view Introduction LongBench is the first benchmark for bilingual, multitask, and comprehensive assessment of long context understanding capabilities of large language models. LongBench includes different languages (Chinese and English) to provide a more comprehensive evaluation of the large models' multilingual capabilities on long contexts. In addition, LongBench is composed of six major categories and twenty one different tasks, covering key long-text application scenarios such as… See the full description on the dataset page: https://huggingface.co/datasets/leideng/longbench-view.textquestion-answering1K<n<10K0 likes205 downloads6mo agoHugging Face141TuanPham /KTO-mix-14k-vietnamese-groqOriginal dataset: https://huggingface.co/datasets/trl-lib/kto-mix-14k This dataset is a KTO-formatted version of argilla/dpo-mix-7k. Please cite the original dataset if you find it useful in your work. Translated to Vietnamese with context-aware using Groq Llama3.3 70B* via this repo: https://github.com/vTuanpham/Large_dataset_translator. Roughly 9 hours for 2k examples. Usage from datasets import load_dataset kto_mix_14k_vi =… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/KTO-mix-14k-vietnamese-groq.textquestion-answering10K<n<100K1 likes176 downloads2y agoHugging Face15kolodkin /pcl-viewer-kitti-movie pcl-viewer KITTI movies Draco-compressed LiDAR frames for the pcl-viewer demo, in two folders: geometry/ — sweeps from KITTI raw drive 2011_09_26_drive_0005, positions plus per-point intensity (Draco color green channel). seg/ — SemanticKITTI sequence slice with a per-point class id (Draco color red channel) and intensity (green channel), plus boxes.json (one axis-aligned 3D box per thing instance per frame). Attribution & license Source: KITTI / SemanticKITTI… See the full description on the dataset page: https://huggingface.co/datasets/kolodkin/pcl-viewer-kitti-movie.othern<1K0 likes167 downloads3mo agoHugging Face165CD-AI /Vietnamese-yfcc15m-OpenAICLIPimageimage-to-text10M<n<100M12 likes164 downloads3y agoHugging Face17pre-view /CS50-rawaudio10K<n<100K0 likes162 downloads2y agoHugging Face18aaaaliou /pi-sessions-viewer Coding agent session traces for aaaaliou/pi-sessions-viewer This dataset contains redacted coding agent session traces collected while working on git@github.com:aliou/pi-sessions-viewer.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a… See the full description on the dataset page: https://huggingface.co/datasets/aaaaliou/pi-sessions-viewer.tabulartext-generationn<1K0 likes154 downloads5mo agoHugging Face19ThanhVu101 /Vietnamese-Legal-QA Vietnamese Legal QA — Question Specificity Phân loại độ cụ thể của câu hỏi pháp luật dân sự Việt Nam: broad (hỏi khái quát, phải tổng hợp nhiều chế định) hay narrow (hỏi vào một tình huống / một điều luật xác định). Dùng để định tuyến truy vấn trong hệ RAG pháp luật. Cấu trúc Mỗi dòng là một câu hỏi kèm vết gán nhãn. Hai dòng cùng pair_id là một cặp đối chứng sinh từ cùng một điều luật — một broad, một narrow. Trường Ý nghĩa item_id, pair_id… See the full description on the dataset page: https://huggingface.co/datasets/ThanhVu101/Vietnamese-Legal-QA.texttext-classification1K<n<10K0 likes133 downloads8d agoHugging Face205CD-AI /Vietnamese-alpaca-gpt4-gg-translatedtextquestion-answering10K<n<100K20 likes128 downloads3y agoHugging Face21nhminh107 /VietEmbed-RAG-Science VietEmbed-RAG Science VietEmbed-RAG Science is a Vietnamese retrieval dataset containing 68,567 query-document examples across seven scientific and technical domains. Each record consists of: A Vietnamese query (anchor) A relevant passage (positive) A semantically related but non-answering passage (hard_negative) Topic and domain metadata The dataset is designed for training and domain adaptation of Vietnamese text embedding, semantic retrieval, and Retrieval-Augmented… See the full description on the dataset page: https://huggingface.co/datasets/nhminh107/VietEmbed-RAG-Science.textsentence-similarity10K<n<100K1 likes125 downloads19d agoHugging Face22VietAlphaLabs /vi-en-mathematics-dictionaryVietAlpha English–Vietnamese Mathematics Dictionary Research page · VietAlpha Lab · Source scan The VietAlpha English–Vietnamese Mathematics Dictionary turns a 709-page printed reference work into a machine-readable bilingual lexicon. It contains 26,205 English and Vietnamese mathematics entries digitized from Cung Kim Tiến's Từ Điển Toán Học Anh – Việt, Việt – Anh and organized as JSON Lines. What is in the dataset Direction Entries English to Vietnamese… See the full description on the dataset page: https://huggingface.co/datasets/VietAlphaLabs/vi-en-mathematics-dictionary.texttranslation10K<n<100K1 likes118 downloads5d agoHugging Face23vilm /OpenOrca-Viet 🇻🇳 Vietnamese OpenOrca is here 🐋 Dive into the Vietnamese linguistic landscape with OpenOrca, a cutting-edge dataset crafted through a pioneering partnership between Virtual Interactive and Alignment Lab AI. Drawing inspiration and methodology from the renowned Orca paper, we've expanded our horizons to distill knowledge from a more eclectic mix of leading LLMs including GPT-4, PaLM-2, and Claude. Our vision with this dataset is to fuel research and development that will… See the full description on the dataset page: https://huggingface.co/datasets/vilm/OpenOrca-Viet.text100K<n<1M16 likes114 downloads3y agoHugging Face24OpenVoiceOS /ovos-wake-word-bench-picovoice-view-glass OVOS wake_word bench — picovoice-view-glass Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the arena's… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-view-glass.tabularn<1K0 likes112 downloads17d agoHugging Face25sailor2 /Vietnamese_RAG Dataset Card for Dataset Name Vi's RAG is an comprehensive Vietnamese dataset optimized for RAG Evaluation, build by ZD AI lab and release under Apache license 2.0. Dataset Details There are four datasets in this card : Vietnamese version of Expert QA that we utilize the strong translation ability of GPT-4 for translation task RAG ViQuAD which was carefully chosen from UIT-ViQuAD2.0 with additional context column filtered by title Legal RAG and BKAI_RAG are long form RAG… See the full description on the dataset page: https://huggingface.co/datasets/sailor2/Vietnamese_RAG.text1K<n<10K10 likes108 downloads2y agoHugging Face26nguyenthanhvuh /vietprofs VietProfs VietProfs is a community-maintained directory of Vietnamese and Vietnamese-diaspora academics at universities and eligible public or nonprofit scholarly research institutes worldwide. This repository publishes the current complete, unmodified roster from VietProfs. It is a JSON array whose records describe a person's current academic appointment and may include research areas, education, honors, public profile links, and portrait source URLs. Portrait image files are… See the full description on the dataset page: https://huggingface.co/datasets/nguyenthanhvuh/vietprofs.image1K<n<10K0 likes108 downloads16d agoHugging Face275CD-AI /Vietnamese-Salesforce-xlam-function-calling-60k-gg-translatedtextquestion-answering10K<n<100K8 likes107 downloads2y agoHugging Face28ngwgsang /vietquill-qcpg Introduction This is the primary training data source used to fine-tune VietQuill, the Vietnamese controlled paraphrase generation model presented in our paper. The dataset is designed to support research on Vietnamese paraphrase generation, particularly controlled paraphrase generation with lexical, syntactic, and semantic transformations. Citation If you use VietQuill-QCPG in your research, please cite the corresponding VietQuill-QCPG paper or repository.… See the full description on the dataset page: https://huggingface.co/datasets/ngwgsang/vietquill-qcpg.tabular100K<n<1M0 likes104 downloads1mo agoHugging Face291TuanPham /Vietnamese-magpie-ultra-v0.1Original dataset: https://huggingface.co/datasets/argilla/magpie-ultra-v0.1 ### Dataset Summary `magpie-ultra` it's a synthetically generated dataset for supervised fine-tuning using the new Llama 3.1 405B-Instruct model, together with other Llama models like `Llama-Guard-3-8B` and `Meta-Llama-3.1-8B-Instruct`. The dataset contains challenging instructions and responses for a wide variety of tasks, such as Coding & debugging, Math, Data analysis, Creative Writing, advice seeking, or… See the full description on the dataset page: https://huggingface.co/datasets/1TuanPham/Vietnamese-magpie-ultra-v0.1.textquestion-answering10K<n<100K1 likes100 downloads2y agoHugging Face305CD-AI /Vietnamese-nampdn-ai-tiny-webtext-gg-translatedtextquestion-answering1M<n<10M10 likes99 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.