datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Vietnamese-THUIR-T2Ranking-gg-translated
📚 5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated
📝 Overview
Vietnamese-THUIR-T2Ranking-gg-translated is a large-scale dataset for passage ranking in Vietnamese.It is translated from the original THUIR/T2Ranking [1] using Google Translate, inspired by the approach of mMARCO [2].The dataset aims to provide a large-scale dataset for research and applications in Information Retrieval (IR) in Vietnamese.
In IR, passage ranking is an essential and challenging task… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Vietnamese-THUIR-T2Ranking-gg-translated.vietnamese-toxic-commentvietnamese-healthcare-dataset
Vietnamese Healthcare Synthetic Patient Records
This dataset contains synthetic Vietnamese healthcare identity records from multiple source systems, plus a canonical synthetic patient table used by the generator.
All records are synthetic and are intended for entity resolution, record linkage, and Vietnamese identity-field preprocessing experiments.
Included Files
Only the following CSV files are included in this upload:
File
Rows
Description… See the full description on the dataset page: https://huggingface.co/datasets/adachankawai/vietnamese-healthcare-dataset.hanzi-sino-vietnamese
HSK × Sino-Vietnamese (Hán-Việt) character dataset
768 HSK characters joined with their Sino-Vietnamese (Hán-Việt) readings, radical breakdowns and hand-written memory hooks in Vietnamese.
Open HSK wordlists are plentiful. The Sino-Vietnamese layer is what is missing from all of them — and it is the layer that matters most for the ~1 million Vietnamese speakers studying Chinese, because roughly 60% of Vietnamese vocabulary descends from Chinese. A learner meeting 学 (xué) already… See the full description on the dataset page: https://huggingface.co/datasets/kaihanzi/hanzi-sino-vietnamese.vietnamese-legal-corpus-20k-rawvietnamese-caucu-comments
Vietnamese Cau Cuu Facebook Comments
Dataset Summary
This dataset contains Vietnamese Facebook comments collected from a natural-disaster discussion thread and auto-labeled for binary emergency detection.
The target task is to detect whether a comment is a real-time rescue request (cau_cuu) versus a non-emergency comment (khong_phai_cau_cuu).
This release is intended as a bootstrap dataset for triage modeling and should be treated as a weakly supervised resource. Human… See the full description on the dataset page: https://huggingface.co/datasets/dat201204/vietnamese-caucu-comments.vietnamese_toxic_corevietnamese-classification-label-number
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/hiuman/vietnamese-classification-label-number.
