VietAI/vi_pubmed
Dataset Summary 20M Vietnamese PubMed biomedical abstracts translated by the state-of-the-art English-Vietnamese Translation project. The data has been used as unlabeled dataset for pretraining a Vietnamese Biomedical-domain Transformer model. image source: Enriching Biomedical Knowledge for Vietnamese Low-resource Language Through Large-Scale Translation Language English: Original biomedical abstracts from Pubmed Vietnamese: Synthetic abstract translated by a… See the full description on the dataset page: https://huggingface.co/datasets/VietAI/vi_pubmed.
261.3k
