CoolFace
Datasetpublic

jknafou/TransCorpus-bio

TransCorpus-bio TransCorpus-bio is a large-scale, parallel biomedical corpus consisting of PubMed abstracts (title + abstract), translated with the TransCorpus Toolkit using NLLB-200. It is designed to enable high-quality multi-lingual biomedical language modeling and downstream NLP research. This dataset was restructured from five separate single-language repositories into one dataset with a config (tab in the dataset viewer) per language, and with each row carrying its source… See the full description on the dataset page: https://huggingface.co/datasets/jknafou/TransCorpus-bio.

sourceHugging Facemitupdated 2d agoView on Hugging Face
0likes300downloads
../
filetrain-00000-of-00008.parquet1.74 GBdownload
filetrain-00001-of-00008.parquet2.09 GBdownload
filetrain-00002-of-00008.parquet2.18 GBdownload
filetrain-00003-of-00008.parquet2.22 GBdownload
filetrain-00004-of-00008.parquet2.34 GBdownload
filetrain-00005-of-00008.parquet2.42 GBdownload
filetrain-00006-of-00008.parquet2.51 GBdownload
filetrain-00007-of-00008.parquet501.0 MBdownload

jknafou/TransCorpus-bio · main · files are served by the source, never re-hosted here