datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the-vault-inlineThe Vault is a multilingual code-text dataset with over 34 million pairs covering 10 popular programming languages.
It is the largest corpus containing parallel code-text data. By building upon The Stack, a massive raw code sample collection,
the Vault offers a comprehensive and clean resource for advancing research in code understanding and generation. It provides a
high-quality dataset that includes code-text pairs at multiple levels, such as class and inline-level, in addition to the function level.
The Vault can serve many purposes at multiple levels.PMC-Inline
PMC-Inline Dataset
PMC-Inline Dataset
Daraset Structure
Sample
This is the text parts and the figure parts can be dowloaded from https://pan.baidu.com/s/1Src_rhXsaOFp8zJ_3zMFsQ?pwd=p3ne.
Dataset Structure
PMC-Inline (PMC papers with inline figures).
We collect the cc lincense papers from pubmed central and remoce the bib, author info, table and iamge captions in the original paper xml files.
Based on the inline figure ref, we link back 11M images into the paper… See the full description on the dataset page: https://huggingface.co/datasets/chaoyi-wu/PMC-Inline.
