CoolFace
Datasetpublic

codeparrot/github-jupyter-text-code-pairs

This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
7likes225downloads

codeparrot/github-jupyter-text-code-pairs · main · files are served by the source, never re-hosted here