CoolFace
Datasetpublic

codeparrot/github-jupyter-text-code-pairs

This is a parsed version of github-jupyter-parsed, with markdown and code pairs. We provide the preprocessing script in preprocessing.py. The data is deduplicated and consists of 451662 examples. For similar datasets with text and Python code, there is CoNaLa benchmark from StackOverflow, with some samples curated by annotators.

sourceHugging Faceotherupdated 4y agoView on Hugging Face
7likes225downloads
settings

This repository belongs to codeparrot on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namegithub-jupyter-text-code-pairs
visibilitypublic
licenceother
gatedno
ownercodeparrot
Account settings
codeparrot/github-jupyter-text-code-pairs · CoolFace