CoolFace
Datasetpublic

Mir-2002/python_code_docstring_ast_corpus

Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.

sourceHugging Faceupdated 1y agoView on Hugging Face
1likes137downloads
train.json4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:2ded7aa725ec6468c9ffa16abdb478e358dacd34cc3d0d2ce2273b1e7296eb183size 237523924