Mir-2002/python_code_docstring_ast_corpus
Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.
1137
1version https://git-lfs.github.com/spec/v12oid sha256:443843b11f7c9f5fd993409c59ea3f0781f1a117dbb48cf7f1e1022e703cc3693size 51312324 