Mir-2002/python_code_docstring_ast_corpus
Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.
1137
1version https://git-lfs.github.com/spec/v12oid sha256:2ded7aa725ec6468c9ffa16abdb478e358dacd34cc3d0d2ce2273b1e7296eb183size 237523924 