Mir-2002/python_code_docstring_ast_corpus
Overview This dataset contains 34,000+ rows of code-docstring-ast data along with additional metadata. Data was gathered from various Python libraries and frameworks and their publicly available GitHub repos. This dataset was created for the purpose of training the CodeT5+ transformer on AST-enhanced code-to-doc tasks. Sources The dataset was gathered from various GitHub repos sampled from this repo by Vinta. The 26 repos are: matplotlib pytorch cryptography… See the full description on the dataset page: https://huggingface.co/datasets/Mir-2002/python_code_docstring_ast_corpus.
Update README.md
Upload 4 files
Delete val.json
Delete train.json
Delete test.json
Delete preprocessing_stats.json
Update README.md
Update README.md
Upload 4 files
Delete val.json
Delete train.json
Delete test.json
Delete preprocessing_stats.json
Update README.md
Upload 4 files
Delete validation.json
Delete train.json
Delete test.json
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Uploaded improved versions
Delete validation.json
Delete train.json
Delete test.json
Delete validation.jsonl
Delete train.jsonl
Delete test.jsonl
Uploaded JSONL versions
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Update README.md
Create README.md file
Upload 3 files
initial commit
