ajibawa-2023/Python-Code-Large
Python-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.
19599
1version https://git-lfs.github.com/spec/v12oid sha256:685fcbf6a8d446e4a82681575321e1b965030b9daa5b2c27dd9e82ddb69500a53size 3304583624 