ajibawa-2023/Python-Code-Large
Python-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.
19599
1version https://git-lfs.github.com/spec/v12oid sha256:0369215971255e4073eb7f5aa1dee6315720ec397c250214cc6a79d8da3a49e83size 2143588734 