ajibawa-2023/Python-Code-Large
Python-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.
19599
1version https://git-lfs.github.com/spec/v12oid sha256:820ea8d957929d232d7886a5327cee9a344d0fe0a1ab584c5080c4d9730b8daf3size 1938450624 