ajibawa-2023/Python-Code-Large
Python-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.
19599
1version https://git-lfs.github.com/spec/v12oid sha256:fe46a8650db851f782ff122d010639f259316bb7223777fe99229f2f6505c6203size 3463128854 