ajibawa-2023/Python-Code-Large
Python-Code-Large Python-Code-Large is a large-scale corpus of Python source code comprising more than 2 million rows of Python code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the Python ecosystem. By providing a high-volume, language-specific corpus, Python-Code-Large enables systematic experimentation in Python-focused model training, domain adaptation, and downstream… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Python-Code-Large.
19599
1version https://git-lfs.github.com/spec/v12oid sha256:9de9a4586e9cde8141f86e5dabd74b5a8ae482c247d0830726ced6972026d1863size 2290234994 