XEUIPR/Java-Code-Large-text-only
Java-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.
02.3k
1version https://git-lfs.github.com/spec/v12oid sha256:9f27a2a81a0f8e934001b5583c1c30867a4d7a273c4582f5d8ffff836e89c99d3size 11075715624 