CoolFace
Datasetpublic

ajibawa-2023/Java-Code-Large

Java-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large.

sourceHugging Facemitupdated 7mo agoView on Hugging Face
34likes1.1kdownloads
java_only_0001.jsonl4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:30be6f3b45bab7b6cde9b12fc8e2efd8efcbe705b0e2d1b157ba8c419fcc9d503size 13387965404