CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SWE-bench /SWE-smith-javatext1K<n<10K0 likes16k downloads7mo agoHugging Face02AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face03bagasshw /n-hypo-java100K<n<1M0 likes3.6k downloads1y agoHugging Face04hrishizone /Java-GitHub-Codestext1M<n<10M1 likes2.9k downloads1y agoHugging Face05ajibawa-2023 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/JavaScript-Code-Large.texttext-generation1M<n<10M33 likes2.4k downloads7mo agoHugging Face06angie-chen55 /javascript-github-codetext10M<n<100M18 likes2.3k downloads4y agoHugging Face07XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes2.3k downloads15d agoHugging Face08ajibawa-2023 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large.texttext-generation10M<n<100M34 likes1.1k downloads7mo agoHugging Face09susnato /java_PRstabular100K<n<1M0 likes913 downloads3y agoHugging Face10JWei05 /SWE-smith-java-6704-filtered-for-problem-statementstext1K<n<10K0 likes875 downloads7mo agoHugging Face11AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes807 downloads1y agoHugging Face12sanjaykz /Java-codetexttext-generation100K<n<1M2 likes794 downloads11mo agoHugging Face13Kitxuuu /AIXCC-Java-Challenge 🧠 AIXCC Challenge Benchmark – Java Challenges 📌 Overview AIXCC Challenge Benchmark (Java Challenges) is a curated subset of the full AIXCC Challenge Benchmark focused solely on C-based challenges. This benchmark is built upon the official AIXCC Challenge. Each Java challenge consists of either a delta (focused diff) or full (whole project) test case. 📎 Reference Implementation This benchmark is designed to work with our open-source CRS system:… See the full description on the dataset page: https://huggingface.co/datasets/Kitxuuu/AIXCC-Java-Challenge.0 likes741 downloads1y agoHugging Face14LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes725 downloads2y agoHugging Face15thedruid831 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/thedruid831/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes648 downloads7mo agoHugging Face16marcelo1234 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marcelo1234/Java-Code-Large.texttext-generation10M<n<100M2 likes616 downloads4mo agoHugging Face17minhhien0811 /java_evaluation_benchmarkstext1K<n<10K0 likes610 downloads4mo agoHugging Face18JWei05 /SWE-smith-java-6450-filteredtext1K<n<10K0 likes583 downloads7mo agoHugging Face19hchautran /javascripttext100K<n<1M0 likes560 downloads4y agoHugging Face20Shuu12121 /github-file-programs-dataset-javatext1M<n<10M0 likes554 downloads9mo agoHugging Face21javadtaghia /deewaiREALCN-training Repo git@hf.co:datasets/telcom/deewaiREALCN-training DeewaiREALCN Training Data Image–text pairs for training captioning or vision–language models. Each image is a 1024×1024 RGB JPEG portrait with a short English description. Contents data/train/: 9,000 pairs for training. images/: JPEG files (090000.jpg, …). captions.jsonl: one JSON object per line with file_name and text. data/val/: 1,000 pairs for validation with the same layout. Example… See the full description on the dataset page: https://huggingface.co/datasets/javadtaghia/deewaiREALCN-training.image10K<n<100K2 likes550 downloads9mo agoHugging Face22nomic-ai /cornstack-java-v1 CoRNStack Python Dataset The CoRNStack Dataset, accepted to ICLR 2025, is a large-scale high quality training dataset specifically for code retrieval across multiple programming languages. This dataset comprises of <query, positive, negative> triplets used to train nomic-embed-code, CodeRankEmbed, and CodeRankLLM. CoRNStack Dataset Curation Starting with the deduplicated Stackv2, we create text-code pairs from function docstrings and respective code. We filtered out… See the full description on the dataset page: https://huggingface.co/datasets/nomic-ai/cornstack-java-v1.text10M<n<100M3 likes491 downloads1y agoHugging Face23ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes470 downloads10mo agoHugging Face24jitx /Methods2Test_java_unit_test_code Dataset Description Microsoft created this large dataset of Java Junit test cases with its corresponding focal methods. It contains 780k pairs of JUnit test cases and focal methods which were extracted from a total of 91K Java open source project hosted on GitHub. The mapping between test case and focal methods are based heuristics rules and Java developer's best practice. More information could be found here: methods2test Github repo Methods2Test: A dataset of focal methods… See the full description on the dataset page: https://huggingface.co/datasets/jitx/Methods2Test_java_unit_test_code.texttext-generation100K<n<1M18 likes456 downloads3y agoHugging Face25JWei05 /swe_smith_java_6457_filteredtext1K<n<10K0 likes452 downloads6mo agoHugging Face26semeru /code-code-translation-java-csharp Dataset is imported from CodeXGLUE and pre-processed using their script. Where to find in Semeru: The dataset can be found at /nfs/semeru/semeru_datasets/code_xglue/code-to-code/code-to-code-trans in Semeru CodeXGLUE -- Code2Code Translation Task Definition Code translation aims to migrate legacy software from one programming language in a platform toanother. In CodeXGLUE, given a piece of Java (C#) code, the task is to translate the code into C#… See the full description on the dataset page: https://huggingface.co/datasets/semeru/code-code-translation-java-csharp.text10K<n<100K2 likes451 downloads3y agoHugging Face27Mo7art /Stack2Graph_KG_java Java StackOverflow Knowledge Graph Summary This Hugging Face dataset repository contains the Java shard of the Stack2Graph StackOverflow Knowledge Graph. Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper. The artifact is optimized for graph-based retrieval, SPARQL analytics, and retrieval-augmented question answering over Stack Overflow content. Stack2Graph source:… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_KG_java.100M<n<1B0 likes445 downloads2mo agoHugging Face28hongliu9903 /stack_edu_javatabular10M<n<100M0 likes436 downloads1y agoHugging Face29tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes419 downloads3y agoHugging Face30dmsovetov /codeparrot-javatext10M<n<100M0 likes409 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.