CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes2.4k downloads17d agoHugging Face02ajibawa-2023 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/JavaScript-Code-Large.texttext-generation1M<n<10M33 likes2.3k downloads7mo agoHugging Face03ajibawa-2023 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/ajibawa-2023/Java-Code-Large.texttext-generation10M<n<100M34 likes1.1k downloads7mo agoHugging Face04marcelo1234 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/marcelo1234/Java-Code-Large.texttext-generation10M<n<100M2 likes598 downloads4mo agoHugging Face05thedruid831 /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/thedruid831/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes589 downloads7mo agoHugging Face06sanjaykz /Java-codetexttext-generation100K<n<1M2 likes534 downloads11mo agoHugging Face07liuhangbiao /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/liuhangbiao/JavaScript-Code-Large.texttext-generation1M<n<10M0 likes344 downloads6mo agoHugging Face08Ujjwal-Tyagi /JavaScript-Code-LargeJavaScript-Code-Large JavaScript-Code-Large is a large-scale corpus of JavaScript source code comprising around 5 million JavaScript files. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis for the JavaScript ecosystem. By providing a high-volume, language-specific corpus, JavaScript-Code-Large enables systematic experimentation in JavaScript-focused model training, domain adaptation… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/JavaScript-Code-Large.texttext-generation1M<n<10M1 likes312 downloads6mo agoHugging Face09Ujjwal-Tyagi /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Ujjwal-Tyagi/Java-Code-Large.texttext-generation10M<n<100M0 likes310 downloads6mo agoHugging Face10HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes303 downloads2mo agoHugging Face11LarsEckart /approvaltests-java-sessions Coding agent session traces for LarsEckart/approvaltests-java-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.tabulartext-generationn<1K2 likes283 downloads6mo agoHugging Face12random-long-int /Java_method2test_chatml Java Method to Test ChatML This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}]. Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here: To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters: The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.textquestion-answering100K<n<1M2 likes127 downloads2y agoHugging Face13javohirmat /uzbek-legal-corpus Uzbek Legal Corpus (Oʻzbek huquqiy korpusi) 25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles. Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI. Contents Path What it is data/raw/*.txt Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.texttext-retrieval1K<n<10K1 likes111 downloads2mo agoHugging Face14Ayan987 /Java-Code-LargeJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Ayan987/Java-Code-Large.texttext-generation10M<n<100M0 likes102 downloads7mo agoHugging Face15Naholav /llama3.2-java-codegen-90sft-10meta-claude-v1 LLaMA 3.2 Java Code Generation Dataset (90% SFT, 10% Meta Annotated with Claude) This dataset contains 100,000 examples for Java method generation based on natural language instructions. It is built from the CodeXGLUE text-to-code dataset and designed to support both pure supervised fine-tuning (SFT) and reflection-based meta-learning approaches using Claude 4 Sonnet as the critique model. 🚀 Trained Models Two models have been trained on this dataset: SFT Model:… See the full description on the dataset page: https://huggingface.co/datasets/Naholav/llama3.2-java-codegen-90sft-10meta-claude-v1.texttext-generation100K<n<1M1 likes88 downloads1y agoHugging Face16israellaguan /axay-javascript-dataset-pn DPO JavaScript Dataset This repository contains a modified version of the JavaScript dataset originally sourced from axay/javascript-dataset-pn. The dataset has been adapted to fit the DPO (Dynamic Programming Object) format, making it compatible with the LLaMA-Factory project. License This dataset is licensed under the Apache 2.0 License. Dataset Overview The dataset consists of JavaScript code snippets that have been restructured and enhanced for use in… See the full description on the dataset page: https://huggingface.co/datasets/israellaguan/axay-javascript-dataset-pn.texttext-generationn<1K0 likes33 downloads2y agoHugging Face17javatask /eidas eIDAS Terminology Dataset Dataset Description Overview The EiDAS Terminology dataset is a comprehensive collection of terms and abbreviations related to electronic identification and trust services for electronic transactions in the European Single Market (eIDAS). This dataset provides clear definitions and explanations of various terms, making it an essential resource for researchers and practitioners in digital identity and security. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/javatask/eidas.tabulartext-generationn<1K0 likes28 downloads3y agoHugging Face18nlpctx /java_optimisation Java Optimisation Dataset 📌 Overview The Java Optimisation Dataset is a curated collection of paired Java code snippets designed for training and evaluating machine learning models on code optimization and refactoring tasks. Each example consists of: Input: Unoptimized or verbose Java code Output: Optimized, cleaner, or more efficient equivalent code This dataset is intended for training sequence-to-sequence models such as CodeT5 for automated code improvement.… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/java_optimisation.texttext-generation1K<n<10K0 likes25 downloads6mo agoHugging Face19rakhman-llm /XCodeEval-Java-Datasettexttext-generation100K<n<1M0 likes24 downloads2y agoHugging Face20nadiamaqbool81 /java_code_instructions_1.178k_alpacaInstruction set for text to java code generation. This is the subset of concode dataset. texttext-generation1K<n<10K3 likes23 downloads3y agoHugging Face21junwatu /javanese-multilingual-lexicongated Javanese Multilingual Lexicon A multilingual lexicon dataset of 202,912 entries derived from Sastra.org, a digital archive of Javanese literary heritage maintained by Yayasan Sastra Lestari. The dataset compiles 32 historical lexicographic sources spanning from the 1830s to 2010, covering Javanese, Old Javanese (Kawi), Indonesian, Dutch, English, and French. Dataset Structure The dataset is provided as JSONL files. The full corpus is in data/leksikon_all.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/javanese-multilingual-lexicon.texttranslation100K<n<1M1 likes16 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.