CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face02susnato /java_PRstabular100K<n<1M0 likes913 downloads3y agoHugging Face03AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes807 downloads1y agoHugging Face04LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes725 downloads2y agoHugging Face05ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes470 downloads10mo agoHugging Face06hongliu9903 /stack_edu_javatabular10M<n<100M0 likes436 downloads1y agoHugging Face07tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes419 downloads3y agoHugging Face08ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes363 downloads3y agoHugging Face09JWei05 /swe_smith_java_qwen3.5_35b_trajs_4369tabular1K<n<10K0 likes298 downloads6mo agoHugging Face10HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes289 downloads2mo agoHugging Face11LarsEckart /approvaltests-java-sessions Coding agent session traces for LarsEckart/approvaltests-java-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.tabulartext-generationn<1K2 likes279 downloads6mo agoHugging Face12hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes222 downloads1y agoHugging Face13iris-sast /CWE-Bench-Java CWE-Bench-Java This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.tabular1K<n<10K1 likes200 downloads1y agoHugging Face14claudios /java-trace-datasettabular100K<n<1M0 likes172 downloads3y agoHugging Face15TheFinAI /github-java-corpus github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes153 downloads6mo agoHugging Face16javasoup /koch_test_2025_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 36, "total_frames": 15986, "total_tasks": 1, "total_videos": 72, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.tabularrobotics10K<n<100K0 likes148 downloads1y agoHugging Face17AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes147 downloads3y agoHugging Face18loubnabnl /stack-filtered-pii-1M-java Dataset Card for "stack-filtered-pii-1M-java" More Information needed tabular1M<n<10M0 likes115 downloads4y agoHugging Face19shyamsubbu /java_open_ds Dataset Card for "java_open_ds" More Information needed tabular100K<n<1M0 likes110 downloads3y agoHugging Face20Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes107 downloads7mo agoHugging Face21Reset23 /the-stack-v2-javatabular1M<n<10M0 likes103 downloads2y agoHugging Face22Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes82 downloads6mo agoHugging Face23Zaib /java-vulnerabilitytabular1K<n<10K8 likes51 downloads4y agoHugging Face24javasoup /eval_act_koch_test_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 855, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.tabularrobotics1K<n<10K0 likes50 downloads1y agoHugging Face25ProgramGym /Candidates_unmatched_javatabular1K<n<10K0 likes43 downloads28d agoHugging Face26KaiLv /UDR_Java Dataset Card for "UDR_Java" More Information needed tabular100K<n<1M1 likes42 downloads3y agoHugging Face27Reset23 /the-stack-v2-new-javatabular1M<n<10M0 likes42 downloads2y agoHugging Face28javasoup /act_koch_binky_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 10, "total_frames": 8411, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/act_koch_binky_1.tabularrobotics1K<n<10K0 likes38 downloads1y agoHugging Face29qiubinjun /CWE-Bench-Java CWE-Bench-Java This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/qiubinjun/CWE-Bench-Java.tabular1K<n<10K0 likes38 downloads4mo agoHugging Face30JavaneseHonorifics /Unggah-Ungguh Javanese Honorifics Dataset (Unggah-Ungguh - Released Version) The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context. Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.tabular1K<n<10K0 likes37 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.