datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
migration-bench-java-full
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.java_PRsmigration-bench-java-selected
MigrationBench
1. 📖 Overview
🤗 MigrationBench
is a large-scale code migration benchmark dataset at the repository level,
across multiple programming languages.
Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.Stackless_Java_V2
Dataset Summary
This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability.
arc-stack-javascriptstack_edu_javarepobench_java_v1.1
RepoBench v1.1 (Java)
Introduction
This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns.
Resources and Links
Paper
GitHub
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.the-stack-java-clean
Dataset 1: TheStack - Java - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language.
Target Language: Java
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.swe_smith_java_qwen3.5_35b_trajs_4369JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.approvaltests-java-sessions
Coding agent session traces for LarsEckart/approvaltests-java-sessions
This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review.
Data description
Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.stack_edu_javascriptCWE-Bench-Java
CWE-Bench-Java
This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.java-trace-datasetgithub-java-corpus
github-java-corpus
Summary
This dataset contains Java source-code text samples prepared for pretraining.
Repository
TheFinAI/github-java-corpus
Required Columns
Source: dataset name
Date: year
Text: the pure text of each sample
Token_count: the token count computed with tiktoken
Schema
Source (string)
Date (int32)
Text (string)
Token_count (int32)
Construction
The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.koch_test_2025_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 36,
"total_frames": 15986,
"total_tasks": 1,
"total_videos": 72,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:36"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.arxiv_java_research_code
Dataset Card for "arxiv_java_research_code"
More Information needed
stack-filtered-pii-1M-java
Dataset Card for "stack-filtered-pii-1M-java"
More Information needed
java_open_ds
Dataset Card for "java_open_ds"
More Information needed
javanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.the-stack-v2-javaRebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
java-vulnerabilityeval_act_koch_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 1,
"total_frames": 855,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.Candidates_unmatched_javaUDR_Java
Dataset Card for "UDR_Java"
More Information needed
the-stack-v2-new-javaact_koch_binky_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 10,
"total_frames": 8411,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/act_koch_binky_1.CWE-Bench-Java
CWE-Bench-Java
This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/qiubinjun/CWE-Bench-Java.Unggah-Ungguh
Javanese Honorifics Dataset (Unggah-Ungguh - Released Version)
The Javanese language, spoken by over 98 million people, features a distinctive honorific system known as Unggah-Ungguh Basa. In this dataset we present UNGGAH-UNGGUH, a carefully curated dataset designed to encapsulate the nuances of Unggah-Ungguh Basa, the Javanese speech etiquette framework that dictates the choice of words and phrases based on social hierarchy and context.
Paper: https://arxiv.org/pdf/2502.20864… See the full description on the dataset page: https://huggingface.co/datasets/JavaneseHonorifics/Unggah-Ungguh.
