datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
java_PRsStackless_Java_V2
Dataset Summary
This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability.
repobench_java_v1.1
RepoBench v1.1 (Java)
Introduction
This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns.
Resources and Links
Paper
GitHub
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.arc-stack-javascriptstack_edu_javathe-stack-java-clean
Dataset 1: TheStack - Java - Cleaned
Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language.
Target Language: Java
Dataset Size:
Training: 900,000 files
Validation: 50,000 files
Test: 50,000 files
Preprocessing:
Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.swe_smith_java_qwen3.5_35b_trajs_4369stack_edu_javascriptjava-trace-datasetarxiv_java_research_code
Dataset Card for "arxiv_java_research_code"
More Information needed
github-java-corpus
github-java-corpus
Summary
This dataset contains Java source-code text samples prepared for pretraining.
Repository
TheFinAI/github-java-corpus
Required Columns
Source: dataset name
Date: year
Text: the pure text of each sample
Token_count: the token count computed with tiktoken
Schema
Source (string)
Date (int32)
Text (string)
Token_count (int32)
Construction
The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.koch_test_2025_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 36,
"total_frames": 15986,
"total_tasks": 1,
"total_videos": 72,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:36"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.stack-filtered-pii-1M-java
Dataset Card for "stack-filtered-pii-1M-java"
More Information needed
java_open_ds
Dataset Card for "java_open_ds"
More Information needed
the-stack-v2-javajavanese-Komodo-pixelgpt
Javanese PixelGPT Dataset
This dataset contains preprocessed Javanese text data for training PixelGPT models.
Dataset Statistics
Language: Javanese (jawa)
Total samples: 401,542
Train samples: 400,726
Test samples: 816
Tokenizers
[TO BE EDITED]
Features
text_id: Document identifier
chunk_id: Chunk identifier within document
pixel_values: Rendered pixel representation of aksara text
<tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.Rebuttal-javanese-pixelgpt
Javanese PixelGPT Tokenizer Ablation Dataset
Optimized with Font Size 6 and Dynamic Trimming.
Tokenizer Schema
tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer)
tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT)
tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1)
tok_mt5: Google Multilingual Unigram (google/mt5-small)
eval_act_koch_test_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 1,
"total_frames": 855,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.UDR_Java
Dataset Card for "UDR_Java"
More Information needed
the-stack-v2-new-javaact_koch_binky_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 10,
"total_frames": 8411,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/act_koch_binky_1.megavul-vulnerability-detection-javajavascript_obfuscated_DPO_maxkoch_binky_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "koch",
"total_episodes": 1,
"total_frames": 1327,
"total_tasks":1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_binky_2.the-stack-v2-filtered-javaSoco_java_test_C2code_vulnerability_javaso101_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 1935,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_2.so101_4This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 1863,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_4.java_renaming_patch
Dataset Card for "java_renaming_patch"
More Information needed
