CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01susnato /java_PRstabular100K<n<1M0 likes914 downloads3y agoHugging Face02LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes725 downloads2y agoHugging Face03tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes484 downloads3y agoHugging Face04ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes470 downloads11mo agoHugging Face05hongliu9903 /stack_edu_javatabular10M<n<100M0 likes440 downloads1y agoHugging Face06ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes354 downloads3y agoHugging Face07JWei05 /swe_smith_java_qwen3.5_35b_trajs_4369tabular1K<n<10K0 likes297 downloads6mo agoHugging Face08hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes223 downloads1y agoHugging Face09claudios /java-trace-datasettabular100K<n<1M0 likes212 downloads3y agoHugging Face10AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes183 downloads3y agoHugging Face11TheFinAI /github-java-corpus github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes168 downloads6mo agoHugging Face12javasoup /koch_test_2025_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 36, "total_frames": 15986, "total_tasks": 1, "total_videos": 72, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.tabularrobotics10K<n<100K0 likes147 downloads1y agoHugging Face13loubnabnl /stack-filtered-pii-1M-java Dataset Card for "stack-filtered-pii-1M-java" More Information needed tabular1M<n<10M0 likes115 downloads4y agoHugging Face14shyamsubbu /java_open_ds Dataset Card for "java_open_ds" More Information needed tabular100K<n<1M0 likes111 downloads3y agoHugging Face15Reset23 /the-stack-v2-javatabular1M<n<10M0 likes103 downloads2y agoHugging Face16Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes100 downloads7mo agoHugging Face17Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes83 downloads6mo agoHugging Face18javasoup /eval_act_koch_test_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 855, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.tabularrobotics1K<n<10K0 likes46 downloads1y agoHugging Face19KaiLv /UDR_Java Dataset Card for "UDR_Java" More Information needed tabular100K<n<1M1 likes45 downloads3y agoHugging Face20Reset23 /the-stack-v2-new-javatabular1M<n<10M0 likes42 downloads2y agoHugging Face21javasoup /act_koch_binky_1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 10, "total_frames": 8411, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/act_koch_binky_1.tabularrobotics1K<n<10K0 likes38 downloads1y agoHugging Face22athrv /megavul-vulnerability-detection-javatabular10K<n<100K1 likes35 downloads11mo agoHugging Face23mxzoo /javascript_obfuscated_DPO_maxtabular10K<n<100K1 likes30 downloads8mo agoHugging Face24javasoup /koch_binky_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 1327, "total_tasks":1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_binky_2.tabularrobotics1K<n<10K0 likes29 downloads11mo agoHugging Face25Reset23 /the-stack-v2-filtered-javatabular100K<n<1M0 likes28 downloads1y agoHugging Face26buelfhood /Soco_java_test_C2tabular1K<n<10K0 likes26 downloads3y agoHugging Face27maddyrucos /code_vulnerability_javatabularn<1K0 likes25 downloads2y agoHugging Face28javadcc /so101_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 1935, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_2.tabularrobotics1K<n<10K0 likes25 downloads10mo agoHugging Face29javadcc /so101_4This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 1863, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javadcc/so101_4.tabularrobotics1K<n<10K0 likes25 downloads10mo agoHugging Face30thiomajid /java_renaming_patch Dataset Card for "java_renaming_patch" More Information needed tabularn<1K0 likes24 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.