CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face02susnato /java_PRstabular100K<n<1M0 likes914 downloads3y agoHugging Face03AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes862 downloads1y agoHugging Face04LaughingLogits /Stackless_Java_V2 Dataset Summary This is the dataset used for the training of the AP-MAE models, it is a subset of The Heap, we release it for reproducability. tabular1M<n<10M0 likes725 downloads2y agoHugging Face05tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes484 downloads3y agoHugging Face06ThomasTheMaker /arc-stack-javascripttabular10M<n<100M0 likes470 downloads11mo agoHugging Face07hongliu9903 /stack_edu_javatabular10M<n<100M0 likes440 downloads1y agoHugging Face08ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes354 downloads3y agoHugging Face09JWei05 /swe_smith_java_qwen3.5_35b_trajs_4369tabular1K<n<10K0 likes297 downloads6mo agoHugging Face10HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes294 downloads2mo agoHugging Face11LarsEckart /approvaltests-java-sessions Coding agent session traces for LarsEckart/approvaltests-java-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.tabulartext-generationn<1K2 likes283 downloads6mo agoHugging Face12javifer /ultra_short_form_generations_labeled_2tabular10K<n<100K0 likes250 downloads1y agoHugging Face13hongliu9903 /stack_edu_javascripttabular10M<n<100M0 likes223 downloads1y agoHugging Face14claudios /java-trace-datasettabular100K<n<1M0 likes212 downloads3y agoHugging Face15iris-sast /CWE-Bench-Java CWE-Bench-Java This repository contains the dataset CWE-Bench-Java presented in the paper LLM-Assisted Static Analysis for Detecting Security Vulnerabilities. At a high level, this dataset contains 120 CVEs spanning 4 CWEs, namely path-traversal, OS-command injection, cross-site scripting, and code-injection. Each CVE includes the buggy and fixed source code of the project, along with the information of the fixed files and functions. We provide the seed information for each CVE in… See the full description on the dataset page: https://huggingface.co/datasets/iris-sast/CWE-Bench-Java.tabular1K<n<10K1 likes200 downloads1y agoHugging Face16AlgorithmicResearchGroup /arxiv_java_research_code Dataset Card for "arxiv_java_research_code" More Information needed tabular100K<n<1M1 likes183 downloads3y agoHugging Face17TheFinAI /github-java-corpus github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes168 downloads6mo agoHugging Face18JavisVerse /JavisBench JavisBench: A Challenging Benchmark for for Joint Audio-Video Generation (JAVG) Evaluation As released in HuggingFace, JavisBench is a comprehensive and challenging benchmark for evaluating text-to-audio-video generation models.It covers multiple aspects of generation quality, semantic alignment, and temporal synchrony, enabling thorough assessment in both controlled and real-world scenarios. Installation Install necessary packages: cd /path/to/JavisDiT pip install… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/JavisBench.tabular10K<n<100K0 likes163 downloads1y agoHugging Face19JavisVerse /AV-DPO AV-DPO AV-DPO is the audio-video preference dataset used in Stage 3 of JavisDiT++. It contains chosen-rejected sounding-video pairs selected jointly across audio quality and text alignment, video quality and text alignment, and audio-video semantic and temporal alignment. Release summary Item Count Preference pairs (train.csv) 23,671 Self-contained generated-only pairs (train_generated_only.csv) 6,138 Unique released generated sounding videos 29,809… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/AV-DPO.tabulartext-to-video10K<n<100K0 likes163 downloads24d agoHugging Face20javasoup /koch_test_2025_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 36, "total_frames": 15986, "total_tasks": 1, "total_videos": 72, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:36" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/koch_test_2025_2.tabularrobotics10K<n<100K0 likes147 downloads1y agoHugging Face21loubnabnl /stack-filtered-pii-1M-java Dataset Card for "stack-filtered-pii-1M-java" More Information needed tabular1M<n<10M0 likes115 downloads4y agoHugging Face22shyamsubbu /java_open_ds Dataset Card for "java_open_ds" More Information needed tabular100K<n<1M0 likes111 downloads3y agoHugging Face23Reset23 /the-stack-v2-javatabular1M<n<10M0 likes103 downloads2y agoHugging Face24Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes100 downloads7mo agoHugging Face25Exqrch /Rebuttal-javanese-pixelgpt Javanese PixelGPT Tokenizer Ablation Dataset Optimized with Font Size 6 and Dynamic Trimming. Tokenizer Schema tok_grapheme: Language-specific Grapheme BPE (izzako/javanese-llama-tokenizer) tok_llama2: Standard Llama-2 BPE (ernie-research/DualGPT) tok_komodo: SEA-Optimized BPE (yellow-ai-central/komodo-7b-v1) tok_mt5: Google Multilingual Unigram (google/mt5-small) tabular100K<n<1M0 likes83 downloads6mo agoHugging Face26Javtor /biomedical-topic-categorization-cased Dataset Card for "biomedical-topic-categorization-cased" More Information needed tabular1M<n<10M0 likes79 downloads4y agoHugging Face27Zaib /java-vulnerabilitytabular1K<n<10K8 likes51 downloads4y agoHugging Face28javiersgjavi /fineweb-1BT FineWeb-1BT: 1 Billion Token Subset Dataset Description FineWeb-1BT is a carefully curated 1 billion token subset of the HuggingFaceFW/fineweb dataset, sampled exclusively from the official 10BT FineWeb subset. It was created using true random sampling to ensure unbiased representation across the entire 10BT corpus. Provenance: This 1BT subset is a uniform sample drawn from the official FineWeb 10BT split (not from the full raw stream), preserving its language/quality… See the full description on the dataset page: https://huggingface.co/datasets/javiersgjavi/fineweb-1BT.tabulartext-generation1M<n<10M0 likes47 downloads1y agoHugging Face29javasoup /eval_act_koch_test_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "koch", "total_episodes": 1, "total_frames": 855, "total_tasks": 1, "total_videos": 2, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/javasoup/eval_act_koch_test_2.tabularrobotics1K<n<10K0 likes46 downloads1y agoHugging Face30KaiLv /UDR_Java Dataset Card for "UDR_Java" More Information needed tabular100K<n<1M1 likes45 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.