CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazonScience /migration-bench-java-full MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-full.tabulartext-generation1K<n<10K4 likes4.4k downloads1y agoHugging Face02AmazonScience /migration-bench-java-selected MigrationBench 1. 📖 Overview 🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages. Current and initial release includes java 8 repositories with the maven build system… See the full description on the dataset page: https://huggingface.co/datasets/AmazonScience/migration-bench-java-selected.tabulartext-generationn<1K7 likes807 downloads1y agoHugging Face03tianyang /repobench_java_v1.1 RepoBench v1.1 (Java) Introduction This dataset presents the Java portion of RepoBench v1.1 (ICLR 2024). The data encompasses a collection from GitHub, spanning the period from October 6th to December 31st, 2023. With a commitment to data integrity, we've implemented a deduplication process based on file content against the Stack v2 dataset (coming soon), aiming to mitigate data leakage and memorization concerns. Resources and Links Paper GitHub Dataset… See the full description on the dataset page: https://huggingface.co/datasets/tianyang/repobench_java_v1.1.tabulartext-generation10K<n<100K0 likes419 downloads3y agoHugging Face04ammarnasr /the-stack-java-clean Dataset 1: TheStack - Java - Cleaned Description: This dataset is drawn from TheStack Corpus, an open-source code dataset with over 3TB of GitHub data covering 48 programming languages. We selected a small portion of this dataset to optimize smaller language models for Java, a popular statically typed language. Target Language: Java Dataset Size: Training: 900,000 files Validation: 50,000 files Test: 50,000 files Preprocessing: Selected Java as the target language due to its… See the full description on the dataset page: https://huggingface.co/datasets/ammarnasr/the-stack-java-clean.tabulartext-generation100K<n<1M13 likes363 downloads3y agoHugging Face05HTJ008 /JavaError-QA JErrRAG-Eval-800 JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record. This Hugging Face repository contains: java_error_qa_v2/: the canonical public benchmark package paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles SHA256SUMS.txt: release-side hash anchors referenced by the paper Dataset Summary Total records: 800 Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.documentquestion-answeringn<1K2 likes289 downloads2mo agoHugging Face06LarsEckart /approvaltests-java-sessions Coding agent session traces for LarsEckart/approvaltests-java-sessions This dataset contains redacted coding agent session traces collected while working on git@github.com:approvals/ApprovalTests.Java.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where… See the full description on the dataset page: https://huggingface.co/datasets/LarsEckart/approvaltests-java-sessions.tabulartext-generationn<1K2 likes279 downloads6mo agoHugging Face07TheFinAI /github-java-corpus github-java-corpus Summary This dataset contains Java source-code text samples prepared for pretraining. Repository TheFinAI/github-java-corpus Required Columns Source: dataset name Date: year Text: the pure text of each sample Token_count: the token count computed with tiktoken Schema Source (string) Date (int32) Text (string) Token_count (int32) Construction The dataset was built from streamed archive processing into… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/github-java-corpus.tabulartext-generation1M<n<10M0 likes153 downloads6mo agoHugging Face08Exqrch /javanese-Komodo-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers [TO BE EDITED] Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel representation of aksara text <tokenizer_name>_token_ids: Token IDs from grapheme-based… See the full description on the dataset page: https://huggingface.co/datasets/Exqrch/javanese-Komodo-pixelgpt.tabulartext-generation100K<n<1M0 likes107 downloads7mo agoHugging Face09javatask /eidas eIDAS Terminology Dataset Dataset Description Overview The EiDAS Terminology dataset is a comprehensive collection of terms and abbreviations related to electronic identification and trust services for electronic transactions in the European Single Market (eIDAS). This dataset provides clear definitions and explanations of various terms, making it an essential resource for researchers and practitioners in digital identity and security. Languages The… See the full description on the dataset page: https://huggingface.co/datasets/javatask/eidas.tabulartext-generationn<1K0 likes15 downloads3y agoHugging Face10izzako /javanese-pixelgpt Javanese PixelGPT Dataset This dataset contains preprocessed Javanese text data for training PixelGPT models. Dataset Statistics Language: Javanese (jawa) Total samples: 401,542 Train samples: 400,726 Test samples: 816 Tokenizers Grapheme tokenizer: izzako/javanese-llama-tokenizer LLaMA tokenizer: ernie-research/DualGPT Features text_id: Document identifier chunk_id: Chunk identifier within document pixel_values: Rendered pixel… See the full description on the dataset page: https://huggingface.co/datasets/izzako/javanese-pixelgpt.tabulartext-generation100K<n<1M0 likes12 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.