datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.javadoc
Java Method to JavaDoc Dataset
Overview
This dataset is designed for the specific task of fine-tuning a model to generate JavaDoc documentation for Java methods.
The dataset contains pairs of Java methods and their corresponding JavaDoc comments, facilitating the model's learning of the relationship between code structure and its descriptive documentation.
Data Collection
The data is collected from various open-source Java projects hosted on platforms such as… See the full description on the dataset page: https://huggingface.co/datasets/Michael22/javadoc.Java_method2test_chatml
Java Method to Test ChatML
This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}].
Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here:
To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters:
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.Synthetic_Java_Dialog_And_Programs_LLM_TrainingThe full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Java Programming Examples Dataset
Dataset Description
This dataset contains 8 distinct Java programs with 10 conversational examples each, synthetically generated from a larger dataset of 80+ programs. Each program has 10,000 variants, providing a diverse set of Java code examples covering various programming… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/Synthetic_Java_Dialog_And_Programs_LLM_Training.RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawioasst-javanese
Dataset Summary
We translated the OpenAssistant Conversations (OASST) dataset into Javanese using Meta's No Language Left Behind (NLLB) model.
Why Javanese?
Javanese is spoken by over 90 million people on the island of Java in Indonesia. While its prevalence is comparable to other widely spoken languages, such as Vietnamese and Turkish, its representation in current large language model (LLM) chatbots remains limited. By translating this dataset, we aim to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/richardcsuwandi/oasst-javanese.Stack2Graph_VD_javascript
Javascript StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Javascript shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_javascript.east_java_dialect_instruct
Complaints From The East Javanese Dialect community
This dataset created manually by humans with reference to public complaints in the comments column of the local government's Instagram account and another platform like X and TikTok Comments.
Stack2Graph_VD_java
Java StackOverflow Vector Dataset
Summary
This Hugging Face dataset repository contains the Java shard of the Stack2Graph vector-database component as restorable Qdrant artifacts plus portable Parquet fallback files.
Hugging Face uses one dataset repository per programming language, so this repository is directly cloneable without an extra top-level archive wrapper.
The artifacts are intended for semantic and hybrid retrieval, graph entry-point finding, and… See the full description on the dataset page: https://huggingface.co/datasets/Mo7art/Stack2Graph_VD_java.javanese-hotel-receptionist-qna
Dataset Card for Alpaca-Cleaned
Repository: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna
Dataset Description
This synthetic dataset is designed for training and fine-tuning language models to handle customer service inquiries in a hotel setting using Javanese language. The data has been generated in the Alpaca format to assist in building models that can follow customer service-related instructions and generate appropriate responses. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna.
