datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
javadoc
Java Method to JavaDoc Dataset
Overview
This dataset is designed for the specific task of fine-tuning a model to generate JavaDoc documentation for Java methods.
The dataset contains pairs of Java methods and their corresponding JavaDoc comments, facilitating the model's learning of the relationship between code structure and its descriptive documentation.
Data Collection
The data is collected from various open-source Java projects hosted on platforms such as… See the full description on the dataset page: https://huggingface.co/datasets/Michael22/javadoc.RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.instruction-dataset-indo-java-sunda-bali-gayo-batak-alas-minang-betawioasst-javanese
Dataset Summary
We translated the OpenAssistant Conversations (OASST) dataset into Javanese using Meta's No Language Left Behind (NLLB) model.
Why Javanese?
Javanese is spoken by over 90 million people on the island of Java in Indonesia. While its prevalence is comparable to other widely spoken languages, such as Vietnamese and Turkish, its representation in current large language model (LLM) chatbots remains limited. By translating this dataset, we aim to enhance the… See the full description on the dataset page: https://huggingface.co/datasets/richardcsuwandi/oasst-javanese.
