datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
JavaError-QA
JErrRAG-Eval-800
JErrRAG-Eval-800 is the public benchmark release aligned with the paper's final canonical dataset and non-anonymous archival record.
This Hugging Face repository contains:
java_error_qa_v2/: the canonical public benchmark package
paper_online_artifacts/: the paper-facing supplementary artifacts and reproduction bundles
SHA256SUMS.txt: release-side hash anchors referenced by the paper
Dataset Summary
Total records: 800
Split sizes: train=639… See the full description on the dataset page: https://huggingface.co/datasets/HTJ008/JavaError-QA.ClimateFund
ClimateFund: An Annotated Dataset of Climate Mitigation Projects for Supporting Question Answering
This repository contains a dataset based on funding proposals of 21 climate mitigation projects, submitted to the Green Climate Fund (GCF).
Climate mitigation documentation is challenging to parse and understand, due to the length of this documents, their multi-modality (commonly comprising tables, figures and
free text), and their highly technical and domain-specific content.… See the full description on the dataset page: https://huggingface.co/datasets/JavierSanzCruza/ClimateFund.Java_method2test_chatml
Java Method to Test ChatML
This dataset is based on the methods2test dataset from Microsoft. It follows the ChatML template format: [{'role': '', 'content': ''}, {...}].
Originally, methods2test contains only Java methods at different levels of granularity along with their corresponding test cases. The different focal method segmentations are illustrated here:
To simulate a conversation between a Java developer and an AI assistant, I introduce two key parameters:
The prompt… See the full description on the dataset page: https://huggingface.co/datasets/random-long-int/Java_method2test_chatml.uzbek-legal-corpus
Uzbek Legal Corpus (Oʻzbek huquqiy korpusi)
25 cleaned, audited, machine-readable texts of the Republic of Uzbekistan's Constitution, 20 codes, and 4 major laws — in Uzbek (Latin script), sourced from the official National Database of Legislation (lex.uz). Snapshot: July 2026. 7,368 articles.
Built for NLP / LLM / RAG work on Uzbek legal text. Released by Tomaris AI.
Contents
Path
What it is
data/raw/*.txt
Cleaned plain text, one file per code/law… See the full description on the dataset page: https://huggingface.co/datasets/javohirmat/uzbek-legal-corpus.javanese-hotel-receptionist-qna
Dataset Card for Alpaca-Cleaned
Repository: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna
Dataset Description
This synthetic dataset is designed for training and fine-tuning language models to handle customer service inquiries in a hotel setting using Javanese language. The data has been generated in the Alpaca format to assist in building models that can follow customer service-related instructions and generate appropriate responses. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/7out/javanese-hotel-receptionist-qna.
