datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ragtruth-qa-ko-alpaca
Dataset Card for Dataset Name
ragtruth-qa 데이터셋을 gpt-4o를 이용하여 한글로 번역 한 데이터셋.
ragtruth-qa-ko 데이터 셋을 alpaca 포맷으로 변환.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Language(s) (NLP): [한국어]
License: [미정]
Dataset Sources [optional]
Repository: [https://huggingface.co/datasets/flowaicom/formatted-ragtruth-qa]
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use
[More… See the full description on the dataset page: https://huggingface.co/datasets/aiyets/ragtruth-qa-ko-alpaca.toolace_ragtruth_synthetic
ToolACE RAGTruth style synthetic hallucination dataset
This repository contains a synthetic span annotated hallucination detection dataset for tool calling responses.
The data was derived from Team ACE ToolACE conversations. Only grounded instances with a user query, tool call, tool output, and final assistant answer were used.
Motivation
The dataset targets hallucination detection in tool augmented generation, where a final assistant response should be grounded in the… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace_ragtruth_synthetic.ragtruth-qa-ko
Dataset Card for Dataset Name
ragtruth-qa 데이터셋을 gpt-4o를 이용하여 한글로 번역 한 데이터셋.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Language(s) (NLP): [한국어]
License: [미정]
Dataset Sources [optional]
Repository: [https://huggingface.co/datasets/flowaicom/formatted-ragtruth-qa]
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/Yettiesoft/ragtruth-qa-ko.toolace-ragtruth-style-hallucinations
ToolACE RAGTruth-style Tool Hallucination Dataset
This dataset was built from ToolACE tool-use dialogues and converted into a RAGTruth-style format for hallucination detection in tool calling.
Task
Given:
query: user query
context: tool response
output: final assistant answer
the goal is to classify whether the answer is grounded in the tool output or belongs to one of three hallucination types.
Labels
clean
tool_output_conflict
overgeneration… See the full description on the dataset page: https://huggingface.co/datasets/Ali-Bhai/toolace-ragtruth-style-hallucinations.RAGTruth-Hallucinations
ToolACE
ToolACE is an automatic agentic pipeline designed to generate Accurate, Complex, and divErse tool-learning data.
ToolACE leverages a novel self-evolution synthesis process to curate a comprehensive API pool of 26,507 diverse APIs.
Dialogs are further generated through the interplay among multiple agents, guided by a formalized thinking process.
To ensure data accuracy, we implement a dual-layer verification system combining rule-based and model-based checks.
More details… See the full description on the dataset page: https://huggingface.co/datasets/drond0174/RAGTruth-Hallucinations.lynx-70b-instruct-ragtruth-generationsRAGHal-RAGTruth-auto-en-v1
RAGHal-RAGTruth-auto-en-v1
Auto-annotated token-level hallucination training set for RAG faithfulness detection.
Built on top of RAGTruth train responses (original answers + prompts), with automatic span labels produced by our annotation pipeline.No human span labels are used in this train set. Human annotations appear only in the official RAGTruth test split (for evaluation of models trained on this data).
This is the training corpus behind ZaandaTeika/RAGHal-large-en-v1.… See the full description on the dataset page: https://huggingface.co/datasets/ZaandaTeika/RAGHal-RAGTruth-auto-en-v1.
