datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Retrieval-Augmented-Question-Answering
🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context
Dataset Summary
Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language.
This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.API_Discovery_Retrieval_Augmented_Calling
🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset
Dataset Summary
Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution.
The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.
