CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zxliu /ReAPR-Automatic-Program-Repair-via-Retrieval-Augmented-Large-Language-ModelsThis is the Retrieval dataset used in the paper "ReAPR: Automatic Program Repair via Retrieval-Augmented Large Language Models" text100K<n<1M3 likes162 downloads2y agoHugging Face02MultiMind-SemEval2025 /Augmented_MultiClaim_FactCheck_Retrieval Augmented MultiClaim FactCheck Retrieval Dataset 1. Dataset Summary This dataset is a collection of social media posts that have been augmented using a large language model (GPT-4o). The original dataset was sourced from the paper Multilingual Previously Fact-Checked Claim Retrieval by Matúš Pikuliak et al. (2023). You can access the original dataset from here. The dataset is used for improving the ability to comprehend content across multiple languages by integrating… See the full description on the dataset page: https://huggingface.co/datasets/MultiMind-SemEval2025/Augmented_MultiClaim_FactCheck_Retrieval.text10K<n<100K1 likes141 downloads1y agoHugging Face03farabi-lab /Retrieval-Augmented-Question-Answeringgated 🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context Dataset Summary Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language. This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.textquestion-answering10K<n<100K0 likes11 downloads2mo agoHugging Face04farabi-lab /API_Discovery_Retrieval_Augmented_Callinggated 🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset Dataset Summary Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution. The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.texttext-generation1K<n<10K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.