CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Emulated-Inc /long-context-retrieval-training-pool Long context retrieval training pool Long prompts with short, checkable answers. Each row is one complete message: a task instruction, a long body of text that hides what the question is about, and the question itself, together with every string an answer has to contain for it to be right. The bodies run from four thousand to thirty-two thousand tokens. Three sources, laid out twice. Train on either layer or on both. pool.jsonl Every source rewritten into one… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/long-context-retrieval-training-pool.texttext-generation10K<n<100K1 likes125 downloads10d agoHugging Face02AmanPriyanshu /tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source Tool-Reasoning SFT — RLVR Retrieval Source Trajectories 156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question. Author: Aman Priyanshu Source Environments Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.tabulartext-generation100K<n<1M0 likes74 downloads6mo agoHugging Face03DinoDS /retrieval_grounding Dino Data Retrieval Grounding Preview What This Dataset Is This dataset is a focused retrieval-grounding preview built from four Dino Data capability slices: search trigger detection grounded search integration history search trigger history search integration The goal is to train or inspect assistant behavior around two connected problems: deciding when retrieval or history lookup is needed generating answers that stay grounded to supplied evidence or prior thread… See the full description on the dataset page: https://huggingface.co/datasets/DinoDS/retrieval_grounding.tabularquestion-answeringn<1K0 likes39 downloads5mo agoHugging Face04ReactiveAI /passkey-retrieval ReactiveAI / passkey-retrieval (Interactions Format) Conversational (in RxLM Interactions Format) retrieval (Passkey / Needle In a Haystack type) dataset, filtered and transformed from grimulkan/passkey-retrieval Subsets to-4k - 3-step instruct examples with first (context) message with up to 4k tokens to-4k-reasoning - 3-step reasoning examples with first (context) query with up to 4k tokens and all the interaction (with reasoning) up to 8k tokens to-8k - 3-step… See the full description on the dataset page: https://huggingface.co/datasets/ReactiveAI/passkey-retrieval.texttext-retrieval1K<n<10K1 likes32 downloads5mo agoHugging Face05leonli66 /stage3-synthetic-structured-retrieval Stage 3 Synthetic Structured-Retrieval Agents Native search-tool trajectories generated by Qwen/Qwen3-235B-A22B-Instruct-2507 for LCLM Stage-3 agent post-training. The default config contains only traces that passed programmatic evidence and answer verification. Harvest Accepted traces: 82 Native search calls: 179 Compressed tool-observation traces: 41 Uncompressed traces: 41 Task-ID overlap between pilot and collection batch: 0 Family counts are 20 latest-state… See the full description on the dataset page: https://huggingface.co/datasets/leonli66/stage3-synthetic-structured-retrieval.tabulartext-generationn<1K0 likes32 downloads1mo agoHugging Face06CausalLM /Retrieval-SFT-Chatgated Retrieval-Based Multi-Turn Chat SFT Synthetic Data A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture. In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.textquestion-answering100K<n<1M61 likes20 downloads2y agoHugging Face07nics-efc /MoA_Long_Retrievaltabularquestion-answering1K<n<10K4 likes17 downloads2y agoHugging Face08farabi-lab /Retrieval-Augmented-Question-Answeringgated 🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context Dataset Summary Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language. This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.textquestion-answering10K<n<100K0 likes16 downloads2mo agoHugging Face09farabi-lab /API_Discovery_Retrieval_Augmented_Callinggated 🇰🇿 Kazakh API Discovery and Tool Retrieval Dataset Dataset Summary Kazakh API Discovery and Tool Retrieval Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in agentic AI workflows that require API discovery, tool documentation retrieval, function calling, and multi-step tool execution. The dataset focuses on scenarios where the assistant must first inspect or retrieve API documentation before calling the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/API_Discovery_Retrieval_Augmented_Calling.texttext-generation1K<n<10K0 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.