CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face02proxectonos /cpt_instruction_datasets Instruction datasets Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3. Dataset creation Datasets were created using two different techniques: Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.tabulartext-generation100K<n<1M0 likes344 downloads5mo agoHugging Face03zhangdw /to-tool-call-datasets 🛠️ To-Tool-Call Datasets A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training &nbsp;&nbsp;&nbsp;&nbsp; To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention. Quick Start · At a Glance · Format · Sources · Training Notes [!IMPORTANT] This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.texttext-generation1K<n<10K3 likes304 downloads4mo agoHugging Face04pikpikcu /airecon-datasets AIRecon Security Datasets Curated security knowledge datasets for AIRecon — an AI-powered security reconnaissance tool that runs 100% locally with Ollama. These datasets augment the LLM agent's knowledge for penetration testing, reconnaissance, vulnerability analysis, and security research workflows. How AIRecon Uses These Datasets Dataset → Phase Mapping Dataset Primary Phase What It Provides recon-playbook RECON Agent methodology, phase tactics… See the full description on the dataset page: https://huggingface.co/datasets/pikpikcu/airecon-datasets.texttext-generation10K<n<100K1 likes171 downloads5mo agoHugging Face05Kandil7 /Athar-Datasets 🕌 Athar Islamic QA Datasets 18.7M passages from classical Islamic books spanning 1,400 years of scholarship A comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, and more — sourced from the Shamela library and enriched with scholarly metadata for RAG-based Islamic QA systems. Based on the Fanar-Sadiq Architecture for grounded, citation-backed Islamic question answering. 📊 Dataset Summary Metric Value… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Datasets.tabularquestion-answering10M<n<100M8 likes152 downloads5mo agoHugging Face06DatasetsEval /RusLang-edu-1000 RusLang-Edu-1000 — an educational Russian-language QA dataset RusLang-Edu-1000 is an expert-curated dataset of 1,000 instruction-format records ("question — detailed educational answer") covering the Russian language and linguistics: from phonetics and orthography to dialectology and theoretical linguistics. Every record contains a detailed answer (on average ≈1,100 characters), a short reference answer, a concise statement of the rule, and rich annotation (subject area, task… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/RusLang-edu-1000.textquestion-answering1K<n<10K2 likes107 downloads1mo agoHugging Face07chopratejas /headroom-datasets Headroom Pilot — prove context-compression value in minutes A curated set of Anthropic Messages API payloads built to demonstrate what Headroom — the context-optimization layer for LLM apps — is good at: shrinking the large tool outputs that dominate agentic and data-processing prompts, without changing the answer. Each row is a complete /v1/messages request (system + tools + messages) whose final turn is a bloated tool_result — exactly the content Headroom compresses. Every row… See the full description on the dataset page: https://huggingface.co/datasets/chopratejas/headroom-datasets.tabulartext-generationn<1K0 likes95 downloads3mo agoHugging Face08NGARiAI /ngari-datasets NGARi Training Datasets NGARi-authored training datasets (Apache 2.0), generated on sovereign edge hardware with zero cloud dependency. These power the NGARi edge models: ngari-ft-distilled and ngari-tool. The NGARi data engine Data quality is the bottleneck for capable small models. NGARi uses a large teacher model (27B-class, e.g. qwen3.8-27B) as an automated data engine — generating diverse edge-cases, complex instructions, and niche domain knowledge — then… See the full description on the dataset page: https://huggingface.co/datasets/NGARiAI/ngari-datasets.texttext-generation1K<n<10K0 likes61 downloads20d agoHugging Face09DatasetsEval /aime-2026-fable-5-answers Homepage and repository Homepage: https://matharena.ai/ Repository: https://github.com/eth-sri/matharena Dataset Summary aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример. Data Fields Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.textquestion-answeringn<1K0 likes54 downloads2mo agoHugging Face10islamic-datasets /Masadir-Maliki-Dataset 📖 Dataset Summary | ملخص القاعدة This dataset is a bibliographic Question–Answer (QA) corpus derived from the reference workSources of Maliki Jurisprudence: Usūlan wa Furūʿan by Shaykh Abū ʿĀṣim Bashīr Ḍayf (d. 1429 AH / 2008 CE). The dataset provides structured access to: Core Maliki fiqh sources Authorial lineages Methodological classifications Printed and manuscript works across the Islamic East and West It is intended for Islamic studies research, bibliographic analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Masadir-Maliki-Dataset.textquestion-answering1K<n<10K0 likes38 downloads9mo agoHugging Face11islamic-datasets /maliki-terminology اصطلاحات أعلام المالكية وألقابهم قاعدة بيانات متخصصة في رموز وألقاب المدرسة المالكية 📖 وصف البيانات تحتوي هذه القاعدة على استخراج دقيق للألقاب والرموز العلمية المستخدمة في كتب الفقه المالكي (مثل: الشيخ، الأخوان، الصادقان، المحمدون...). تم تحويل المادة العلمية إلى صيغة سؤال وجواب (QA) لتسهيل تدريب نماذج الذكاء الاصطناعي على فهم السياق التاريخي والعلمي للمذهب. 📂 محتوى الملف عدد القيود: 29 اصطلاحاً رئيسياً. الصيغة: JSONL. الحقول: - question: السؤال… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/maliki-terminology.textquestion-answeringn<1K0 likes32 downloads9mo agoHugging Face12cookey39 /Five_Phases_Mindset_datasetsWelcome to our Traditional Chinese Medicine (TCM) Consultation Dataset! This dataset contains approximately one hundred thousand TCM consultation dialogue records, aiming to provide a rich resource for research and development in the field of TCM. These dialogue data cover various TCM diseases, diagnoses, and treatment methods, serving as an important reference for TCM research and clinical practice. The dataset was created using a method that combines manual annotation with extraction from… See the full description on the dataset page: https://huggingface.co/datasets/cookey39/Five_Phases_Mindset_datasets.textquestion-answering100K<n<1M2 likes31 downloads2y agoHugging Face13islamic-datasets /fiqh-maliki-talqinAl-Talqīn: A Digitized Dataset of Mālikī Jurisprudence قاعدة بيانات كتاب «التلقين في الفقه المالكي» – رقمنة تراث فقهي About the Book | عن الكتاب 🇬🇧 English Al-Talqīn (التلقين) is one of the most authoritative concise manuals in Mālikī jurisprudence. It was authored by Al-Qāḍī Abū Muḥammad ʿAbd al-Wahhāb ibn ʿAlī al-Baghdādī al-Mālikī (d. 422 AH). The book is renowned for its precision, clarity, and systematic organization, making it a foundational reference for students and scholars of the… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/fiqh-maliki-talqin.textquestion-answering1K<n<10K3 likes31 downloads9mo agoHugging Face14jdh-algo /joy_common_sft_datasetstextquestion-answering10K<n<100K1 likes17 downloads2y agoHugging Face15Mattimax /Fusion_Ita_Datasets 📚 Mattimax/Fusion_Ita_Datasets 📌 Descrizione Mattimax/Fusion_Ita_Datasets è un dataset in italiano ottenuto dalla fusione, pulizia e normalizzazione di sei dataset pubblici di conversazioni e istruzioni, pensato per l’addestramento di modelli di linguaggio in italiano. Include dati di alta qualità da QA, conversazioni multi-turno, domande in stile Quora e StackOverflow, filtrati per lingua e deduplicati per garantire coerenza e ridurre il rumore. 🛠… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/Fusion_Ita_Datasets.texttext-generation100K<n<1M0 likes14 downloads1y agoHugging Face16MatinaAI /alignment_datasetsgated 🧠 Persian Cultural Alignment Dataset for LLMs This repository contains a high-quality, Alignment dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation. 📚 Dataset Overview Domain Methods Used Culinary… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/alignment_datasets.textquestion-answering10K<n<100K0 likes9 downloads1y agoHugging Face17Mattimax /Fusion_Ita_Datasets_2 📚 Mattimax/Fusion_Ita_Datasets_2 📌 Descrizione Mattimax/Fusion_Ita_Datasets_v2 è un dataset in italiano creato dalla fusione e normalizzazione di diversi dataset pubblici di conversazioni, istruzioni e QA. Include dati di alta qualità in lingua italiana, filtrati per rimuovere valori nulli e duplicati, pronti per l’addestramento di modelli di linguaggio per completamento di testi, domande/risposte e dialoghi multi-turno. 🛠 Origine dei dati I dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/Fusion_Ita_Datasets_2.texttext-generation100K<n<1M0 likes9 downloads1y agoHugging Face18ratno /datasets_ub Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ratno/datasets_ub.textquestion-answering1K<n<10K0 likes5 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.