datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.cpt_instruction_datasets
Instruction datasets
Collection of synthetic instruction datasets used during the continued pretraining of Model-small-instr-1, Model-small-instr-2 and Model-small-instr-3. You can currently find these models under: Llama-3.1-Carballo-Instr1 and Llama-3.1-Carballo-Instr3.
Dataset creation
Datasets were created using two different techniques:
Adapting already existing datasets or corpora by modifying their format to make them suitable for including instructions during… See the full description on the dataset page: https://huggingface.co/datasets/proxectonos/cpt_instruction_datasets.to-tool-call-datasets
🛠️ To-Tool-Call Datasets
A unified Qwen3-style tool-call corpus for SFT, GRPO, and agent training
To-Tool-Call Datasets is a curated mirror of public tool-call and function-calling corpora, re-serialized into one training-ready messages JSONL convention.
Quick Start ·
At a Glance ·
Format ·
Sources ·
Training Notes
[!IMPORTANT]
This repository is a format-harmonization layer, not a new claim of ownership over the… See the full description on the dataset page: https://huggingface.co/datasets/zhangdw/to-tool-call-datasets.airecon-datasets
AIRecon Security Datasets
Curated security knowledge datasets for AIRecon — an AI-powered security reconnaissance tool that runs 100% locally with Ollama.
These datasets augment the LLM agent's knowledge for penetration testing, reconnaissance, vulnerability analysis, and security research workflows.
How AIRecon Uses These Datasets
Dataset → Phase Mapping
Dataset
Primary Phase
What It Provides
recon-playbook
RECON
Agent methodology, phase tactics… See the full description on the dataset page: https://huggingface.co/datasets/pikpikcu/airecon-datasets.Athar-Datasets
🕌 Athar Islamic QA Datasets
18.7M passages from classical Islamic books spanning 1,400 years of scholarship
A comprehensive collection of Islamic texts covering Quran, Hadith, Fiqh, Tafsir, Aqeedah, Seerah, and more — sourced from the Shamela library and enriched with scholarly metadata for RAG-based Islamic QA systems.
Based on the Fanar-Sadiq Architecture for grounded, citation-backed Islamic question answering.
📊 Dataset Summary
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/Kandil7/Athar-Datasets.RusLang-edu-1000
RusLang-Edu-1000 — an educational Russian-language QA dataset
RusLang-Edu-1000 is an expert-curated dataset of 1,000 instruction-format records ("question — detailed educational answer") covering the Russian language and linguistics: from phonetics and orthography to dialectology and theoretical linguistics. Every record contains a detailed answer (on average ≈1,100 characters), a short reference answer, a concise statement of the rule, and rich annotation (subject area, task… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/RusLang-edu-1000.headroom-datasets
Headroom Pilot — prove context-compression value in minutes
A curated set of Anthropic Messages API payloads built to demonstrate what
Headroom — the context-optimization
layer for LLM apps — is good at: shrinking the large tool outputs that dominate
agentic and data-processing prompts, without changing the answer.
Each row is a complete /v1/messages request (system + tools + messages) whose
final turn is a bloated tool_result — exactly the content Headroom
compresses. Every row… See the full description on the dataset page: https://huggingface.co/datasets/chopratejas/headroom-datasets.ngari-datasets
NGARi Training Datasets
NGARi-authored training datasets (Apache 2.0), generated on sovereign edge hardware with zero cloud dependency. These power the NGARi edge models: ngari-ft-distilled and ngari-tool.
The NGARi data engine
Data quality is the bottleneck for capable small models. NGARi uses a large teacher model (27B-class, e.g. qwen3.8-27B) as an automated data engine — generating diverse edge-cases, complex instructions, and niche domain knowledge — then… See the full description on the dataset page: https://huggingface.co/datasets/NGARiAI/ngari-datasets.aime-2026-fable-5-answers
Homepage and repository
Homepage: https://matharena.ai/
Repository: https://github.com/eth-sri/matharena
Dataset Summary
aime-2026-formatted-fable — это обработанный и структурированный датасет на основе задач AIME 2026 из бенчмарка MathArena. Датасет сохранён в формате JSONL и помимо условий задач с финальными ответами содержит сгенерированные цепочки рассуждений (think) с ограничением объёма до 2048 токенов на пример.
Data Fields
Каждая запись в… See the full description on the dataset page: https://huggingface.co/datasets/DatasetsEval/aime-2026-fable-5-answers.Masadir-Maliki-Dataset
📖 Dataset Summary | ملخص القاعدة
This dataset is a bibliographic Question–Answer (QA) corpus derived from the reference workSources of Maliki Jurisprudence: Usūlan wa Furūʿan by Shaykh Abū ʿĀṣim Bashīr Ḍayf (d. 1429 AH / 2008 CE).
The dataset provides structured access to:
Core Maliki fiqh sources
Authorial lineages
Methodological classifications
Printed and manuscript works across the Islamic East and West
It is intended for Islamic studies research, bibliographic analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/Masadir-Maliki-Dataset.maliki-terminology
اصطلاحات أعلام المالكية وألقابهم
قاعدة بيانات متخصصة في رموز وألقاب المدرسة المالكية
📖 وصف البيانات
تحتوي هذه القاعدة على استخراج دقيق للألقاب والرموز العلمية المستخدمة في كتب الفقه المالكي (مثل: الشيخ، الأخوان، الصادقان، المحمدون...). تم تحويل المادة العلمية إلى صيغة سؤال وجواب (QA) لتسهيل تدريب نماذج الذكاء الاصطناعي على فهم السياق التاريخي والعلمي للمذهب.
📂 محتوى الملف
عدد القيود: 29 اصطلاحاً رئيسياً.
الصيغة: JSONL.
الحقول: - question: السؤال… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/maliki-terminology.Five_Phases_Mindset_datasetsWelcome to our Traditional Chinese Medicine (TCM) Consultation Dataset! This dataset contains approximately one hundred thousand TCM consultation dialogue records, aiming to provide a rich resource for research and development in the field of TCM. These dialogue data cover various TCM diseases, diagnoses, and treatment methods, serving as an important reference for TCM research and clinical practice.
The dataset was created using a method that combines manual annotation with extraction from… See the full description on the dataset page: https://huggingface.co/datasets/cookey39/Five_Phases_Mindset_datasets.fiqh-maliki-talqinAl-Talqīn: A Digitized Dataset of Mālikī Jurisprudence
قاعدة بيانات كتاب «التلقين في الفقه المالكي» – رقمنة تراث فقهي
About the Book | عن الكتاب
🇬🇧 English
Al-Talqīn (التلقين) is one of the most authoritative concise manuals in Mālikī jurisprudence. It was authored by Al-Qāḍī Abū Muḥammad ʿAbd al-Wahhāb ibn ʿAlī al-Baghdādī al-Mālikī (d. 422 AH).
The book is renowned for its precision, clarity, and systematic organization, making it a foundational reference for students and scholars of the… See the full description on the dataset page: https://huggingface.co/datasets/islamic-datasets/fiqh-maliki-talqin.joy_common_sft_datasetsFusion_Ita_Datasets
📚 Mattimax/Fusion_Ita_Datasets
📌 Descrizione
Mattimax/Fusion_Ita_Datasets è un dataset in italiano ottenuto dalla fusione, pulizia e normalizzazione di sei dataset pubblici di conversazioni e istruzioni, pensato per l’addestramento di modelli di linguaggio in italiano.
Include dati di alta qualità da QA, conversazioni multi-turno, domande in stile Quora e StackOverflow, filtrati per lingua e deduplicati per garantire coerenza e ridurre il rumore.
🛠… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/Fusion_Ita_Datasets.alignment_datasets
🧠 Persian Cultural Alignment Dataset for LLMs
This repository contains a high-quality, Alignment dataset for cultural alignment of large language models (LLMs) in the Persian language. The dataset is curated using hybrid strategies that incorporate culturally grounded generation, multi-turn dialogues, translation, and augmentation methods, making it suitable for SFT, DPO, RLHF, and alignment evaluation.
📚 Dataset Overview
Domain
Methods Used
Culinary… See the full description on the dataset page: https://huggingface.co/datasets/MatinaAI/alignment_datasets.Fusion_Ita_Datasets_2
📚 Mattimax/Fusion_Ita_Datasets_2
📌 Descrizione
Mattimax/Fusion_Ita_Datasets_v2 è un dataset in italiano creato dalla fusione e normalizzazione di diversi dataset pubblici di conversazioni, istruzioni e QA.
Include dati di alta qualità in lingua italiana, filtrati per rimuovere valori nulli e duplicati, pronti per l’addestramento di modelli di linguaggio per completamento di testi, domande/risposte e dialoghi multi-turno.
🛠 Origine dei dati
I dataset… See the full description on the dataset page: https://huggingface.co/datasets/Mattimax/Fusion_Ita_Datasets_2.datasets_ub
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ratno/datasets_ub.
