datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Customer_support_faqs_datasetDataset Name: Customer Support FAQs Dataset
Description:
This dataset contains a collection of 200 frequently asked questions (FAQs) and their corresponding answers, designed to assist in customer support scenarios. The questions cover a wide range of common customer inquiries related to account management, payment methods, order tracking, shipping, returns, and more. This dataset is intended for use in developing and training AI models for customer support chatbots, automated response systems… See the full description on the dataset page: https://huggingface.co/datasets/MakTek/Customer_support_faqs_dataset.faq-bacen
FaqBacenRetrieval
Retrieve the correct answer to a citizen question about Brazilian financial/banking regulation, from the Banco Central do Brasil (BACEN) public FAQ. 373 test questions over a pool of 1673 unique regulatory answers. Native PT-BR; financial/government domain.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark. Task type: Retrieval · Language: Brazilian Portuguese (mined from real-world sources) · Domains: Financial, Government, Written.… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/faq-bacen.faquad-ir
FaQuADIR
FaQuAD reformulated as PT-BR academic retrieval: given a question about Brazilian higher education, retrieve the source paragraph containing the answer. 900 questions over 244 unique paragraphs, drawn from 18 official documents of a federal-university CS program plus 21 Wikipedia articles about Brazil's higher-education system. Complements legal/web retrieval domains with academic/educational content.
Part of MTEB-BR — the native Brazilian-Portuguese MTEB sub-benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/MTEB-BR/faquad-ir.sap_faqMental_Health_FAQ
License & Attribution
MTEB-format derivative of tolu07/Mental_Health_FAQ. Licensed under MIT (same as source). Text encoding repaired with ftfy.
faq
QA4FAQ @ EVALITA 2016
Original dataset information available here
Data format
The data has been converted to be used as a questin answering task.
There are two splits, test-1 and test-2, each containing the same data processed in slightly different ways.
test-1
The data is in jsonl format, where each line is a json object with the following fields:
id: a unique identifier for the question
question: the question
A, B, C, D: the possible answers to the question… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/faq.bitcoin-wallet-recovery-faq
Bitcoin Wallet Recovery FAQ Dataset v1.0
A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery.
Dataset Summary
Total records: 500
Language: English
Answer length: 150–300 words per record
Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.georgian-faqFrequently asked questions (FAQs) and answers mined from Georgian websites via Common Crawl
hc3-faq-fr-gouv
License & Attribution
MTEB-format derivative of the faq_fr_gouv subset of almanach/hc3_french_ood (French government FAQ). Query = question; corpus = answer. Licensed under CC-BY-SA-4.0 (same as source).
faq
Council of AI — FAQ door
FAQ door. Living GET only. No 13-of-14. No 24/7. No “all 313 verify.”
Live: https://councilof.ai/faq
Council OS: https://councilof.ai/os
Council Space: https://councilof.ai/gspc-arena
Measurement, not certification. Empty slots are not for sale. No scores on this card.
Jail is a measured floor, not a 16th pane.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of… See the full description on the dataset page: https://huggingface.co/datasets/csoai/faq.faquad-nli
Dataset Card for FaQuAD-NLI
Dataset Summary
FaQuAD is a Portuguese reading comprehension dataset that follows the format of the Stanford Question Answering Dataset (SQuAD). It is a pioneer Portuguese reading comprehension dataset using the challenging format of SQuAD. The dataset aims to address the problem of abundant questions sent by academics whose answers are found in available institutional documents in the Brazilian higher education system. It consists of 900… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/faquad-nli.FDA_Pharmaceuticals_FAQ
FDA Pharmaceutical Q&A Dataset
Description
This dataset contains a collection of question-and-answer pairs related to pharmaceutical regulatory compliance provided by the Food and Drug Administration (FDA). It is designed to support research and development in the field of natural language processing, particularly for tasks involving information retrieval, question answering, and conversational agents within the pharmaceutical domain.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/Jaymax/FDA_Pharmaceuticals_FAQ.Ecommerce_FAQEcommerce FAQ Chatbot Dataset
Overview
The Ecommerce FAQ Chatbot Dataset is a valuable collection of questions and corresponding answers, meticulously curated for training and evaluating chatbot models in the context of an Ecommerce environment. This dataset is designed to assist developers, researchers, and data scientists in building effective chatbots that can handle customer inquiries related to an Ecommerce platform.
Contents
The dataset comprises a total of 79 question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/Andyrasika/Ecommerce_FAQ.swedish-construction-faq
Swedish Construction FAQ — Open Q&A Dataset
Open bilingual (Swedish/English) Q&A dataset for the Swedish construction
industry (byggbranschen). 503 Q&A pairs across 39 categories, every answer
grounded in Swedish primary law and authoritative guidance.
Maintained by Zaragoza AB, Helsingborg, Sweden.
DOI: 10.5281/zenodo.19630803 ·
Wikidata: Q139393633
Try it first
Live search demo: huggingface.co/spaces/DecDEPO/swedish-construction-faq-search
Colab quickstart: Open in… See the full description on the dataset page: https://huggingface.co/datasets/DecDEPO/swedish-construction-faq.synthetic-persian-chatbot-rag-faq-retrieval
Dataset Summary
Synthetic Persian Chatbot RAG FAQ Retrieval (SynPerChatbotRAGFAQRetrieval) is a Persian (Farsi) dataset built for the Retrieval task in Retrieval-Augmented Generation (RAG)-based chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini. The dataset is designed to evaluate how well models retrieve relevant FAQ entries based on a user's message and prior conversation context.
Language(s):… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-retrieval.E-Commerce_FAQsbanking-faq-hi-en-speech
Banking FAQ Hindi-English Speech Dataset
A multilingual banking FAQ dataset that combines an English FAQ source with Hindi translation and synthetic speech generation.
Overview
This project starts from a public English banking FAQ dataset and converts it into a Hindi-English speech dataset for research and experimentation in conversational AI and multilingual speech systems.
Source dataset
Kaggle:… See the full description on the dataset page: https://huggingface.co/datasets/ashirbadsahu/banking-faq-hi-en-speech.domeggook_faqsynthetic-persian-chatbot-rag-faq-pair-classification
Dataset Summary
Synthetic Persian Chatbot RAG FAQ Pair Classification (SynPerChatbotRAGFAQPC) is a Persian (Farsi) dataset for the Pair Classification task, specifically designed for evaluating Retrieval-Augmented Generation (RAG) chatbot systems. It is part of the FaMTEB (Farsi Massive Text Embedding Benchmark) and was synthetically generated using GPT-4o-mini.
The dataset measures a model’s ability to assess whether a given FAQ (question–answer pair) is relevant to a new user… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/synthetic-persian-chatbot-rag-faq-pair-classification.Mental_Health_FAQContent
Mental health includes our emotional, psychological, and social well-being. Mental health is integral to living a healthy, balanced life. It affects how we think, feel, and act. It also helps determine how we handle stress, relate to others, and make choices. Emotional and mental health is important because it’s a vital part of your life and impacts your thoughts, behaviors and emotions. Being healthy emotionally can promote productivity and effectiveness in activities like work… See the full description on the dataset page: https://huggingface.co/datasets/tolu07/Mental_Health_FAQ.c4-faqs
Dataset Card for [Dataset Name]
Dataset Summary
This dataset comprises of open-domain question-answer pairs obtained from extracting 150K FAQ URLs from C4 dataset. Please refer to the original paper and dataset card for more details.
You can load C4-FAQs as follows:
from datasets import load_dataset
c4_faqs_dataset = load_dataset("vishal-burman/c4-faqs")
Supported Tasks and Leaderboards
C4-FAQs is mainly intended for open-domain end-to-end question… See the full description on the dataset page: https://huggingface.co/datasets/vishal-burman/c4-faqs.faqih_sft_dataset
💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples)
مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026)
100% Curated & Unabridged MAFQA Multi-Hop Dataset
تم تنقيح وتدقيق البيانات بالكامل:
1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر.
2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية.
3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.clinicguide-faq-corpus
ClinicGuide FAQ corpus
Synthetic, generic clinic-style FAQs for a medical-tourism pre-consultation assistant. The corpus is not copied from a named hospital and is not a substitute for a real clinic's published policies.
Languages: English, Arabic (MSA), Persian.
Intended use
Retrieval-augmented answers for visa, stay, companion, hotel, airport transfer, starting-from costs, documents, booking, and “what to ask the doctor”
Safety / refusal examples: diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/clinicguide-faq-corpus.Future_Education_Nepal_MBBS_FAQ_Dataset
Future Education Nepal — MBBS FAQ Dataset (Nepali)
Overview
This dataset (future_education_nepal_mbbs_faq_nepali_sharegpt.jsonl) is a collection of 38 question-answering conversation pairs in Nepali, covering frequently asked questions about studying MBBS (medicine) in Nepal — primarily aimed at Indian students considering Nepal as a study destination. Each record is a single-turn human↔gpt exchange in ShareGPT-style format: a Nepali-language question about MBBS… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Future_Education_Nepal_MBBS_FAQ_Dataset.educational_vedantu_FAQs_Nepali_sft_dataset
Vedantu CUET FAQs — Nepali SFT Dataset
A single-turn instruction-following (Q&A) dataset in Nepali, built from Vedantu's CUET (Common University Entrance Test) 2026 FAQ page. The dataset is formatted for supervised fine-tuning (SFT) in a conversations-style chat schema.
Summary
Records
49
Language
Nepali (ne / npi)
Script
Devanagari (Deva)
Format
JSONL, conversations (human/human turns)
Domain
Education — CUET exam FAQs
Task type… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/educational_vedantu_FAQs_Nepali_sft_dataset.FAQ_BACENThis dataset was used in the article: https://arxiv.org/abs/2311.11331
Pharmacy_Licensing_Exam_FAQ
💊 Pharmacy Licensing Exam FAQ (Nepali) — Dataset README
A 100% Nepali-language synthetic reasoning dataset built from a single row-pair of official Nepal Department of Health Services exam-result data — turned into 100 question-answer pairs covering counts, percentages, ratios, statistics, and "what-if" arithmetic.
🔖 TL;DR (At a Glance)
What
Answer
Total records
100
File size
~198 KB
Language / script
Nepali (ne / npi) — Devanagari
Format… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Pharmacy_Licensing_Exam_FAQ.study-in-india-faq
Study-in-India FAQ Dataset
This dataset, study-in-india-faq, is designed for fine-tuning language models to answer frequently asked questions about studying in India. It includes 200,000 pairs of questions and answers covering topics such as admissions, scholarships, accommodation, cultural adjustments, and visa requirements.
Dataset Summary
The study-in-india-faq dataset is a comprehensive resource for students seeking information about studying in India, whether they… See the full description on the dataset page: https://huggingface.co/datasets/millat/study-in-india-faq.luciolescribe-transcription-faq
🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0
📋 Description
Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception.
🎯 Caractéristiques clés
📊 Taille: 182+ paires question-réponse (expansion continue)
🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles
🌍 Langue: Français (France)
📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.ShareHub_FAQ_Nepali_Dataset
ShareHub FAQ Nepali Dataset
A Nepali-language, instruction-following (question–answer) dataset of 2,448 real frequently-asked-questions about securities listed on the Nepal Stock Exchange (NEPSE) — covering common stocks, debentures, and mutual funds — sourced from ShareHub. Every record is a single-turn human → gpt conversation in Devanagari script, and every record carries rich provenance and generation metadata.
File analyzed: sharehub_faq.jsonl
1. Quick Facts… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/ShareHub_FAQ_Nepali_Dataset.
