datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
faq
QA4FAQ @ EVALITA 2016
Original dataset information available here
Data format
The data has been converted to be used as a questin answering task.
There are two splits, test-1 and test-2, each containing the same data processed in slightly different ways.
test-1
The data is in jsonl format, where each line is a json object with the following fields:
id: a unique identifier for the question
question: the question
A, B, C, D: the possible answers to the question… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/faq.bitcoin-wallet-recovery-faq
Bitcoin Wallet Recovery FAQ Dataset v1.0
A high-quality Question & Answer dataset focused exclusively on Bitcoin wallet recovery and self-custody best practices. It is designed for training, fine-tuning, and evaluating LLMs and retrieval-augmented generation (RAG) systems in the domain of bitcoin security, seed backup, device loss, and fund recovery.
Dataset Summary
Total records: 500
Language: English
Answer length: 150–300 words per record
Categories: 39… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-recovery-faq.faquad-nli
Dataset Card for FaQuAD-NLI
Dataset Summary
FaQuAD is a Portuguese reading comprehension dataset that follows the format of the Stanford Question Answering Dataset (SQuAD). It is a pioneer Portuguese reading comprehension dataset using the challenging format of SQuAD. The dataset aims to address the problem of abundant questions sent by academics whose answers are found in available institutional documents in the Brazilian higher education system. It consists of 900… See the full description on the dataset page: https://huggingface.co/datasets/ruanchaves/faquad-nli.faquadAcademic secretaries and faculty members of higher education institutions face a common problem:
the abundance of questions sent by academics
whose answers are found in available institutional documents.
The official documents produced by Brazilian public universities are vast and disperse,
which discourage students to further search for answers in such sources.
In order to lessen this problem, we present FaQuAD:
a novel machine reading comprehension dataset
in the domain of Brazilian higher education institutions.
FaQuAD follows the format of SQuAD (Stanford Question Answering Dataset) [Rajpurkar et al. 2016].
It comprises 900 questions about 249 reading passages (paragraphs),
which were taken from 18 official documents of a computer science college
from a Brazilian federal university
and 21 Wikipedia articles related to Brazilian higher education system.
As far as we know, this is the first Portuguese reading comprehension dataset in this format.swedish-construction-faq
Swedish Construction FAQ — Open Q&A Dataset
Open bilingual (Swedish/English) Q&A dataset for the Swedish construction
industry (byggbranschen). 503 Q&A pairs across 39 categories, every answer
grounded in Swedish primary law and authoritative guidance.
Maintained by Zaragoza AB, Helsingborg, Sweden.
DOI: 10.5281/zenodo.19630803 ·
Wikidata: Q139393633
Try it first
Live search demo: huggingface.co/spaces/DecDEPO/swedish-construction-faq-search
Colab quickstart: Open in… See the full description on the dataset page: https://huggingface.co/datasets/DecDEPO/swedish-construction-faq.c4-faqs
Dataset Card for [Dataset Name]
Dataset Summary
This dataset comprises of open-domain question-answer pairs obtained from extracting 150K FAQ URLs from C4 dataset. Please refer to the original paper and dataset card for more details.
You can load C4-FAQs as follows:
from datasets import load_dataset
c4_faqs_dataset = load_dataset("vishal-burman/c4-faqs")
Supported Tasks and Leaderboards
C4-FAQs is mainly intended for open-domain end-to-end question… See the full description on the dataset page: https://huggingface.co/datasets/vishal-burman/c4-faqs.clinicguide-faq-corpus
ClinicGuide FAQ corpus
Synthetic, generic clinic-style FAQs for a medical-tourism pre-consultation assistant. The corpus is not copied from a named hospital and is not a substitute for a real clinic's published policies.
Languages: English, Arabic (MSA), Persian.
Intended use
Retrieval-augmented answers for visa, stay, companion, hotel, airport transfer, starting-from costs, documents, booking, and “what to ask the doctor”
Safety / refusal examples: diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/clinicguide-faq-corpus.luciolescribe-transcription-faq
🎙️ LucioleScribe Transcription FAQ - Dataset Production v2.0
📋 Description
Dataset enrichi de questions-réponses FAQ sur la transcription IA 100% locale avec LucioleScribe, première plateforme française de transcription conforme RGPD par conception.
🎯 Caractéristiques clés
📊 Taille: 182+ paires question-réponse (expansion continue)
🏷️ Métadonnées: Enrichi avec catégories, intents, buyer stages, styles
🌍 Langue: Français (France)
📑 Format: JSON Lines… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/luciolescribe-transcription-faq.FAQ_BACENThis dataset was used in the article: https://arxiv.org/abs/2311.11331
study-in-india-faq
Study-in-India FAQ Dataset
This dataset, study-in-india-faq, is designed for fine-tuning language models to answer frequently asked questions about studying in India. It includes 200,000 pairs of questions and answers covering topics such as admissions, scholarships, accommodation, cultural adjustments, and visa requirements.
Dataset Summary
The study-in-india-faq dataset is a comprehensive resource for students seeking information about studying in India, whether they… See the full description on the dataset page: https://huggingface.co/datasets/millat/study-in-india-faq.faqih_sft_dataset
💎 MAFQA: Perfected Multi-Hop Arabic Fatwa QA Dataset (388 Samples)
مجموعة بيانات الاستدلال الفقهي المركب ومتعدد الخطوات (جامعة الملك سعود / MDPI 2026)
100% Curated & Unabridged MAFQA Multi-Hop Dataset
تم تنقيح وتدقيق البيانات بالكامل:
1. إزالة جميع التقطيعات النصية وإيراد النصوص والأدلة كاملة دون بتر.
2. تفعيل خطوة التركيب والترجيح النهائي (الخطوة 4) بربط استدلالي حقيقي بين المسائل الفرعية.
3. تصحيح الأخطاء المطبعية في دلالات الحل والحرمة.… See the full description on the dataset page: https://huggingface.co/datasets/hozifa1/faqih_sft_dataset.ameba_faq_search
AMEBA Blog FAQ Search Dataset
This data was obtained by crawling this website.
The FAQ Data was processed to remove HTML tags and other formatting after crawling, and entries containing excessively long content were excluded.
The Query Data was generated using a Large Language Model (LLM). Please refer to the following blog for information about the generation process.
https://www.ai-shift.co.jp/techblog/3710
https://www.ai-shift.co.jp/techblog/3761
Column description… See the full description on the dataset page: https://huggingface.co/datasets/ai-shift/ameba_faq_search.romanian-legal-faq-2026
Dataset: Romanian Legal FAQ 2026 (Coltuc Legal Knowledge Base)
Descriere
Set de date structurat în limba română conținând instrucțiuni, întrebări frecvente și soluții procedurale din dreptul civil, drept bancar (clauze abuzive, executări silite), dreptul muncii și dreptul pensiilor.
Dataset-ul este optimizat pentru fine-tuning LLM, sisteme RAG (Retrieval-Augmented Generation) și modele de asistență juridică automată.
Autor și Proprietate Intelectuală… See the full description on the dataset page: https://huggingface.co/datasets/Coltuc2026/romanian-legal-faq-2026.faqs-rrhh
FAQ de RRHH, registro horario y canal de denuncias (España)
Conjunto de preguntas frecuentes con respuesta sobre software de RRHH, registro horario (RD-ley 8/2019) y canal de denuncias (Ley 2/2023) para pymes españolas, publicado por Nucleo360. Las 8 primeras respuestas coinciden literalmente con las publicadas en la web oficial; el resto amplía la FAQ del registro horario obligatorio en 2026.
Contenido
faqs.csv — una pregunta por fila, con las columnas:… See the full description on the dataset page: https://huggingface.co/datasets/Nucleo360/faqs-rrhh.wiki-faq-10kfaquad-nli-parquet
Dataset Card for FaQuAD-NLI
THIS IS A TEMPORARY COPY OF THE ORIGINAL ruanchaves/faquad-nli.
WHY? As of datasets==4.0, loading scripts and trust_remote_code are no longer supported.
This breaks things, like the lm-evaluation-harness-pt, which people who work with Portuguese LLMs need for running evals.
As soon as ruanchaves updates his version, I'll delete this copy.
Dataset Summary
FaQuAD is a Portuguese reading comprehension dataset that follows the format of the… See the full description on the dataset page: https://huggingface.co/datasets/nicholasKluge/faquad-nli-parquet.Legal-FAQ
Legal FAQ Dataset
Dataset Description
The Legal FAQ Dataset is a collection of frequently asked legal questions along with their respective answers. This dataset is particularly useful for building question-answering systems, legal chatbots, or other applications in the domain of law and justice.
Dataset Details
Dataset ID: gurumurthy3/Legal-FAQ
Total Entries: 2091
Columns:
Question: A legal-related question.
Answer: The corresponding legal response.… See the full description on the dataset page: https://huggingface.co/datasets/gurumurthy3/Legal-FAQ.MonEspaceSante-FAQ-QA
MonEspaceSanté FAQ — paires Q/R (réelles + synthétiques)
Jeu de paires question / réponse en français ayant servi à fine-tuner l'assistant
MonEspaceSanté-FAQ-Mistral-Small-24B-GGUF,
spécialisé sur la FAQ du service public Mon espace santé.
Le jeu mélange les questions/réponses officielles de la FAQ (vérité terrain) et une
augmentation synthétique ancrée : des reformulations variées générées par un modèle
enseignant, dont chaque réponse est strictement justifiée par le texte… See the full description on the dataset page: https://huggingface.co/datasets/fenyo/MonEspaceSante-FAQ-QA.ReactJS_FAQ_Datasetkosodate-faq-pairs-ja
子育てAIチャットボット 汎用FAQ (anchor/positive pairs, ja)
Dataset Summary
本データセットは、「子育てオープンデータ協議会」が 2020 年に公開した 「子育てAIチャットボットの利活用促進に向けた検討 2020年報告書」附属の FAQ データセット を加工し、Sentence Transformers 互換の (anchor, positive) positive pair 形式 に変換したものです。
anchor: ユーザー側のサンプル問合せ文(自然な表現での質問)
positive: それに対応する FAQ 項目 / サンプル応答文(正解ドキュメント)
日本語テキスト埋め込みモデル(例: 日本語 BERT/SBERT, SPLADE, ColBERT, ELSER 類似の日本語スパース検索モデルなど)の fine-tune / 評価 用途や、自治体 FAQ を対象とした RAG システムの Retrieval コーパスとしてそのまま利用できます。… See the full description on the dataset page: https://huggingface.co/datasets/mahiyama/kosodate-faq-pairs-ja.FAQ_NelsMarketplaceThis dataset was created to test two different things:
First, check LLM's capabilities of augmenting data in a coherent way.
Second, create a dataset to finetune LLMs for the QA task.
The dataset contains the frequently asked questions and their answers of a made-up online fashion marketplace called: Nels Marketplace.
FAQ_BACENThis dataset was the used in the paper https://arxiv.org/abs/2311.11331
license: apache-2.0
mortgage-faq-2026
Mortgage FAQ 2026
203 expert-verified mortgage FAQs across 13 categories.
Details
Records: 203
Format: JSONL
License: CC-BY-4.0
Last Updated: 2026-07-26 (data-integrity correction)
Publisher: Good News Lending
Data-Integrity Correction (2026-07-26)
Prior versions of this dataset carried a fabricated NMLS number (#1573055) attributed to Wendy Thompson across 9 divorce and reverse-mortgage FAQ entries. The canonical NMLS for Wendy Thompson is… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/mortgage-faq-2026.HCMUT_FAQheath_disease_faqfaqFAQ (Frequently Asked Questions) - это датасет с вопросами, которые чаще всего задают пользователи, и с автоматизированным анализом этих вопросов, позволяющим нейронной сети выбирать из нескольких вероятных ответов. Важно иметь в виду, что эти ответы могут быть неточными или размытыми, так как они являются статистически вероятными ответами на заданный вопрос.
ecommerce-faq-llama2-QAarchitecture_faqsJapanese construction themes FAQs scraped from https://www.city.yokohama.lg.jp/business/bunyabetsu/kenchiku/annai/faq/qa.html.
Downloaded using the following code:
import requests
from lxml import html
import pandas as pd
from datasets import Dataset
hrefs = [
"/business/bunyabetsu/kenchiku/annai/faq/ji-annnai.html",
"/business/bunyabetsu/kenchiku/tetsuduki/kakunin/qa-kakunin.html",
"/business/bunyabetsu/kenchiku/tetsuduki/teikihoukoku/seido/01.html"… See the full description on the dataset page: https://huggingface.co/datasets/lightblue/architecture_faqs.ecommerce-faq-llama2-chatlegal-faq-paraphrased
Legal FAQ Query Reformulation Benchmark
Dataset Summary
This dataset is derived from the Legal-FAQ dataset and is designed for evaluating the robustness of retrieval and embedding models under query reformulation.
Each original FAQ question has been expanded into multiple query styles:
Formal
Casual
Search
Hard Reformulation
The objective is to measure how well retrieval systems can retrieve the correct answer despite substantial changes in wording.… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/legal-faq-paraphrased.
