datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bangladesh-legal-qa-dataset
Bangladesh Legal QA Dataset: Bangla-English Law and Fine-Tuning
The Bangladesh Legal QA Dataset is a bilingual Bangla-English dataset for
Bangladesh law question answering, legal NLP, LLM fine-tuning, instruction
tuning, and retrieval-augmented generation (RAG). It provides 2,165
context-grounded legal QA records, direct-answer and IRAC chat-format training
data, and structured statutory text from six Bangladesh Acts and three
schedules.
This is the 2,165-record paper-aligned… See the full description on the dataset page: https://huggingface.co/datasets/momahadi/bangladesh-legal-qa-dataset.NaijaMed_QA_Dataset
Nigerian Healthcare Forum Q&A Datase
Dataset Summary
This dataset contains questions from Nigerians on a dedicated healthcare forum and responses provided exclusively by licensed and trained healthcare professionals. It reflects health concerns within the Nigerian context, incorporating English and local colloquialisms. All answers are reliable, as the platform restricted responses to verified healthcare professionals, ensuring the quality and credibility of the… See the full description on the dataset page: https://huggingface.co/datasets/Ayomidejoe/NaijaMed_QA_Dataset.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Medical-Health-QA-Articles-Dataset.Medical-Health-QA-Articles-Dataset
Medical Health Q&A & Articles Dataset — iCliniq, HealthTap & WebMD
A multi-source medical Q&A and health articles dataset combining doctor-answered questions and medically reviewed content from iCliniq, HealthTap, and WebMD. Built for LLM fine-tuning, medical chatbot training, clinical NLP research, and healthcare AI development.
Dataset Overview
Field
Details
Sources
iCliniq, HealthTap, WebMD
Total Records
1,000 (sample) — 50,000+ full dataset… See the full description on the dataset page: https://huggingface.co/datasets/david-sprague/Medical-Health-QA-Articles-Dataset.safety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.qa-dataset-llm-judge-flattened
Q&A Dataset - LLM-as-Judge Analyzed (Flattened)
Dataset Description
This dataset contains 5,008 high-quality question-answer pairs extracted from regulatory and policy documents, analyzed and quality-assessed using LLM-as-Judge methodology with parallel processing.
Key Features
Source: Official regulatory documents including policy directions, guidelines, and circulars
Quality Assessment: Each Q&A pair evaluated by LLM-as-Judge on multiple criteria
Answer… See the full description on the dataset page: https://huggingface.co/datasets/Magneto/qa-dataset-llm-judge-flattened.aee-qa-dataset
Jaliyanimantha/aee-qa-dataset
Structured question-answering dataset built from annual-report style source material.
Schema
id: Stable row id added during upload preparation.
context: Source passage used to answer the question.
has_graph: Whether the source page contains graph content.
page: Source page number when available.
question_no: Original question number within the source file.
question_type: Extracted label such as HARD FACT or STRATEGIC.
question: Cleaned… See the full description on the dataset page: https://huggingface.co/datasets/Jaliyanimantha/aee-qa-dataset.
