datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-question-answering-datasetsNLU-Question-Answering
SEA Question Answering
SEA Question Answering evaluates a model's ability to predict a contiguous span of characters that answers the question about a given passage. It is sampled from TyDi QA-GoldP for Indonesian, IndicQA for Tamil, and XQuaD for Thai and Vietnamese.
Supported Tasks and Leaderboards
SEA Question Answering is designed for evaluating chat or instruction-tuned large language models (LLMs). It is part of the SEA-HELM leaderboard from AI Singapore.… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/NLU-Question-Answering.Turkish-medical-visual-question-answering-LLaVa-dataset
Türkçe Radyoloji Görüntüleme Veri Seti - data_RAD
data_RAD veri seti, radyoloji görüntüleri üzerinde görsel soru-cevaplama (VQA) araştırmaları yapmak amacıyla Türkçeye çevrilmiş ve LLaVa mimarisiyle uyumlu hale getirilmiştir. Bu veri seti, tıbbi görüntü analizi ve yapay zeka destekli radyoloji uygulamalarını geliştirmek için kullanılabilir.
Veri Seti İçeriği
Toplam Görüntü Sayısı: 316
Veri Yapısı: DatasetDict({ train: Dataset({ features: ['image'], num_rows: 316 }) })
Özellikler:… See the full description on the dataset page: https://huggingface.co/datasets/nezahatkorkmaz/Turkish-medical-visual-question-answering-LLaVa-dataset.Question-AnsweringThis is the question answering datasets collected by TextBox, including:
SQuAD (squad)
CoQA (coqa)
Natural Questions (nq)
TriviaQA (tqa)
WebQuestions (webq)
NarrativeQA (nqa)
MS MARCO (marco)
NewsQA (newsqa)
HotpotQA (hotpotqa)
MSQG (msqg)
QuAC (quac).
The detail and leaderboard of each dataset can be found in TextBox page.
Chinese_Question_Answering_Datasetmedical-question-answering-datasetsmedical-question-answering-datasetschaii-hindi-and-tamil-question-answeringQuestion-Answering-Generation-Choices
The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets,
having undergone preprocessing.
QuestionAnswering
Serbian Question-Answering Datasets
This repository provides multiple QA datasets in Serbian, suitable for training LLMs to answer questions, perform tasks, or function as chatbots.
Datasets Overview
SQuAD-sr-md – Manually corrected subset of SQuAD-sr (~7k corrected samples), for higher reliability and accuracy.
SerbianQA-Gen – Synthetic QA dataset (~74k samples) generated from encyclopedia articles, Wikipedia pages, and scientific abstracts. Organized into four… See the full description on the dataset page: https://huggingface.co/datasets/te-sla/QuestionAnswering.bodo-legal-question-answering-ai4bharat
Bodo Legal Question Answering Dataset
Overview
This dataset is a Bodo-language legal Question Answering (QA) resource
created for research in low-resource Natural Language Processing (NLP)
and legal language processing.
The supplied source files contain legal judgment contexts together with
multiple questions and answers. For Hugging Face compatibility and
question-answering model training, each question-answer pair has been
flattened into a separate JSONL example… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-ai4bharat.civil-human-rights-question-answering
Dataset Card for rag-prompt
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/rag-prompt/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/civil-human-rights-question-answering.bodo-legal-question-answering-iiith
Bodo Legal Question Answering Dataset — IIITH Translation
Overview
A Bodo-language legal Question Answering (QA) resource derived from
English legal judgments. Each example contains a judgment context, a
question, and its corresponding answer.
Data Provenance
Original Legal Source
The underlying English legal judgments were extracted from the publicly
accessible Gauhati High Court judgment repository:… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-iiith.medical-question-answering-datasetsQuestion-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.questionanswering-datasetQuestion-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.siddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.medical-question-answering-datasetsmedical-question-answering-datasetsMulti-hop_Question_Answering_Hard
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Tianda7721/Multi-hop_Question_Answering_Hard.github_fetch_huggingface_pdf-tools_terminal_2096-docaudit-7c91-financial-question-answering
Financial Question Answering
Dataset Summary
Question-answer pairs extracted from financial documents and earnings reports.
Dataset Structure
Data fields: context, question, answer.
Licensing Information
This dataset is released under the MIT license (mit).
Question-Answering-V.2
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/kunarmy99/Question-Answering-V.2.Retrieval-Augmented-Question-Answering
🇰🇿 Retrieval-Augmented Question Answering in Kazakh Context
Dataset Summary
Retrieval-Augmented Question Answering (RAG), Kazakh Context is a specialized dataset designed to train Large Language Models (LLMs) to accurately answer complex questions by drawing strictly from provided external knowledge sources in the Kazakh language.
This dataset teaches models to synthesize information from multiple retrieved documents, compare concepts, and ground their answers… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Retrieval-Augmented-Question-Answering.prime-survey-question-answering
PRIME Survey Dataset of Minoritised Ethnic People’s Engagement with Online Services
Our dataset is now publicly available via the university's open access repository:
DOI: 10.17861/db813826-e45d-4274-b4c3-7ecdbf2336a5
License: This dataset is released under the Creative Commons Attribution 4.0 International (CC BY 4.0)license.
Note: Please use the DOI link above to access and download the data.
This directory is designated for dataset documentation, metadata, and any… See the full description on the dataset page: https://huggingface.co/datasets/JunhaoSong/prime-survey-question-answering.question_answering
Dataset Information
This Question Answering dataset is a reading comprehension resource derived from Persian Wikipedia. This crowd-sourced dataset contains over 9,000 entries, each of which can either be an unanswerable question or a question with one or more answers based on the provided context. Similar to the SQuAD2.0 dataset, the inclusion of unanswerable questions allows for the development of systems that "know they don't know the answer." Additionally, the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/azizmatin/question_answering.QuestionAnsweringProblemSolving
Practical Problem Solving QA Dataset
This dataset focuses on question–answer pairs related to structured thinking, problem solving, and decision-making processes.
Dataset Structure
Each record contains:
question: A practical or conceptual question
answer: A concise and logical response
Intended Use
Suitable for:
Question answering models
Reasoning and analysis tasks
Educational and evaluation purposes
General-purpose language models
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/joey4/QuestionAnsweringProblemSolving.SQAD-Sinhala_Question_Answering_DatasetThis dataset is a back-translated version of the SQuAD 2.0 dataset, translated into Sinhala using the Google Cloud Translate API by Sachin Hansaka.
Original dataset by the Stanford QA Group: https://rajpurkar.github.io/SQuAD-explorer/
Original work licensed under CC BY-SA 4.0.
This Sinhala version © 2025 Sachin Hansaka, also licensed under CC BY-SA 4.0.
📚 Dataset Overview
SQAD-Sinhala_Question_Answering_Dataset is a high-quality, back-translated version of the original SQuAD 2.0 dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sachin-Hansaka/SQAD-Sinhala_Question_Answering_Dataset.
