datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CFQA_Chinese_Finance_Question_Answering
Citation
For the complete project, please check Here
If you use CFQA in your research, experiments, benchmarks, or publications, please cite the accompanying paper:
@inproceedings{zhu2026cfqa,
title = {CFQA: A Chinese Financial Question Answering Benchmark From Corporate Annual Reports},
author = {Tianning Zhu and Mo Liu and Murathan Kurfali},
booktitle = {Proceedings of The 7th Financial Narrative Processing Workshop (FNP 2026)},
year = {2026},
address =… See the full description on the dataset page: https://huggingface.co/datasets/ZackZhu00/CFQA_Chinese_Finance_Question_Answering.chaii-hindi-and-tamil-question-answeringquestion-answering-state-of-the-unionquestion-answering-paul-grahambodo-legal-question-answering-ai4bharat
Bodo Legal Question Answering Dataset
Overview
This dataset is a Bodo-language legal Question Answering (QA) resource
created for research in low-resource Natural Language Processing (NLP)
and legal language processing.
The supplied source files contain legal judgment contexts together with
multiple questions and answers. For Hugging Face compatibility and
question-answering model training, each question-answer pair has been
flattened into a separate JSONL example… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-ai4bharat.bodo-legal-question-answering-iiith
Bodo Legal Question Answering Dataset — IIITH Translation
Overview
A Bodo-language legal Question Answering (QA) resource derived from
English legal judgments. Each example contains a judgment context, a
question, and its corresponding answer.
Data Provenance
Original Legal Source
The underlying English legal judgments were extracted from the publicly
accessible Gauhati High Court judgment repository:… See the full description on the dataset page: https://huggingface.co/datasets/Mwnthai/bodo-legal-question-answering-iiith.Dengue_Surveillance_Data_Question_Answering_Dataset
Nepali ShareGPT Clean Final Dataset
Comprehensive Documentation & Analysis Report
Dengue Surveillance Data - Question Answering Dataset
📋 Dataset Overview
This dataset is a curated collection of 256 question-answer pairs focused on Dengue Surveillance in Nepal. It contains data-grounded questions in Nepali language paired with factual, statistical answers sourced from the Department of Health Services (DoHS), Nepal. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Dengue_Surveillance_Data_Question_Answering_Dataset.questionanswering-datasetBenchMAX_Question_Answering
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Question_Answering is a dataset of BenchMAX for evaluating the long-context capability of LLMs in multilingual scenarios.
The subtasks are similar to the subtasks in RULER.
The data is sourcing from UN Parallel Corpus and xquad.
The haystacks… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Question_Answering.Question-Answering_Kazakh
🇰🇿 Question-Answering_Kazakh
A comprehensive Kazakh-language question-answer dataset for fine-tuning
and training language models.Created and maintained by Kurumikz. Free to use with attribution.
📌 Overview
Question-Answering_Kazakh is an open-domain QA dataset written entirely
in the Kazakh language (kk). It covers a wide range of topics — from the
history and geography of Kazakhstan to Kazakh grammar, culture, economy, and
language learning (Kazakh ↔ English).… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-Answering_Kazakh.Question-answeringsmall-ru
Dataset Card for Question Answering Russian Dataset
🧠 Quick Summary
Небольшой, чистый и тестовый датасет, созданный энтузиастом.Содержит базовые фундаментальные знания по математике, странам и тюркским народам.Подходит для обучения и тестирования моделей в образовательных и исследовательских целях.
📚 Dataset Details
Curated by: @kurumikz
Language(s): Russian (ru)
License: CC-BY 4.0 — свободное использование с обязательным указанием автора
Size Category:… See the full description on the dataset page: https://huggingface.co/datasets/kurumikz/Question-answeringsmall-ru.Chinese_Question_Answering_Datasetmedical-question-answering-datasets-alpacaquestion_answering
Dataset Information
This Question Answering dataset is a reading comprehension resource derived from Persian Wikipedia. This crowd-sourced dataset contains over 9,000 entries, each of which can either be an unanswerable question or a question with one or more answers based on the provided context. Similar to the SQuAD2.0 dataset, the inclusion of unanswerable questions allows for the development of systems that "know they don't know the answer." Additionally, the dataset includes… See the full description on the dataset page: https://huggingface.co/datasets/azizmatin/question_answering.medical_question_answering_datasetsBias-Question-Answering
Dataset Card for QA Bias Detection Dataset
Summary
Description: This dataset is designed for the task of bias detection in text, particularly focusing on dimensions of ageism and sentiment analysis. It contains question-answer pairs that assess potential biases in statements.
Purpose: To facilitate research and development in the areas of bias detection, natural language understanding, and sentiment analysis.
Supported Tasks: Bias detection, sentiment analysis, natural language… See the full description on the dataset page: https://huggingface.co/datasets/newsmediabias/Bias-Question-Answering.QuestionAnsweringProblemSolving
Practical Problem Solving QA Dataset
This dataset focuses on question–answer pairs related to structured thinking, problem solving, and decision-making processes.
Dataset Structure
Each record contains:
question: A practical or conceptual question
answer: A concise and logical response
Intended Use
Suitable for:
Question answering models
Reasoning and analysis tasks
Educational and evaluation purposes
General-purpose language models
Data Format… See the full description on the dataset page: https://huggingface.co/datasets/joey4/QuestionAnsweringProblemSolving.SQAD-Sinhala_Question_Answering_DatasetThis dataset is a back-translated version of the SQuAD 2.0 dataset, translated into Sinhala using the Google Cloud Translate API by Sachin Hansaka.
Original dataset by the Stanford QA Group: https://rajpurkar.github.io/SQuAD-explorer/
Original work licensed under CC BY-SA 4.0.
This Sinhala version © 2025 Sachin Hansaka, also licensed under CC BY-SA 4.0.
📚 Dataset Overview
SQAD-Sinhala_Question_Answering_Dataset is a high-quality, back-translated version of the original SQuAD 2.0 dataset… See the full description on the dataset page: https://huggingface.co/datasets/Sachin-Hansaka/SQAD-Sinhala_Question_Answering_Dataset.
