datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Question-Answering-Generation-Choices
The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets,
having undergone preprocessing.
fitness-question-answersA total of 965 q&a pairs i gathered from the web related to physical activity and fitness.
CNTXTAI_Medical_Questions_AnswersThis dataset is highly valuable for medical research, categorization, and analysis. The structured format allows for efficient information retrieval and classification, making it a well-maintained reference for academic and clinical research. A rigorous validation process ensures credibility, making this dataset reliable for further study and application.
General Overview
Total Number of Rows: 50 (excluding headers)
Total Number of Columns: 3
Column Headers and Data Types:
Question: Text… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/CNTXTAI_Medical_Questions_Answers.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.albanian_legal_questions_answerssiddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.llm-answer-set-qa
Answer-Set Consistency Benchmark (ASCB)
Overview
The Answer-Set Consistency Benchmark (ASCB) evaluates whether language models provide mutually consistent answers to related factual enumeration questions. Unlike conventional QA datasets, ASCB focuses on whether generated answer sets satisfy known set-theoretic relations rather than solely on factual accuracy.
ASCB contains 600 English question quadruples (2,400 questions) across primarily static, objective factual… See the full description on the dataset page: https://huggingface.co/datasets/anonymous2026nips/llm-answer-set-qa.qrecc_context_and_answersThis is the QRECC dataset arranged to be used for a query rewriting task. Each line is composed as follow:
INTRODUCTION token is followed by the PREVIOUS TURNS of the same conversation, which means only the previous questions
QUESTION token is followed by the current query the system should rewrite
ANSWER token is followed by the REWRITING of the current query + the given ANSWER
Islam_Question_and_Answer1qrecc_with_context_and_answersThis is the QRECC dataset arranged to be used for a query rewriting task.
Each line is composed as follow:
INTRODUCTION token is followed by the PREVIOUS TURNS of the conversation together WITH THE ANSWERS provided in the dataset
QUESTION token is followed by the current query the system should rewrite
ANSWER token is followed by the rewriting of the current query
yahoo_answers_matchesreligious-questions-and-answers
Main fields
article_id, url, title, question, short_answer, content_html,
content_text, published_at_persian, view_count, and category fields.
is_valid_article marks archive links that resolved to a valid article page.
