datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
extractive_qa_question_answering_hr
Dataset Card
HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research.
Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction.
Dataset Sources
Repository: xwjzds/extractive_qa_question_answering_hr
Paper: HR-MultiWOZ: A Task Oriented Dialogue (TOD)… See the full description on the dataset page: https://huggingface.co/datasets/xwjzds/extractive_qa_question_answering_hr.OWASP-question-answer-datasetFinancial_Question_Answeringtamil-question-answering-datasetthis dataset contains 5 columns
context, question, answer_start, answer_text, source
Column
Description
context
A general small paragraph in tamil language
question
question framed form the context
answer_text
text span that extracted from context
answer_start
index of answer_text
source
who framed this context, question, answer pair
source
team KBA => (Karthi, Balaji, Azeez) these people manually created
CHAII =>a kaggle competition
XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/AswiN037/tamil-question-answering-dataset.question-answerOWASP-and-NVD-question-answer-datasetquestion-answering-ukrainian-json-answersQuestion-Answering-Generation-Choices
The dataset is a merged compilation of QuAIL, RACE, and Cosmos QA datasets,
having undergone preprocessing.
question-answering-ukrainianDermatology-Question-Answer-Dataset-For-Fine-Tuning
Dataset Details
The data set has about 1 Million Tokens for Training and about 1500 question answers.
Dataset Description
This dataset is a comprehensive compilation of questions related to dermatology, spanning inquiries about various skin diseases, their symptoms, recommended medications, and available treatment modalities. Each question is paired with a concise and informative response, making it an ideal resource for training and fine-tuning language models in the… See the full description on the dataset page: https://huggingface.co/datasets/Mreeb/Dermatology-Question-Answer-Dataset-For-Fine-Tuning.cncf-question-and-answer-dataset-for-llm-training
CNCF QA Dataset for LLM Tuning
Description
This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model.
The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-question-and-answer-dataset-for-llm-training.fitness-question-answersA total of 965 q&a pairs i gathered from the web related to physical activity and fitness.
Santali-Ol-Chiki-Agriculture_Question-Answer_DatasetSantali (Ol Chiki) Agriculture Question-Answer Dataset is a curated collection of question–answer pairs focused on agricultural knowledge in the Santali language, written in the Ol Chiki script. The dataset consists of question-answer pairs in the Santali language focusing on agriculture, animal husbandry, and rural health topics. The content covers crop diseases, soil management, livestock care, and farming techniques tailored for tribal communities. This dataset is designed to support… See the full description on the dataset page: https://huggingface.co/datasets/nharshavardhana/Santali-Ol-Chiki-Agriculture_Question-Answer_Dataset.dais-question-answers
DAIS-Question-Answers Dataset
This dataset contains question-answer pairs created using ChatGPT using text data scraped from the Databricks Data and AI Summit 2023 (DAIS 2023) homepage
as well as text from any public page that is linked in that page or is a two-hop linked page.
We have used this dataset to fine-tune our DAIS DLite model, along with our dataset of webpage texts. Feel free to check them out!
Note that, due to the use of ChatGPT to curate these question-answer pairs… See the full description on the dataset page: https://huggingface.co/datasets/aisquared/dais-question-answers.QuestionAnswer_MCQarabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.siddha_vaithiyam_question_answering_chatbot
Medical Home Remedy Chatbot Dataset
Overview
This dataset is designed for a chatbot that answers questions related to medical problems with simple home remedies. The information in this dataset has been sourced from old books containing traditional remedies used in the past.
Contents
Dataset Files:
dataset.csv : The main dataset file containing questions and corresponding home remedy answers.
Data Structure:
Each row in the CSV file… See the full description on the dataset page: https://huggingface.co/datasets/RahulS3/siddha_vaithiyam_question_answering_chatbot.tamil-question-answering-datasetthis dataset contains 5 columns
context, question, answer_start, answer_text, source
Column
Description
context
A general small paragraph in tamil language
question
question framed form the context
answer_text
text span that extracted from context
answer_start
index of answer_text
source
who framed this context, question, answer pair
source
team KBA => (Karthi, Balaji, Azeez) these people manually created
CHAII =>a kaggle competition
XQA => multilingual QA… See the full description on the dataset page: https://huggingface.co/datasets/Subi1152/tamil-question-answering-dataset.OpenMP_Question_Answering
OpenMP Question Answering Dataset
OpenMP Question Answering Dataset is a new OpenMP question answering introduced in paper "LM4HPC: Towards Effective Language Model Application in High-Performance Computing".
It is designed to probe the capabilities of language models in single-turn interactions with users. Similar to other QA datasets, we include
some request-response pairs which are not strictly question-answering pairs. The categories and examples of questions in the OMPQA… See the full description on the dataset page: https://huggingface.co/datasets/chenle015/OpenMP_Question_Answering.Football_Question_AnswersNVD-question-answer-datasetcrewai_docs_question_answer_test1elby_the_elephant_question_and_answersquestion_answerquestion_answer_finetuning_embeddings.csvquestion-without-answersIslam_Question_and_Answer1question_answerquestion_answerquestionAnswer
