datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OWASP-and-NVD-question-answer-datasetchaii-hindi-and-tamil-question-answeringcncf-question-and-answer-dataset-for-llm-training
CNCF QA Dataset for LLM Tuning
Description
This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model.
The dataset includes the… See the full description on the dataset page: https://huggingface.co/datasets/Kubermatic/cncf-question-and-answer-dataset-for-llm-training.ptbr-question-and-answer
Perguntas e Respostas Brasileiras
Esse dataset é uma compilação das perguntas e respostas em português disponíveis em clips/mqa.
Foi realizada uma limpeza e normalização dos dados, mantendo apenas domínios mais relevantes, removendo texto danosos e inadequados.
O código para a limpeza dos dados pode ser acessado aqui
O principal objetivo deste dataset é ajudar modelos de linguagem natural e modelos de embedding em português a gerar textos e cálculos de similaridade
mais precisos e… See the full description on the dataset page: https://huggingface.co/datasets/emdemor/ptbr-question-and-answer.piaf_fr_prompt_context_generation_with_answer_and_question
piaf_fr_prompt_context_generation_with_answer_and_question
Summary
piaf_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 442,752 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_context_generation_with_answer_and_question.arabic_dialects_question_and_answerData Content
The file provided: Q/A Reasoning dataset
contains the following columns:
ID # : Denotes the reference ID for:
a. Question
b. Answer to the question
c. Hint
d. Reasoning
e. Word count for items a to d above
Dialects: Contains the following dialects in separate columns:
a. English
b. MSA
c. Emirati
d. Egyptian
e. Levantine Syria
f. Levantine Jordan
g. Levantine Palestine
h. Levantine Lebanon
Data Generation Process
The following are the steps that were followed to curate the data:… See the full description on the dataset page: https://huggingface.co/datasets/CNTXTAI0/arabic_dialects_question_and_answer.pcos-patient-assist-question-and-answer
Dataset Card for PCOS Patient Assist Question and Answer Dataset
Dataset Details
Dataset Description
The PCOS Patient Assist Question and Answer Dataset is a curated dataset of question–answer pairs designed to represent common questions asked by patients diagnosed with or concerned about Polycystic Ovary Syndrome (PCOS).
The dataset is structured to simulate real patient queries that arise during different stages of the PCOS journey, including diagnosis… See the full description on the dataset page: https://huggingface.co/datasets/Khyatimirani/pcos-patient-assist-question-and-answer.squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question
Summary
squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 1,271,928 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_context_generation_with_answer_and_question.race_high_Read_the_article_and_answer_the_question_no_option_DeepEval-Question-Answer-Dataset-for-RAG-Evaluation-A2A-And-ACP-PDFnewsquadfr_fr_prompt_context_generation_with_answer_and_question
newsquadfr_fr_prompt_context_generation_with_answer_and_question
Summary
newsquadfr_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 101,040 rows that can be used for a context-generation (with answer)task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_context_generation_with_answer_and_question.newsquadfr_fr_prompt_question_generation_with_answer_and_context
newsquadfr_fr_prompt_question_generation_with_answer_and_context
Summary
newsquadfr_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 88,410 rows that can be used for a question generation (with answer and context) task.The original data (without prompts) comes from the dataset newsquadfr and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/newsquadfr_fr_prompt_question_generation_with_answer_and_context.piaf_fr_prompt_question_generation_with_answer_and_context
piaf_fr_prompt_question_generation_with_answer_and_context
Summary
piaf_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 387,408 rows that can be used for a question-generation (with answer and context) task.The original data (without prompts) comes from the dataset PIAF and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
A list of prompts (see below) was then applied in order to… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/piaf_fr_prompt_question_generation_with_answer_and_context.squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context
squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context
Summary
squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 1,112,937 rows that can be used for a question-generation (with answer and context) task.The original data (without prompts) comes from the dataset pragnakalp/squad_v2_french_translated and was augmented by questions in SQUAD 2.0 format in the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/squad_v2_french_translated_fr_prompt_question_generation_with_answer_and_context.YouTube-Transcript-Question-Answer-Dataset-for-RAG-Evaluation-A2A-And-ACP-PDFKaggle-post-and-comments-question-answer-topic
This is a dataset containing 10,000 posts from Kaggle and 60,000 comments related to those posts in the question-answer topic.
Data Fields
kaggle_post
'pseudo', The question authors.
'title', Title of the Post.
'question', The question's body.
'vote', Voting on Kaggle is similar to liking.
'medal', I will share with you the Kaggle medal system, which can be found at https://www.kaggle.com/progression. The system awards medals to users based on their… See the full description on the dataset page: https://huggingface.co/datasets/Raaxx/Kaggle-post-and-comments-question-answer-topic.legal-question-and-answerfquad_fr_prompt_question_generation_with_answer_and_context
fquad_fr_prompt_question_generation_with_answer_and_context
Summary
fquad_fr_prompt_question_generation_with_answer_and_context is a subset of the Dataset of French Prompts (DFP).It contains 502,299 rows that can be used for a question-generation (with answer and context) task.The original data (without prompts) comes from the dataset FQuAD by d'Hoffschmidt et al. and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
As FQuAD's license does not allow… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/fquad_fr_prompt_question_generation_with_answer_and_context.fquad_fr_prompt_context_generation_with_answer_and_question
fquad_fr_prompt_context_generation_with_answer_and_question
Summary
fquad_fr_prompt_context_generation_with_answer_and_question is a subset of the Dataset of French Prompts (DFP).It contains 574,056 rows that can be used for a context-generation (with answer and question) task.The original data (without prompts) comes from the dataset FQuAD by d'Hoffschmidt et al. and was augmented by questions in SQUAD 2.0 format in the FrenchQA dataset.
As FQuAD's license does not allow… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/fquad_fr_prompt_context_generation_with_answer_and_question.race_middle_Read_the_article_and_answer_the_question_no_option_dream_read_the_following_conversation_and_answer_the_question_subsetbiology-120-question-and-answerSINHALA_QUESTION_AND_ANSWER_DATASET
SINHALA QUESTION AND ANSWER DATASET by Indramal
Contact details for access this: Indramal Wansekara Profile Website
Research Paper
Website: https://indramal.github.io/SinhalaQandA/
Paper: https://ieeexplore.ieee.org/document/10841868
Citation
If you want to cite research paper, you can use this:
@INPROCEEDINGS{10841868,
author={Wansekara, Indramal and Jayasekara, A.G.B.P.},
booktitle={2024 4th International Conference on Electrical Engineering (EECon)}… See the full description on the dataset page: https://huggingface.co/datasets/Indramal/SINHALA_QUESTION_AND_ANSWER_DATASET.elby_the_elephant_question_and_answersIslam_Question_and_Answer
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/hmmamalrjoub/Islam_Question_and_Answer.steve-jobs-question-and-answersIslam_Question_and_Answer1mirror-cncf-question-and-answer-dataset-for-llm-training
CNCF QA Dataset for LLM Tuning
Description
This dataset, named cncf-qa-dataset-for-llm-tuning, is designed for fine-tuning large language models (LLMs) and is formatted in a question-answer (QA) style. The data is sourced from PDF and markdown (MD) files extracted from various project repositories within the CNCF (Cloud Native Computing Foundation) landscape. These files were processed and converted into a QA format to be fed into the LLM model.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-cncf-question-and-answer-dataset-for-llm-training.dream_read_the_following_conversation_and_answer_the_questionflan_source_race_high_Read_the_article_and_answer_the_question_no_option__73
