datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arabic-qna
Sadeem QnA: An Arabic QnA Dataset 🌍✨
Welcome to the Sadeem QnA dataset, a vibrant collection designed for the advancement of Arabic natural language processing, specifically tailored for Question Answering (QnA) systems. Sourced from the rich and diverse content of Arabic Wikipedia, this dataset is a gateway to exploring the depths of Arabic language understanding, offering a unique challenge to both researchers and AI enthusiasts alike.
About Sadeem QnA
The Sadeem… See the full description on the dataset page: https://huggingface.co/datasets/sadeem-ai/arabic-qna.online_privacy_qnaOnline Privacy Policy QnA Dataset
Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
QnA_Descriptivemagicmotion
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
Quanhao Li*, Zhen Xing*, Rui Wang, Hui Zhang, Qi Dai, and Zuxuan Wu
* equal contribution
💡 Abstract
Recent advances in video generation have led to remarkable improvements in visual quality and temporal coherence. Upon this, trajectory-controllable video generation has emerged to enable precise object motion control through explicitly defined spatial paths.
However, existing methods… See the full description on the dataset page: https://huggingface.co/datasets/Qnancy/magicmotion.alodokter-qna
Dataset Question Answer Health Indonesian
Dataset Summary
The Question Answer Health Indonesian dataset contains +250,000 question-and-answer pairs related to health topics sourced from the Alodokter website. The dataset spans a collection period from July 2023 to September 2023 (approximately 2 months). It is designed to facilitate research and development in the fields of natural language processing (NLP), particularly for Indonesian language models, health information… See the full description on the dataset page: https://huggingface.co/datasets/agufsamudra/alodokter-qna.tesla-qna-feedback-logs
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/pgurazada1/tesla-qna-feedback-logs.qdrant_doc_qnamyanmar_qna_dataset
Myanmar QnA Dataset v7
Language: Burmese (Myanmar)Total Entries: 22,783 QnA pairsTotal Sentences: ~ 466,330(Counted using the Myanmar sentence-ending symbol "။")License: CC0 1.0 (Public Domain)
Description
This dataset contains Myanmar-language question-answer pairs (QnA) generated with the assistance of ChatGPT-5 for question crafting with English and Gemini 3.0 Pro for Myanmar QnA generation. It is intended for research, AI training, and educational purposes.
Each entry… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar_qna_dataset.indonesia-medical-qna
indonesia-medical-qna
This dataset is configured so the Hugging Face Dataset Viewer loads qna.csv as the primary data file.
The repository also contains all_links.csv, which has a different schema and is kept as an auxiliary file rather than part of the default viewer configuration.
This avoids schema-casting errors caused by the Hub trying to combine both CSV files into one split.
qna-quantum-information
Q&A Quantum Information Dataset
This dataset is created by digesting 500 different papers from the quantum information directory on arXiv,
the papers are located based on their relevance to the quantum information keyword.
Data retrieval
The data is extracted from the .pdf files using PyMuPDF package proxied from langchain.
Then the Q&A pair is generated by:
Generate N questions per page of the .pdf document based on its content.
We will feed each question to an LLM… See the full description on the dataset page: https://huggingface.co/datasets/CoAILab/qna-quantum-information.bangla-law-qnaalodokter-qna
Dataset Question Answer Health Indonesian
Dataset Summary
The Question Answer Health Indonesian dataset contains +250,000 question-and-answer pairs related to health topics sourced from the Alodokter website. The dataset spans a collection period from July 2023 to September 2023 (approximately 2 months). It is designed to facilitate research and development in the fields of natural language processing (NLP), particularly for Indonesian language models, health… See the full description on the dataset page: https://huggingface.co/datasets/coegggg/alodokter-qna.BNF_QNAQnA_for_Non-Technical_Roles
Columns:
Non-Technical Role: Specifies the role being assessed (e.g., Content Developer).
Assessment Domain: Denotes the skill or attribute being evaluated (e.g., Adaptability).
Question: The actual assessment question.
Content Overview:
The dataset is focused on assessing various competencies and skills relevant to non-technical roles.
Questions are tailored to evaluate how individuals in these roles handle various situations and challenges.
Example Entries:
Role: Content Developer… See the full description on the dataset page: https://huggingface.co/datasets/Sumsam/QnA_for_Non-Technical_Roles.workshop_QnAwikipedia-id-qna
Wikipedia ID Synthetic QnA
This dataset contains synthetic question-answer pairs (QnA) generated using DeepSeek from Indonesian Wikipedia articles. The data has been sourced from this Wikipedia dataset, which contains a subset of Indonesian Wikipedia articles. Each entry includes a context, a related question and answer pair, and an unrelated question.
Dataset Structure
The dataset contains the following columns:
id: A unique identifier for each row.
context: A… See the full description on the dataset page: https://huggingface.co/datasets/vitoghif/wikipedia-id-qna.QnAMedicDaatasetqdrant_docs_qna_ragasQnA_20240513_122841nbme_qnaquestion-context-answer format rows for nbme qna fine-tuning
Human-Like-Gut-Health-DPO-QnA
Gut Health DPO Dataset
Overview
This dataset contains 200 carefully curated examples for Direct Preference Optimization (DPO) training in the domain of gut health and digestive wellness. Each example consists of a user prompt, a "chosen" response (preferred), and a "rejected" response (less preferred), designed to train AI models to provide high-quality, medically responsible advice on digestive health topics.
Dataset Structure
The dataset is provided in CSV… See the full description on the dataset page: https://huggingface.co/datasets/Sid3503/Human-Like-Gut-Health-DPO-QnA.network-QnA-datasethr_assistant_qna_datasetwiki-ro-qna
Description
There are more than 550k questions with roughly 53k paragraphs. The questions were built using the ChatGPT 3.5 API.
The dataset is based on the Romanian Wikipedia 2020 June dump, curated by Dumitrescu Stefan.
The paragraphs retained are those between 100 and 410 words (roughly 512 max tokens), using the following script:
# Open the text file
with open('wiki-ro/corpus/wiki-ro/wiki.txt.train', 'r') as file:
# Read the entire content… See the full description on the dataset page: https://huggingface.co/datasets/catalin1122/wiki-ro-qna.customer-review-qna-25
📦 Dataset: customer-review-qna-25
🧾 Summary
A small dataset of 25 GPT-4o-mini generated customer reviews, each containing:
input_question: Prompt or user query
output: Generated customer review
context: Background info (e.g. product, tone)
tags: Labels like positive, complaint, delivery, etc.
✅ safe ✅ Filtered ✅ Compact & usable for review generation or sentiment tasks
🧱 Dataset Structure
Format: CSVSize: 25 samplesFields:
Field
Type… See the full description on the dataset page: https://huggingface.co/datasets/elvanalabs/customer-review-qna-25.Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
tcm-qnaKNF-Methods-QnAWork in progress. Not yet reviewed by domain experts.
Neuro-sama-QnAThis dataset was manually created, line by line, by my tiny hand!
Why? Because I was just bored during my summer.
