datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pidgin-QandA-data-samples
Pidgin Question-Answer Dataset (Sample)
Sample dataset: Nigerian Pidgin conversational Q&A for dialogue systems and language modeling
🤗 Hugging Face • 📊 Figshare • 🌐 Website • 📧 Contact
📋 Overview
The Pidgin Question-Answer Dataset (Sample) is a conversational corpus containing 1,462 question-answer pairs entirely in Nigerian Pidgin English. Created by Bytte AI through AI chatbot interactions with human validation, this sample dataset supports dialogue… See the full description on the dataset page: https://huggingface.co/datasets/Bytte-AI/Pidgin-QandA-data-samples.who-QandA
WHO Q&A Dataset
Dataset Description
This dataset contains question-and-answer pairs sourced from the World Health
Organization (WHO), covering topics such as [e.g. disease outbreaks, vaccination,
nutrition, mental health, health emergencies]. It was built to support tasks
like question answering, retrieval-augmented generation (RAG), and health
information chatbots.
License: [check WHO's actual terms — WHO content is often
CC BY-NC-SA 3.0 IGO; verify before… See the full description on the dataset page: https://huggingface.co/datasets/miguellong/who-QandA.ayurveda-text-based-qandaReasoning-Socratic-QandA
Reasoning-Socratic-QandA Dataset
This dataset is a curated mixture of three high-quality data sources, specifically engineered to train Socratic Tutors. It balances the model's ability to think deeply (Reasoning), guide students through questioning (Pedagogy), and provide direct answers when necessary (Support).
Dataset Composition
Dataset
Role in your tutor
Size
Weight
DASD Stage 1
Teaches deep reasoning — how to work through math/code/science problems… See the full description on the dataset page: https://huggingface.co/datasets/Bialy17/Reasoning-Socratic-QandA.
