datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime-1983-2025
AIME Datasets from 1983 to 2025
This dataset contains the AIME datasets from 1983 to 2025.
For AIME 1983 to 2026 use Pandores/aime-1983-2026
Features Description
Feature
Description
Example
year
The year this problem was released. From 1983 to 2025.
2022
index
The index of the problem for a year and part. From 1 to 15.
12
part
The dataset part if this dataset has multiple parts. Can be AIME, AIME I, AIME II or None. Datasets have multiple parts… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2025.PangeaInstruct
PangeaInstruct
Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages
🇪🇹 🇸🇦 🇧🇬 🇧🇩 🇨🇿 🇩🇪 🇬🇷 🇬🇧 🇺🇸 🇪🇸 🇮🇷 🇫🇷 🇮🇪 🇮🇳 🇮🇩 🇳🇬 🇮🇹 🇮🇱 🇯🇵 🇮🇩 🇰🇷 🇳🇱 🇲🇳 🇲🇾 🇳🇴 🇵🇱 🇵🇹 🇧🇷 🇷🇴 🇷🇺 🇱🇰 🇮🇩 🇰🇪 🇹🇿 🇱🇰 🇮🇳 🇮🇳 🇹🇭 🇹🇷 🇺🇦 🇵🇰 🇮🇳 🇻🇳 🇨🇳 🇹🇼
🏠 Homepage | 🤖 Pangea-7B | 📊 PangeaIns | 🧪 PangeaBench | 💻 Github | 📄 Arxiv | 📕 PDF | 🖥️ Demo
This README provides comprehensive details on the PangeaIns dataset, which… See the full description on the dataset page: https://huggingface.co/datasets/neulab/PangeaInstruct.EngineMT-QA
EngineMT-QA Dataset
Overview
EngineMT-QA is a large-scale, multi-task, multimodal dataset for Time-Series Question Answering (Time-Series QA). It enables research on aligning multivariate time-series signals with natural language through four key cognitive tasks:
Understanding
Perception
Reasoning
Decision-Making
The dataset is built on N-CMAPSS, simulating real-world aero-engine operational and maintenance scenarios. It supports the development and evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/pandalin98/EngineMT-QA.PangeaBench-tydiqa
Dataset Card for "tydiqa"
Dataset Summary
TyDi QA is a question answering dataset covering 11 typologically diverse languages with 204K question-answer pairs.
The languages of TyDi QA are diverse with regard to their typology -- the set of linguistic features that each language
expresses -- such that we expect models performing well on this set to generalize across a large number of the languages
in the world. It contains language phenomena that would not be found in… See the full description on the dataset page: https://huggingface.co/datasets/neulab/PangeaBench-tydiqa.PangeaBench-xmmmuCAGBCredibility-aware Generation Benchmark (CAGB) is a benchmark constructed to evaluate the credibility-aware generation ability of models, dealing with flawed information in the context.
This benchmark encompasses the following three specific scenarios where the integration of credibility is essential:
Open-domain QA
2WikiMultiHopQA
HotpotQA
Musique
RGB
Time-sensitive QA
EvoTempQA
Misinformation Polluted QA
NewsPollutedQA
Read our paper for more insights on credibility-aware generation.
thema-panhellenic-exams
Thema
Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments.
Website · Code · Leaderboard · Method
Dataset summary
Current release
Exam years
2023 to 2026
Published subjects
10 of 10
Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.PangeaBench-cvqa
About CVQA
CVQA is a culturally diverse multilingual VQA benchmark consisting of over 9,000 questions from 33 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories.
This data is designed for use as a test set. Please submit your submission here to evaluate your model performance. CVQA is constructed through a collaborative effort led by a team of researchers from MBZUAI. Read more about CVQA… See the full description on the dataset page: https://huggingface.co/datasets/neulab/PangeaBench-cvqa.pan-african-primary-care-benchmark
Pan-African Primary Care Benchmark (v1)
A multilingual safety and reasoning benchmark for clinical AI in African primary-care contexts
Overview
This dataset is designed to rigorously evaluate how well large language models (and clinical AI agents) perform when patients present in real-world African languages — exactly as they do in clinics across the continent.
v1 contains 300 synthetic, de-identified primary-care scenarios (50 per language × 6 files):
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/nimrodzw/pan-african-primary-care-benchmark.pandora-data
Pandora Benchmark Data
Versioned, processed benchmark annotations for the
Pandora structured knowledge reasoning
framework.
Contents
Dataset
Split
Records
Artifact license
Spider-Syn
test
1,034
MIT
BIRD
dev
1,534
CC BY-SA 4.0
WikiTableQuestions
test
4,344
CC BY-SA 4.0
WikiSQL
test
15,878
BSD-3-Clause
GrailQA
evaluation set from public validation data
6,409
CC BY-SA 4.0 annotations; Freebase CC BY 2.5 BOX data
WebQSP
test
1,598
Freebase CC BY… See the full description on the dataset page: https://huggingface.co/datasets/bahuia/pandora-data.pandas-create-context
Overview
This dataset is built from sql-create-context, which in itself builds from WikiSQL and Spider.
I have used GPT4 to translate the SQL schema into pandas DataFrame schem initialization statements and to translate the SQL queries into pandas queries.
There are 862 examples of natural language queries, pandas DataFrame creation statements, and pandas query answering the question using the DataFrame creation statement as context. This dataset was built with text-to-pandas… See the full description on the dataset page: https://huggingface.co/datasets/hiltch/pandas-create-context.Panini-Benchmarks
Panini: Continual Learning in Token Space via Structured Memory
Extracted Generative Semantic Workspace (GSW) representations and curated evaluation splits from Panini, provided for ease of replication and future research.
Contents
GSW Networks (gsw_networks/)
Structured semantic representations extracted from document corpora using the Panini/GSW framework. Each GSW captures entities, their roles/states, and verb-phrase relationships as question-answer pairs.… See the full description on the dataset page: https://huggingface.co/datasets/roychowdhuryresearch/Panini-Benchmarks.newsophy-v0.1
This dataset was used to train the pansophic-1-preview model
This dataset was created using open-source, permissively licensed models. In addition to providing answers to a diverse set of questions, we leveraged multiple open-source pipelines to generate new tasks and questions, enriching the dataset's variety and complexity. The dataset includes examples that showcase tool usage, contextual understanding, and the application of system prompts.
Topics distribtuion in… See the full description on the dataset page: https://huggingface.co/datasets/pansophic/newsophy-v0.1.aime-1983-2026
AIME Datasets from 1983 to 2026
This dataset contains all the AIME problems from 1983 to 2026. For a total of 1065 problems.
Example
Download
from datasets import load_dataset
dataset = load_dataset("Pandores/aime-1983-2026")
print(dataset["train"][0])
Download and iterate
from datasets import load_dataset
dataset = load_dataset("Pandores/aime-1983-2026", split="train")
for entry in dataset:
print(entry["problem"])… See the full description on the dataset page: https://huggingface.co/datasets/Pandores/aime-1983-2026.PanDomain-V1.3PanDomain-V1 is a high-quality, fully English dataset designed for training generalist language models across all major domains. It serves as the foundational training corpus for the Talon model family, built to support broad capabilities in both reasoning and generation.
Every model sees everything.
PanAfriQA
PanAfriQA
PanAfriQA is a question-answering (QA) dataset focused on African history, designed to provide a resource for exploring and understanding the continent's rich historical narrative.
It includes a collection of questions paired with accurate answers, covering key events, figures, cultures, and developments across African history.
The dataset aims to support research, education, and the development of AI systems by offering a structured and accessible way to engage with… See the full description on the dataset page: https://huggingface.co/datasets/Svngoku/PanAfriQA.Panther-dataset_v1
Dataset Details
This dataset is a modified version of Anthropic/hh-rlhf
This dataset is used in fine tuning Panther - an state of the art LLM funtuned on llama-7b pretrained model.
A very small portion i.e. 5.3% of prompts and responses were taken from this dataset to finetune and train Panther
Dataset Details
Dataset Structure
Train
Train rows : 377k
Validation
Validation rows : 20.3k
Dataset Format
input… See the full description on the dataset page: https://huggingface.co/datasets/Rardilit/Panther-dataset_v1.COVID-QA-el-small
Dataset Card for COVID-QA-el-small
Dataset Summary
The COVID-QA-el-small dataset is a Greek-language subset of 826 examples derived from the COVID-QA-el dataset, translated using machine translation. The dataset follows the SQuADv1.1 fashion style.
The original dataset, COVID-QA: A Question Answering Dataset for COVID-19 (ACL 2020) contains 2,019 question-answer pairs annotated by volunteer biomedical experts on scientific literature about COVID-19.
Data… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-QA-el-small.Iraqi-Arabic-multidomain-QA-text
Iraqi Arabic Multidomain QA Dataset
The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models.
This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.alpaca_pangasinan
🇵🇭 Pangasinan Alpaca Dataset
Dataset Summary
This dataset is a Pangasinan translation of the original Alpaca instruction-following dataset. It is designed to support research and development of instruction-tuned language models for low-resource Philippine languages, particularly Pangasinan.
The dataset retains the original Alpaca structure while providing high-quality translations of instructions, inputs, and outputs.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/PLTAT/alpaca_pangasinan.PanDomain-V1PanDomain-V1 is a high-quality, fully English dataset designed for training generalist language models across all major domains. It serves as the foundational training corpus for the Talon model family, built to support broad capabilities in both reasoning and generation.
Every model sees everything.
Dataset Map Using Nomic
Click here to view the map
Warning: This dataset is contaminated with unloaded instructions. Some rows contain just "-" in the instruction field, while other… See the full description on the dataset page: https://huggingface.co/datasets/talon-community/PanDomain-V1.Multi-Turn-Insurance-Underwriting
Dataset Card for Multi-Turn-Insurance-Underwriting
Dataset Summary
This dataset includes sample traces and associated metadata from multi-turn interactions between a commercial underwriter and AI assistant. We built the system in langgraph with model context protocol and ReAct agents. In each sample, the underwriter has a specific task to solve related to a recent application for insurance by a small business. We created a diverse sample dataset covering 6 distinct types… See the full description on the dataset page: https://huggingface.co/datasets/dennis-panos/Multi-Turn-Insurance-Underwriting.COVID-19_qa_pairs
Dataset Card for COVID-19_qa_pairs dataset
Dataset Summary
This datasets includes 604 question-answer pairs related to COVID-19 pandemic machine translated in Greek language.
The data is extracted from the official website of WHO.
Data Fields
question: Query question
document: Answer to the question
Bias, Risks, and Limitations
This dataset is the result of machine translation.
Licensing Information
The dataset is licensed under the… See the full description on the dataset page: https://huggingface.co/datasets/panosgriz/COVID-19_qa_pairs.rcqa-system-XiaoHong-v1-golden-evalSets
紅樓夢相關古漢語知識問答系統-黃金驗證集
本資料集為「基於 RA-LLMs 架構之古漢語知識問答系統」的整體系統驗證集,包含 60 題以《紅樓夢》為核心的學術題目,依 Bloom 認知分類法設計,用於 RAG + LM-ft 系統之自動化評估管線。
皆為AI(LLM)生成。
pangasinan
🇵🇭 Pangasinan Alpaca Dataset
Dataset Summary
This dataset is a Pangasinan translation of the original Alpaca instruction-following dataset. It is designed to support research and development of instruction-tuned language models for low-resource Philippine languages, particularly Pangasinan.
The dataset retains the original Alpaca structure while providing high-quality translations of instructions, inputs, and outputs.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/leklek02/pangasinan.alpaca_pangasinan
🇵🇭 Pangasinan Alpaca Dataset
Dataset Summary
This dataset is a Pangasinan translation of the original Alpaca instruction-following dataset. It is designed to support research and development of instruction-tuned language models for low-resource Philippine languages, particularly Pangasinan.
The dataset retains the original Alpaca structure while providing high-quality translations of instructions, inputs, and outputs.
Dataset Structure
Data… See the full description on the dataset page: https://huggingface.co/datasets/leklek02/alpaca_pangasinan.
