datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.Panda-CVL-train
Panda-CVL Training Split
Overview
Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool.
Given a question-response pair, the model first judges whether the response is acceptable.
If correction is needed, it must locate the first inappropriate token and replace it with an
appropriate one, so generation can continue from the "correct prefix + corrected token" state
and ultimately produce an acceptable response.
Compared with… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-train.Panda-CVL-test
Panda-CVL Test Split
This dataset is based on the paper onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction. Code is available at GitHub.
Overview
Panda-CVL is a token-level correction dataset and benchmark annotated with the onPanda tool.
Given a question-response pair, the model first judges whether the response is acceptable.
If correction is needed, it must locate the first inappropriate token and replace it… See the full description on the dataset page: https://huggingface.co/datasets/diyer22/Panda-CVL-test.nemotron-nano-30b-miniswe-swebench-verified
Nemotron Nano 30B + mini-swe-agent SWE-bench Verified Trajectories
Agent trajectories from running NVIDIA Nemotron 3 Nano 30B A3B (MoE, 8B active params) on SWE-bench Verified using mini-swe-agent.
⚠️ Incomplete Run
This benchmark was terminated early due to poor performance. The model struggled with the agentic coding task.
Model Information
Attribute
Value
Model
NVIDIA Nemotron 3 Nano 30B A3B
Architecture
MoE (30B total, 8B active)
Serving
vLLM… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-30b-miniswe-swebench-verified.pandora-tool-calling
Pandora Tool Calling
A tool-calling dataset for Supervised fine-tuning of the Pandora Large Language Model (LLM).
The dataset is based on the glaiveai/glaive-function-calling-v2 dataset.
Copyright and license
Copyright (c) 2024, Danilo Peixoto Ferreira. All rights reserved.
Project developed under a BSD-3-Clause license.
livesweagent-devstral2-123b-swebench-verified
Live-SWE-agent + Devstral 2 (123B) SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral 2 (123B) via OpenRouter on the SWE-bench Verified benchmark using the Live-SWE-agent framework.
⚠️ Framework Note
This run uses Live-SWE-agent, NOT standard mini-swe-agent. Live-SWE-agent is a self-evolving agent framework that encourages the model to create custom Python tools during runtime.
Key differences from mini-swe-agent:
Self-evolving… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/livesweagent-devstral2-123b-swebench-verified.nemotron-nano-swebench-verified-traj
NVIDIA Nemotron-3-Nano-30B SWE-bench Verified Trajectories
This dataset contains agent trajectories from running NVIDIA Nemotron-3-Nano-30B-A3B on the SWE-bench Verified benchmark.
Model Information
Attribute
Value
Model
nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
Architecture
Mixture of Experts (MoE)
Parameters
30B total / 3B active per token
Hardware
NVIDIA B200 (1 GPU)
Benchmark Results
Metric
Value
Total Instances
500… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/nemotron-nano-swebench-verified-traj.khmer-text-corpus
Khmer + English Mixed Text Corpus
A cleaned, deduplicated Khmer text corpus with naturally-occurring Khmer/English
code-switching (English tech & finance terms, Latin script, and digits embedded in Khmer
text). It is the training data for the
Panhapich/khmer-sp-8k SentencePiece tokenizer and the
text-only warm-start of a shared Khmer diffusion decoder.
Dataset summary
4,893,739 sentences, one per line, UTF-8, deduplicated and shuffled
(fixed seed 42… See the full description on the dataset page: https://huggingface.co/datasets/Panhapich/khmer-text-corpus.thema-panhellenic-exams
Thema
Thema is an open benchmark based on the Greek Panhellenic university entrance examinations (Πανελλαδικές Εξετάσεις ΓΕΛ). Models answer authentic exam questions in Greek. Responses are graded on the national 0 to 20 scale and can be converted to the admission points used by Greek university departments.
Website · Code · Leaderboard · Method
Dataset summary
Current release
Exam years
2023 to 2026
Published subjects
10 of 10
Complete exam… See the full description on the dataset page: https://huggingface.co/datasets/mpvasilis/thema-panhellenic-exams.panta_instruct_multi_modal_v1
Panta Instruct Multi-Modal v1
Dataset d'instructions multimodal en français : chaque exemple associe une question
(texte + parole + pictogrammes) à une réponse (texte + pictogrammes).
Colonnes
Colonne
Type
Description
audio
Audio (24 kHz, mono)
Enregistrement de la question (text_input)
text_input
string
Question / instruction
text_output
string
Réponse
pictos_input
list[string]
Identifiants des pictogrammes de la question
pictos_output… See the full description on the dataset page: https://huggingface.co/datasets/audibeal74/panta_instruct_multi_modal_v1.pander-score
Pander Score
This is the companion dataset for The Pander Score: A Continuous Measure of
Sycophancy as Epistemic Deference by Alejandro Botas, Paul de Font-Reaulx, and
Luke Hewitt.
The default benchmark config contains the saved responses and credence
judgments for all 18 published target models. Each model contributes one row
for each of the 11,172 fixed prompts. Join these rows to the prompt_attributes
config on sample_id before computing the score.
The companion code, fixed… See the full description on the dataset page: https://huggingface.co/datasets/sophronresearch/pander-score.pan-african-primary-care-benchmark
Pan-African Primary Care Benchmark (v1)
A multilingual safety and reasoning benchmark for clinical AI in African primary-care contexts
Overview
This dataset is designed to rigorously evaluate how well large language models (and clinical AI agents) perform when patients present in real-world African languages — exactly as they do in clinics across the continent.
v1 contains 300 synthetic, de-identified primary-care scenarios (50 per language × 6 files):
File
Language… See the full description on the dataset page: https://huggingface.co/datasets/nimrodzw/pan-african-primary-care-benchmark.pandora-data
Pandora Benchmark Data
Versioned, processed benchmark annotations for the
Pandora structured knowledge reasoning
framework.
Contents
Dataset
Split
Records
Artifact license
Spider-Syn
test
1,034
MIT
BIRD
dev
1,534
CC BY-SA 4.0
WikiTableQuestions
test
4,344
CC BY-SA 4.0
WikiSQL
test
15,878
BSD-3-Clause
GrailQA
evaluation set from public validation data
6,409
CC BY-SA 4.0 annotations; Freebase CC BY 2.5 BOX data
WebQSP
test
1,598
Freebase CC BY… See the full description on the dataset page: https://huggingface.co/datasets/bahuia/pandora-data.panorama
Dataset Summary
Dataset of satirical news from "Panorama", Russian "The Onion".
Dataset Format
Dataset is in JSONLines format, where "title" is the article title, and "body" are contents of the article.
pandas-create-context
Overview
This dataset is built from sql-create-context, which in itself builds from WikiSQL and Spider.
I have used GPT4 to translate the SQL schema into pandas DataFrame schem initialization statements and to translate the SQL queries into pandas queries.
There are 862 examples of natural language queries, pandas DataFrame creation statements, and pandas query answering the question using the DataFrame creation statement as context. This dataset was built with text-to-pandas… See the full description on the dataset page: https://huggingface.co/datasets/hiltch/pandas-create-context.autoscientist-toolcaller-dataset
AutoScientist Tool-Calling Dataset
A curated function-calling / tool-use dataset for the Adaption AutoScientist Challenge. Its
distinguishing feature is a large slice of hard negatives and reliability-focused cases — where the
correct behavior is not a plain tool call.
Adaptive Data quality (real): on the fixed set (c4923b7f…, graded on 1,000 of 2,440 rows
under the free-tier cap) the platform reported 7.0 → 8.1, +15.7%, grade C → B — now confirmed by a
completed, uncapped run… See the full description on the dataset page: https://huggingface.co/datasets/pandeyankit84/autoscientist-toolcaller-dataset.pandora-rlhf
Pandora RLHF
A Reinforcement Learning from Human Feedback (RLHF) dataset for Direct Preference Optimization (DPO) fine-tuning of the Pandora Large Language Model (LLM).
The dataset is based on the anthropic/hh-rlhf dataset.
Copyright and license
Copyright (c) 2024, Danilo Peixoto Ferreira. All rights reserved.
Project developed under a BSD-3-Clause license.
WizardLM_OrcaExplain tuned WizardLM dataset ~55K created using approaches from Orca Research Paper.
We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets.
This helps student models like orca_mini_13b to learn thought process from teacher model, which is ChatGPT (gpt-3.5-turbo-0301 version).
Please see how the System prompt is added before each instruction.
devstral2-123b-swebench-verified-traj
Devstral 2 (123B) SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral 2 (via OpenRouter) on the SWE-bench Verified benchmark using the mini-swe-agent framework.
Model Information
Attribute
Value
Model
Devstral 2
Parameters
123B (dense transformer)
Context Window
256K tokens
Provider
OpenRouter (mistralai/devstral-2512:free)
Specialization
Agentic coding
Framework
mini-swe-agent
Devstral 2 is a… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/devstral2-123b-swebench-verified-traj.kakugo-pan
Kakugo Panjabi dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Panjabi.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Panjabi. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-pan.devstral-24b-swebench-verified-traj
Devstral-24B SWE-bench Verified Trajectories
This dataset contains agent trajectories from running Devstral-24B on the SWE-bench Verified benchmark.
Model Information
Attribute
Value
Model
Devstral-24B
Parameters
24B dense transformer
Provider
Local vLLM
Benchmark Results
Metric
Value
Total Instances
500
Submitted
345 (69.0%)
Resolved
277 (55.4%)
Usage
from huggingface_hub import hf_hub_download… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/devstral-24b-swebench-verified-traj.newsophy-v0.1
This dataset was used to train the pansophic-1-preview model
This dataset was created using open-source, permissively licensed models. In addition to providing answers to a diverse set of questions, we leveraged multiple open-source pipelines to generate new tasks and questions, enriching the dataset's variety and complexity. The dataset includes examples that showcase tool usage, contextual understanding, and the application of system prompts.
Topics distribtuion in… See the full description on the dataset page: https://huggingface.co/datasets/pansophic/newsophy-v0.1.safekeep-paired-data
SafeKeep Paired Direction Data
400 minimally-contrastive pairs of harmful / harmless agent requests, used to extract
a refusal direction (difference-in-means) for tool-calling LLM agents.
Paper: arXiv:2607.29254
Format
One JSON object per line:
field
description
harmful
the original harmful user request
harmless
minimal rewrite of the same request with the harmful span swapped for a benign one
system
the agent system prompt, including the JSON tool… See the full description on the dataset page: https://huggingface.co/datasets/Panminghui/safekeep-paired-data.spelling-bee-pangrams
Spelling Bee-style letter sets and their pangrams
Each row is one puzzle: 7 distinct letters (letters, alphabetical) and the list
of pangrams — words using all 7 letters. Nothing else.
Puzzles come from the dwyl/english-words words_alpha.txt word list (~370k
entries). Included: every 7-letter combination with at least one pangram and at
least 10 valid answers (words of 4+ letters using only the 7 letters).
Source: https://github.com/dwyl/english-words (words_alpha.txt)
orca_minis_uncensored_datasetUncensored explain tuned WizardLM + Alpaca + Dolly V-2 datasets ~104K created using approaches from Orca Research Paper.
We leverage all of the 15 system instructions provided in Orca Research Paper. to generate custom datasets, in contrast to vanilla instruction tuning approaches used by original datasets.
This helps student models like orca_mini_v2_7b to learn thought process from teacher model, which is ChatGPT (gpt-3.5-turbo-0301 version).
Please see how the System prompt is added before… See the full description on the dataset page: https://huggingface.co/datasets/pankajmathur/orca_minis_uncensored_dataset.arm-asmtharu-chat
Tharu Chat Dataset
Overview
The Tharu Chat is a conversational dataset designed to support the development of Large Language Models (LLMs) for the Tharu language, a low-resource Indo-Aryan language spoken across Nepal and India.
This dataset is designed for Supervised Fine-Tuning (SFT) and addresses the digital exclusion of low-resource languages, particularly Tharu, by enabling the development of conversational AI systems grounded in real linguistic usage.
This work is… See the full description on the dataset page: https://huggingface.co/datasets/prajwal-panth/tharu-chat.Panther-dataset_v1
Dataset Details
This dataset is a modified version of Anthropic/hh-rlhf
This dataset is used in fine tuning Panther - an state of the art LLM funtuned on llama-7b pretrained model.
A very small portion i.e. 5.3% of prompts and responses were taken from this dataset to finetune and train Panther
Dataset Details
Dataset Structure
Train
Train rows : 377k
Validation
Validation rows : 20.3k
Dataset Format
input… See the full description on the dataset page: https://huggingface.co/datasets/Rardilit/Panther-dataset_v1.Iraqi-Arabic-multidomain-QA-text
Iraqi Arabic Multidomain QA Dataset
The Iraqi Arabic Multidomain QA Dataset is a curated conversational Arabic dataset designed for training, fine-tuning, benchmarking, and evaluating Large Language Models (LLMs), conversational AI systems, multilingual NLP pipelines, question answering systems, Arabic chatbots, retrieval-augmented generation (RAG), and instruction-tuned AI models.
This dataset focuses specifically on Iraqi Arabic dialectal content, one of the most… See the full description on the dataset page: https://huggingface.co/datasets/Pangeanic/Iraqi-Arabic-multidomain-QA-text.messages
Human Messages Corpus
4,594,008 individual human-written messages. No bot messages or URLs.
File
all_messages.parquet
Column
Column
Description
text
The message content (whitespace cleaned)
Statistics
Total messages: 4,594,008
Average words per message: 5.9
Total words: ~27.2 million
Source
Merged from previous versions of the dataset. Extracted from raw messages, deduplicated, and filtered to retain… See the full description on the dataset page: https://huggingface.co/datasets/Pankaj8922/messages.
