datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pandabench
PandaBench
Paper | Project Page | Code
PandaBench (and PandaSet) is an image distortion benchmark designed for evaluating perceptual comparison and distortion-aware visual reasoning. It introduces the task of learning a Distortion Graph (DG), representing dense degradation information such as distortion type, severity, and quality scores in a compact, interpretable graph structure grounded in image regions.
PandaSet (the train/val folders) is used to train the model (Panda), while… See the full description on the dataset page: https://huggingface.co/datasets/kjanjua26/pandabench.panda-bench
PandaBench
PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies.
The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges.
Dataset Description
This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page: https://huggingface.co/datasets/Beijing-AISI/panda-bench.BRIDGE
BRIDGE — Qwen Image Edit Dataset
Part of the dataset used in BRIDGE: Background Routing and Isolated Discrete Gating for Coarse-Mask Local Editing.
FLUX subject-condition extension (2026-09-19)
The same BRIDGE method is trained with LoRA on Qwen and full-transformer
fine-tuning on FLUX. The FLUX dataset adds a generated subject-reference image
condition. See the extension documentation.
Exact FLUX split: 27,834 training rows, 3,092 test rows.
30,926 selected… See the full description on the dataset page: https://huggingface.co/datasets/PANDATREE/BRIDGE.Machine_Mindset_MBTI_datasetHere are the behavior datasets used for supervised fine-tuning (SFT). And they can also be used for direct preference optimization (DPO).
The exact copy can also be found in Github.
Prefix 'en' denotes the datasets of the English version.
Prefix 'zh' denotes the datasets of the Chinese version.
Dataset introduction
There are four dimension in MBTI. And there are two opposite attributes within each dimension.
To be specific:
Energe: Extraversion (E) - Introversion (I)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/Machine_Mindset_MBTI_dataset.starcoder-numpy-pandasxenium_pandavid_datasetpanda-70m
Panda 70M dataset by Snap Inc
70M video-caption pairs
Code for downloading: https://github.com/snap-research/Panda-70M/dataset_dataloading
xenium_5k_pandavid_dataset_v2xenium_pandavid_dataset4PANEL_IMAGESpanda70m_wdsPandasPlotBench
PandasPlotBench
PandasPlotBench is a benchmark to assess the capability of models in writing the code for visualizations given the description of the Pandas DataFrame.
🛠️ Task. Given the plotting task and the description of a Pandas DataFrame, write the code to build a plot.
The dataset is based on the MatPlotLib gallery.
The paper can be found in arXiv: https://arxiv.org/abs/2412.02764v1.
To score your model on this dataset, you can use the our GitHub repository.
📩 If you have… See the full description on the dataset page: https://huggingface.co/datasets/JetBrains-Research/PandasPlotBench.lm-eval-results-ichigoberry-pandafish-dt-7b-private
Dataset Card for Evaluation run of ichigoberry/pandafish-dt-7b
Dataset automatically created during the evaluation run of model ichigoberry/pandafish-dt-7b
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-ichigoberry-pandafish-dt-7b-private.pandas-questionspanda-70mpandaschinese_law_examples
1000 examples of law items
law_item.jsonl contains 1000 samples of current and effective Chinese laws. e.g.
{"title": "《中华人民共和国劳动合同法(2012修正)》",
"classification": "类别 : 劳动合同营商环境优化 ",
"num": "第十九条",
"contents": "第十九条【试用期】劳动合同期限三个月以上不满一年的,试用期不得超过一个月;劳动合同期限一年以上不满三年的,试用期不得超过二个月;三年以上固定期限和无固定期限的劳动合同,试用期不得超过六个月。同一用人单位与同一劳动者只能约定一次试用期。以完成一定工作任务为期限的劳动合同或者劳动合同期限不满三个月的,不得约定试用期。试用期包含在劳动合同期限内。劳动合同仅约定试用期的,试用期不成立,该期限为劳动合同期限。"}
Using BGE Embedding to compute similarity between query and… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_law_examples.tangram-house-panda-330epThis dataset was created using LeRobot.
Dataset Description
Demonstrations for Tangram-Bench: a Panda arm in MuJoCo assembles a tangram silhouette from a packed square of seven pieces, given the silhouette outline on the table and a one-sentence prompt.
Episodes
330 (330 solved)
Distinct scenes
72; the designed grid holds 72 per figure, so seeds n and n + 72 share a scene
Frames
691,855 at 10 fps, 1153 min of manipulation
Figures
house
Robot
panda… See the full description on the dataset page: https://huggingface.co/datasets/murobotics/tangram-house-panda-330ep.tangram-square-rectangle-house-panda-80epThis dataset was created using LeRobot.
Dataset Description
Demonstrations for Tangram-Bench: a Panda arm in MuJoCo assembles a tangram silhouette from a packed square of seven pieces, given the silhouette outline on the table and a one-sentence prompt.
Episodes
80 (80 solved)
Frames
167,212 at 10 fps, 279 min of manipulation
Figures
square, rectangle, house
Robot
panda, simulated (MuJoCo)
Cameras
observation.images.top, observation.images.wrist… See the full description on the dataset page: https://huggingface.co/datasets/murobotics/tangram-square-rectangle-house-panda-80ep.panda-ordered-sweepThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "panda",
"total_episodes": 90,
"total_frames": 40148,
"total_tasks": 3,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:90"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rdoshi21/panda-ordered-sweep.panda
Dataset Card for PANDA
Dataset Summary
PANDA (Perturbation Augmentation NLP DAtaset) consists of approximately 100K pairs of crowdsourced human-perturbed text snippets (original, perturbed). Annotators were given selected terms and target demographic attributes, and instructed to rewrite text snippets along three demographic axes: gender, race and age, while preserving semantic meaning. Text snippets were sourced from a range of text corpora (BookCorpus, Wikipedia, ANLI… See the full description on the dataset page: https://huggingface.co/datasets/facebook/panda.dataset-panda72M-repro
Dataset Card for dataset-panda72M-repro
These are all the datasets we used for training our scaled-up Panda-72M model (https://huggingface.co/GilpinLab/panda-72M) of Panda: Patched Attention for Nonlinear Dynamics.
We have made all parquet files here streamable by using chunked row groups, instead of the single row group containing all data (as done for our other Panda datasets on HF).
NOTE (on historical changelog):
base_mixedp_ic16
The dataset in… See the full description on the dataset page: https://huggingface.co/datasets/GilpinLab/dataset-panda72M-repro.Panda-70MPANDA-PLUS-Bench
PANDA-PLUS-Bench
A benchmark dataset for evaluating WSI-specific feature collapse in pathology foundation models.
Dataset Description
PANDA-PLUS-Bench contains expert-annotated prostate biopsy patches from 9 whole slide images (9 unique patients) with pixel-level Gleason pattern annotations.
Dataset Summary
Patches: ~2,770 per augmentation condition
Resolution: 224×224 pixels at 20× magnification
Classes: Benign (0), GP3 (1), GP4 (2), GP5 (3)
Slides: 9 (one… See the full description on the dataset page: https://huggingface.co/datasets/dellacorte/PANDA-PLUS-Bench.push_panda_toilet_paper_s2_postprocessed
push_panda_toilet_paper_s2_postprocessed
Post-processed LeRobot dataset for the task: push the panda and the toilet paper.
The repository contains the full LeRobot dataset under the standard meta/, data/, and videos/ directories. A lightweight preview/ subset is configured as the default Hugging Face Dataset Viewer view so the dataset page shows representative videos directly.
Dataset Summary
Task: push the panda and the toilet paper
Format: LeRobot v2.1
Episodes: 32… See the full description on the dataset page: https://huggingface.co/datasets/NONHUMAN-RESEARCH/push_panda_toilet_paper_s2_postprocessed.megazeka-tr-spellfix-pairs
Megazeka · Turkish Spelling Correction Pairs
40,313 (misspelled → correct) Turkish sentence pairs with a full record of every corruption that
was applied, generated from the CC0 Common Voice Turkish Sentence Collector and from project-authored
everyday first/second-person sentences.
The distinguishing feature is the operations field: each pair carries the exact sequence of noise
transformations that produced it, with before/after text at each step. That makes it possible to… See the full description on the dataset page: https://huggingface.co/datasets/pandakingpunc/megazeka-tr-spellfix-pairs.tangram-square-rectangle-house-panda-300epThis dataset was created using LeRobot.
Dataset Description
Demonstrations for Tangram-Bench: a Panda arm in MuJoCo assembles a tangram silhouette from a packed square of seven pieces, given the silhouette outline on the table and a one-sentence prompt.
Episodes
300 (300 solved)
Distinct scenes
300; the designed grid holds 7,776 per figure, seed n + 7776 repeating seed n
Frames
616,513 at 10 fps, 1028 min of manipulation
Figures
square, rectangle, house… See the full description on the dataset page: https://huggingface.co/datasets/Autobrik/tangram-square-rectangle-house-panda-300ep.stance_process_pandas_df
Dataset Card for "stance_process_pandas_df"
More Information needed
pandas-create-context
Overview
This dataset is built from sql-create-context, which in itself builds from WikiSQL and Spider.
I have used GPT4 to translate the SQL schema into pandas DataFrame schem initialization statements and to translate the SQL queries into pandas queries.
There are 862 examples of natural language queries, pandas DataFrame creation statements, and pandas query answering the question using the DataFrame creation statement as context. This dataset was built with text-to-pandas… See the full description on the dataset page: https://huggingface.co/datasets/hiltch/pandas-create-context.chinese_verdict_examples
verdicts examples
verdicts_200.jsonl contains 200 examples of verdicts from Chinese Judgements Online, we process the datasets for semantic retrieval
using BGE to compute similarity between query and verdict
from FlagEmbedding import FlagModel
from datasets import load_dataset
dataset = load_dataset("FarReelAILab/verdicts")
model = FlagModel('BAAI/bge-large-zh-v1.5',
query_instruction_for_retrieval="为这个句子生成表示以用于检索相关文章:",
use_fp16=True)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/chinese_verdict_examples.
