datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
draco
DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
The DRACO Benchmark consists of complex, open-ended research tasks with expert-curated rubrics for evaluating deep research systems. Tasks span 10 domains and require drawing on information sources from 40 countries. Each task is paired with a detailed, task-specific rubric featuring an average of ~40 evaluation criteria across four axes: factual accuracy, breadth and depth of analysis… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/draco.browsesafe-bench
Dataset Card for BrowseSafe-Bench
Dataset Details
Dataset Description
BrowseSafe-Bench is a comprehensive security benchmark designed to evaluate the robustness of AI browser agents against prompt injection attacks embedded in realistic HTML environments. Unlike prior benchmarks that focus on simple text injections, BrowseSafe-Bench emphasizes environmental realism, incorporating complex HTML structures, diverse attack semantics, and benign "distractor"… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/browsesafe-bench.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters"
More Information needed
languagewise_perplexity_aya_datasetwandr
WANDR
Overview and provenance
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, structured,
high-volume web research tasks. This dataset is a task-and-verification corpus,
not a question/answer collection: it contains no solver outputs or reference
answer sets. WANDR evaluation refetches cited pages and judges submitted records
against task-specific, reference-free specifications.
See the paper, the
blog,
and the evaluation repository.… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/wandr.PII-TRACE
PII-TRACE
PII-TRACE is a synthetic dataset for privacy-focused named entity recognition (NER) of personally identifiable information (PII), released as a 500-conversation subset of multi-turn dialogues with exact span annotations.
Dataset summary
This release contains 500 conversations with 2,653 annotated PII spans across nine labels. Among them, 450 conversations contain PII spans and 50 contain none.
Data format
Each record contains:
Field… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/PII-TRACE.perplexity_analysis
Perplexity Analysis
This repository contains the data, scripts, and generated figures used for
perplexity analysis experiments.
Contents
data/Qwen3: rollout data for Qwen3 1.7B and 4B Base, GRPO, and MaxRL
models on AIME25 and BeyondAIME.
data/Maze/perplexity: maze rollout data and derived perplexity analysis
artifacts.
outputs: generated JSON summaries and figures for Qwen3 analyses.
*.py: analysis and plotting scripts.
See data/README.md for additional data details… See the full description on the dataset page: https://huggingface.co/datasets/max-rl/perplexity_analysis.jailbreaks_dataset_with_perplexity_bigcode_starcoder2-3b_bigcode_starcoder2-7bChatGPT-Gemini-Claude-Perplexity-Human-Evaluation-Multi-Aspects-Review-Dataset
ChatGPT Gemini Claude Perplexity Human Evaluation Multi Aspect Review Dataset
Introduction
Human evaluation and reviews with scalar score of AI Services responses are very usefuly in LLM Finetuning, Human Preference Alignment, Few-Shot Learning, Bad Case Shooting, etc, but extremely difficult to collect.
This dataset is collected from DeepNLP AI Service User Review panel (http://www.deepnlp.org/store), which is an open review website for users to give reviews and upload… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/ChatGPT-Gemini-Claude-Perplexity-Human-Evaluation-Multi-Aspects-Review-Dataset.augmented_images_perplexity
Dataset Card for "augmented_images_perplexity"
More Information needed
sinhala-perplexity-test-dataset
Sinhala Perplexity Test Dataset
Dataset Description
This dataset contains 500 sentence pairs in Sinhala, provided in two scripts: Unicode Sinhala and Romanized Sinhala. It is intended for evaluating perplexity and benchmarking language models on Sinhala text.
This dataset was introduced and used in the following benchmark study:
Rajapakse, M., & Weerasinghe, R. (2025). Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala.… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-perplexity-test-dataset.olm-october-2022-tokenized-1024-perplexity-filters
Dataset Card for "olm-october-2022-tokenized-1024-perplexity-filters"
More Information needed
hc3-and-gpt-wiki-intro-with-perplexity-and-32-windowchat_perplexity_scoredperplexity_evaluation
SaulLM-7B: Pioneering the first Legal Large Language Model
Perplexity Analysis
This dataset presents the data used in the paper "SaulLM-7B: Pioneering the first Legal Large Language Model" in "6.3 Perplexity Analysis" section.
The dataset contains the perplexity scores of SaulLM-7B, Llama2-7B and Mistral-7B across a corpora of recent text.
Cleaning
We proceeded to standardize the data by removing any special characters using unicodedata normalization.
We also… See the full description on the dataset page: https://huggingface.co/datasets/Equall/perplexity_evaluation.language-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-1-OP-False-train-perplexitydyck-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-False-train-perplexityhc3-and-gpt-wiki-intro-with-perplexity
Dataset Card for "hc3-and-gpt-wiki-intro-with-perplexity"
More Information needed
hc2-wiki-perplexity-stride-32-maxlen-128wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-True-train-perplexityopenhermes2.5-Perplexity_filtered_top30
OpenHermes 2.5 - Perlexity Filtered (Top 30%)
A filtered subset of OpenHermes 2.5
containing the top 30% highest perplexity samples scored by Qwen2.5-3B-Instruct
Dataset Summary
Source teknium/OpenHermes-2.5
Size 300466 samples (from ~1M original)
Filter method Cross-entropy loss scored by Qwen2.5-3B-Instruct (4-bit NF4)
Kept samples above the 70th percentile loss threshold
Why this dataset?
High perplexity samples are the examples a model finds… See the full description on the dataset page: https://huggingface.co/datasets/Osye/openhermes2.5-Perplexity_filtered_top30.pippa_perplexityjust a pippa dataset I can have in hand to replace wikitext
wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-True-train-perplexityperplexity__aya_dataset__trainpseudo-perplexityperplexity__aya_dataset__trainhc3-and-gpt-wiki-intro-with-perplexity-and-128-windowexperiment_perplexity_instruction_llama3_8b_response
