datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
browsesafe-bench
Dataset Card for BrowseSafe-Bench
Dataset Details
Dataset Description
BrowseSafe-Bench is a comprehensive security benchmark designed to evaluate the robustness of AI browser agents against prompt injection attacks embedded in realistic HTML environments. Unlike prior benchmarks that focus on simple text injections, BrowseSafe-Bench emphasizes environmental realism, incorporating complex HTML structures, diverse attack semantics, and benign "distractor"… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/browsesafe-bench.olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters
Dataset Card for "olm-CC-MAIN-2022-40-sampling-ratio-0.15894621295-perplexity-filters"
More Information needed
PII-TRACE
PII-TRACE
PII-TRACE is a synthetic dataset for privacy-focused named entity recognition (NER) of personally identifiable information (PII), released as a 500-conversation subset of multi-turn dialogues with exact span annotations.
Dataset summary
This release contains 500 conversations with 2,653 annotated PII spans across nine labels. Among them, 450 conversations contain PII spans and 50 contain none.
Data format
Each record contains:
Field… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/PII-TRACE.jailbreaks_dataset_with_perplexity_bigcode_starcoder2-3b_bigcode_starcoder2-7baugmented_images_perplexity
Dataset Card for "augmented_images_perplexity"
More Information needed
sinhala-perplexity-test-dataset
Sinhala Perplexity Test Dataset
Dataset Description
This dataset contains 500 sentence pairs in Sinhala, provided in two scripts: Unicode Sinhala and Romanized Sinhala. It is intended for evaluating perplexity and benchmarking language models on Sinhala text.
This dataset was introduced and used in the following benchmark study:
Rajapakse, M., & Weerasinghe, R. (2025). Script Sensitivity: Benchmarking Language Models on Unicode, Romanized and Mixed-Script Sinhala.… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-perplexity-test-dataset.olm-october-2022-tokenized-1024-perplexity-filters
Dataset Card for "olm-october-2022-tokenized-1024-perplexity-filters"
More Information needed
hc3-and-gpt-wiki-intro-with-perplexity-and-32-windowchat_perplexity_scoredlanguage-perplexitydyck-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-1-OP-False-train-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-False-train-perplexityhc3-and-gpt-wiki-intro-with-perplexity
Dataset Card for "hc3-and-gpt-wiki-intro-with-perplexity"
More Information needed
hc2-wiki-perplexity-stride-32-maxlen-128wikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP_-wikipedia-dpr-k-2-OP-True-train-perplexitywikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-True-train-perplexityhc3-and-gpt-wiki-intro-with-perplexity-and-128-windowexperiment_perplexity_instruction_llama3_8b_responsewikitext-2-raw-v1-preprocessed-200-PI_KFI-FK_claude-train-PI_KFI_IK-perplexityReasoningSet-Perplexity-Distill-Llama70b
Development Process
question dataset from facebook/natural_reasoning
We used perplexity-ai/r1-1776-distill-llama-70b
License
facebook/natural_reasoning : https://spdx.org/licenses/CC-BY-NC-4.0
perplexity-ai/r1-1776-distill-llama-70b : https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/mit.md
Acknowledgement
This research is supported by TPU Research Cloud program.
pippa_perplexityjust a pippa dataset I can have in hand to replace wikitext
wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-False-train-perplexitywikitext-2-raw-v1-preprocessed-200-PI_KFI-FK_-train-perplexitywikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-False-train-PI_KFI-perplexitywikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-2-OP-False-train-perplexitytest-dpo-perplexity
test-dpo-perplexity
This DPO (Direct Preference Optimization) dataset was generated using watermarked text generation
with red/blue token parity sampling. Each prompt has both a red and blue answer for creating
preference pairs.
Watermark Configuration
Sampler Type: soft_watermark
Soft Mode: True
Soft Threshold: 0.95
Sampling Parameters
Temperature: 0.7
Top-p: 0.8
Top-k: 20
Max Tokens: 64
Statistics
Total Prompts: 3
Avg Red Parity Ratio: 0.5521… See the full description on the dataset page: https://huggingface.co/datasets/eac123/test-dpo-perplexity.ko-perplexity-corpus
Lumia101/ko-perplexity-corpus
This dataset was created to measure the perplexity of an LLM trained on a Korean dataset.
Dataset Source
HAERAE-HUB/KOREAN-WEBTEXT
maxidl/FineNews-unfiltered
wikimedia/wikipedia
wikitext-2-raw-v1-preprocessed-200-PI_KFI_-wikipedia-dpr-k-1-OP-True-train-PI_KFI-perplexitywikitext-2-raw-v1-preprocessed-200-PP_RIP_PTP-train-perplexity
