datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.PsTuts-VQA
PsTuts-VQA-Dataset
Note: This is a mirror of Adobe Research dataset hosted at https://github.com/adobe-research/PsTuts-VQA-Dataset
and released under a permissive CC license.
https://sites.google.com/view/pstuts-vqa/home
PsTuts-VQA is a video question answering dataset on the narrated instructional videos for an image editing software. All the videos and their manual transcripts in English were obtained from the official website of the software. To collect a high-quality question… See the full description on the dataset page: https://huggingface.co/datasets/mbudisic/PsTuts-VQA.pstuts_rag_qa
📊 PsTuts-RAG Q&A Dataset
This dataset contains question-answer pairs generated using RAGAS
from Photoshop tutorial video transcripts published in PsTuts-VQA Dataset.
It's designed for training and evaluating RAG (Retrieval-Augmented Generation) systems focused on Photoshop tutorials.
📝 Dataset Description
Dataset Summary
The dataset contains 100 question-answer pairs related to Photoshop usage, generated from video transcripts using RAGAS's… See the full description on the dataset page: https://huggingface.co/datasets/mbudisic/pstuts_rag_qa.ukraine-liveblog
Dataset Card
Dataset Summary
The "ukraine-liveblog" dataset contains a collection of news articles published on the liveblog of the popular German news website, tagesschau.de. The dataset covers the period from February 2022 to February 2023, and includes every news feed published during this time that covers the ongoing war in Ukraine.
Supported Tasks and Leaderboards
--
Languages
The language of the dataset is German.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pstuerner/ukraine-liveblog.
