datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.VietDBMS-MCQGen
Dataset: Vietnamese MCQ / DPO
Description
Dataset dùng để huấn luyện và đánh giá mô hình sinh câu hỏi trắc nghiệm tiếng Việt.
Splits
train: dữ liệu DPO từ nhiều nguồn (HĐH, CSDL)
test_hdh: tập đánh giá Hệ điều hành
test_csdl: tập đánh giá Cơ sở dữ liệu
ukraine-liveblog
Dataset Card
Dataset Summary
The "ukraine-liveblog" dataset contains a collection of news articles published on the liveblog of the popular German news website, tagesschau.de. The dataset covers the period from February 2022 to February 2023, and includes every news feed published during this time that covers the ongoing war in Ukraine.
Supported Tasks and Leaderboards
--
Languages
The language of the dataset is German.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pstuerner/ukraine-liveblog.
