datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron-ferc-pst
Enron FERC email corpus in native PST
The EDRM Enron v2 email corpus in Microsoft PST format, modified to reduce personal privacy risk. Mailbox structure, MAPI metadata, message bodies, and retained attachments are preserved.
The release contains 171 PST files in data/, with one or more files per custodian.
Count
Version
v1
Messages
1,226,178
Attachments
453,832
PST files
171
Possible uses include email research, e-discovery testing, information retrieval… See the full description on the dataset page: https://huggingface.co/datasets/intellekthq/enron-ferc-pst.pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.cc100-latin
Latin part of cc100 corpus
This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface.
Preprocessing
I undertook the following preprocessing steps:
Removal of all "pseudo-Latin" text ("Lorem ipsum ...").
Use of CLTK for sentence splitting and normalisation.
Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.VSMRC-mrc-ABCD
VSMRC/mrc (bản tách cột A/B/C/D)
Dataset này là gì
Đây là bản định dạng lại (reformatted / derived) của dataset gốc VSMRC/mrc — phần multiple-choice reading comprehension trong bộ VSMRC (Vietnamese Text Segmentation and Multiple-Choice Reading Comprehension Dataset), do nhóm tác giả tại Đại học
Công nghệ, ĐHQGHN công bố.
Dataset gốc đã có sẵn cột choices (list Python) và correctchoice (số nguyên 0-3) — bản này chỉ map lại thành các cột A, B, C, D, answer cho khớp… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/VSMRC-mrc-ABCD.pst-audioViCS-MCQ
ViCS-MCQ: Vietnamese Computer Science Multiple-Choice QA Dataset
ViCS-MCQ (Vietnamese Computer Science Multiple-Choice Questions) là bộ dữ liệu câu hỏi trắc nghiệm tiếng Việt phục vụ nghiên cứu xử lý ngôn ngữ tự nhiên (NLP), đánh giá năng lực suy luận và đọc hiểu của các mô hình ngôn ngữ lớn (LLM Benchmark), cũng như xây dựng các hệ thống hỗ trợ học tập trong lĩnh vực Khoa học Máy tính & Công nghệ Thông tin.
Bộ dữ liệu tập trung vào 2 môn học nền tảng:
Hệ điều hành (Operating… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/ViCS-MCQ.pstuts_rag_qa
📊 PsTuts-RAG Q&A Dataset
This dataset contains question-answer pairs generated using RAGAS
from Photoshop tutorial video transcripts published in PsTuts-VQA Dataset.
It's designed for training and evaluating RAG (Retrieval-Augmented Generation) systems focused on Photoshop tutorials.
📝 Dataset Description
Dataset Summary
The dataset contains 100 question-answer pairs related to Photoshop usage, generated from video transcripts using RAGAS's… See the full description on the dataset page: https://huggingface.co/datasets/mbudisic/pstuts_rag_qa.estonian-blimp-ind-pst-3sg-to-1sg-experimentalestonian-blimp-ind-pst-3pl-to-1pl-experimentalukraine-liveblog
Dataset Card
Dataset Summary
The "ukraine-liveblog" dataset contains a collection of news articles published on the liveblog of the popular German news website, tagesschau.de. The dataset covers the period from February 2022 to February 2023, and includes every news feed published during this time that covers the ongoing war in Ukraine.
Supported Tasks and Leaderboards
--
Languages
The language of the dataset is German.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pstuerner/ukraine-liveblog.estonian-blimp-ind-pst-3pl-to-2pl-experimentalestonian-blimp-ind-pst-3sg-to-2sg-experimental
