pst
Datasets
All datasets matching “pst”enron-ferc-pst
Enron FERC email corpus in native PST
The EDRM Enron v2 email corpus in Microsoft PST format, modified to reduce personal privacy risk. Mailbox structure, MAPI metadata, message bodies, and retained attachments are preserved.
The release contains 171 PST files in data/, with one or more files per custodian.
Count
Version
v1
Messages
1,226,178
Attachments
453,832
PST files
171
Possible uses include email research, e-discovery testing, information retrieval… See the full description on the dataset page: https://huggingface.co/datasets/intellekthq/enron-ferc-pst.pstu-synthetic-secrets
PSTU Synthetic Secrets Dataset
Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper:
Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal
Hoda Fakhar — ECML PKDD 2026
Dataset Description
175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric.
All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.PST50
PST50: Benchmark for Photorealistic Style Transfer
PST50 is the first benchmark dataset designed for rigorous evaluation of Photorealistic Style Transfer (PST). It includes high-resolution, professionally curated content-style pairs with ground truth stylizations, suitable for both paired and unpaired evaluation protocols.
This dataset is introduced in the paper:SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
Project Page • Paper • Code
📁… See the full description on the dataset page: https://huggingface.co/datasets/KlyaT/PST50.pst-demo
PST MugBox-v2 Teleoperation Demonstrations
UR5e + Inspire Hand(右手・6 能動関節)による MugBox-v2 タスクの遠隔操作デモ集です。
Leap Motion で捉えた人の手を Inspire Hand へリターゲティングし、手首姿勢は Pinocchio CLIK
(IK)で UR5e に追従させて ManiSkill / SAPIEN 上で収集しました。学習用の成功試行(success=True)に
加え、オペレータの習熟過程を解析できるよう 失敗試行も raw/ 以下に収録しています。
研究の主題は「非専門家オペレータから高品質なデモを大量に集めるにはどうすればよいか」です。
収集時にオペレータへ提示する視覚表現(ロボットハンド上の点群アノテーション)を 3 条件に分け、
専門家データと比較できるように構成されています。
コード: https://github.com/omron-sinicx/particle-skill-transfer
学習済みモデル:… See the full description on the dataset page: https://huggingface.co/datasets/yamaama/pst-demo.PST50
PST50: Benchmark for Photorealistic Style Transfer
PST50 is the first benchmark dataset designed for rigorous evaluation of Photorealistic Style Transfer (PST). It includes high-resolution, professionally curated content-style pairs with ground truth stylizations, suitable for both paired and unpaired evaluation protocols.
This dataset is introduced in the paper:SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer
Project Page • Paper • Code
📁… See the full description on the dataset page: https://huggingface.co/datasets/zrgong/PST50.cc100-latin
Latin part of cc100 corpus
This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface.
Preprocessing
I undertook the following preprocessing steps:
Removal of all "pseudo-Latin" text ("Lorem ipsum ...").
Use of CLTK for sentence splitting and normalisation.
Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.
