CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01intellekthq /enron-ferc-pst Enron FERC email corpus in native PST The EDRM Enron v2 email corpus in Microsoft PST format, modified to reduce personal privacy risk. Mailbox structure, MAPI metadata, message bodies, and retained attachments are preserved. The release contains 171 PST files in data/, with one or more files per custodian. Count Version v1 Messages 1,226,178 Attachments 453,832 PST files 171 Possible uses include email research, e-discovery testing, information retrieval… See the full description on the dataset page: https://huggingface.co/datasets/intellekthq/enron-ferc-pst.text1M<n<10M1 likes631 downloads2mo agoHugging Face02Hodfa71 /pstu-synthetic-secrets PSTU Synthetic Secrets Dataset Synthetic secrets benchmark for evaluating LLM memorization and unlearning, from the paper: Not All Secrets Are Equal: Type-Aware Unlearning for Language Model Secret Removal Hoda Fakhar — ECML PKDD 2026 Dataset Description 175 synthetic secrets across 25 types, each paired with 100 structurally similar decoys for computing the Carlini exposure metric. All data is synthetically generated. No real credentials, PII, or sensitive information… See the full description on the dataset page: https://huggingface.co/datasets/Hodfa71/pstu-synthetic-secrets.texttext-generationn<1K0 likes201 downloads6mo agoHugging Face03pstroe /cc100-latin Latin part of cc100 corpus This dataset contains parts of the Latin part of the cc100 dataset. It was used to train a RoBERTa-based LM model with huggingface. Preprocessing I undertook the following preprocessing steps: Removal of all "pseudo-Latin" text ("Lorem ipsum ..."). Use of CLTK for sentence splitting and normalisation. Retaining only lines containing letters of the Latin alphabet, numerals, and certain punctuation (--> grep -P '^[A-z0-9ÄÖÜäöüÆæŒœᵫĀāūōŌ.,;:?!\-… See the full description on the dataset page: https://huggingface.co/datasets/pstroe/cc100-latin.textn<1K9 likes150 downloads4y agoHugging Face04p-storm /VSMRC-mrc-ABCD VSMRC/mrc (bản tách cột A/B/C/D) Dataset này là gì Đây là bản định dạng lại (reformatted / derived) của dataset gốc VSMRC/mrc — phần multiple-choice reading comprehension trong bộ VSMRC (Vietnamese Text Segmentation and Multiple-Choice Reading Comprehension Dataset), do nhóm tác giả tại Đại học Công nghệ, ĐHQGHN công bố. Dataset gốc đã có sẵn cột choices (list Python) và correctchoice (số nguyên 0-3) — bản này chỉ map lại thành các cột A, B, C, D, answer cho khớp… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/VSMRC-mrc-ABCD.tabularquestion-answering10K<n<100K0 likes68 downloads1mo agoHugging Face05pkd /pst-audioaudio10K<n<100K0 likes66 downloads2y agoHugging Face06p-storm /ViCS-MCQ ViCS-MCQ: Vietnamese Computer Science Multiple-Choice QA Dataset ViCS-MCQ (Vietnamese Computer Science Multiple-Choice Questions) là bộ dữ liệu câu hỏi trắc nghiệm tiếng Việt phục vụ nghiên cứu xử lý ngôn ngữ tự nhiên (NLP), đánh giá năng lực suy luận và đọc hiểu của các mô hình ngôn ngữ lớn (LLM Benchmark), cũng như xây dựng các hệ thống hỗ trợ học tập trong lĩnh vực Khoa học Máy tính & Công nghệ Thông tin. Bộ dữ liệu tập trung vào 2 môn học nền tảng: Hệ điều hành (Operating… See the full description on the dataset page: https://huggingface.co/datasets/p-storm/ViCS-MCQ.textmultiple-choice1K<n<10K0 likes52 downloads1mo agoHugging Face07mbudisic /pstuts_rag_qa 📊 PsTuts-RAG Q&A Dataset This dataset contains question-answer pairs generated using RAGAS from Photoshop tutorial video transcripts published in PsTuts-VQA Dataset. It's designed for training and evaluating RAG (Retrieval-Augmented Generation) systems focused on Photoshop tutorials. 📝 Dataset Description Dataset Summary The dataset contains 100 question-answer pairs related to Photoshop usage, generated from video transcripts using RAGAS's… See the full description on the dataset page: https://huggingface.co/datasets/mbudisic/pstuts_rag_qa.textquestion-answeringn<1K0 likes11 downloads1y agoHugging Face08liu-nlp /estonian-blimp-ind-pst-3sg-to-1sg-experimentaltextn<1K0 likes11 downloads1y agoHugging Face09liu-nlp /estonian-blimp-ind-pst-3pl-to-1pl-experimentaltextn<1K0 likes10 downloads1y agoHugging Face10pstuerner /ukraine-liveblog Dataset Card Dataset Summary The "ukraine-liveblog" dataset contains a collection of news articles published on the liveblog of the popular German news website, tagesschau.de. The dataset covers the period from February 2022 to February 2023, and includes every news feed published during this time that covers the ongoing war in Ukraine. Supported Tasks and Leaderboards -- Languages The language of the dataset is German. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/pstuerner/ukraine-liveblog.texttext-generation10K<n<100K1 likes9 downloads4y agoHugging Face11liu-nlp /estonian-blimp-ind-pst-3pl-to-2pl-experimentaltextn<1K0 likes6 downloads1y agoHugging Face12liu-nlp /estonian-blimp-ind-pst-3sg-to-2sg-experimentaltextn<1K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.