datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enron-emailsARCO-EMARS
EMARS — ARCO Format
Analysis-Ready and Cloud Optimized Zarr conversion of the EMARS v1.0 reanalysis (Greybush et al., 2019) covering Mars Years 24–33.
Citation
Bhattacharya, A. (2026). ARCO-EMARS: Analysis-Ready Cloud-Optimized EMARS Reanalysis. https://doi.org/10.57967/hf/8859
Greybush et al., (2018), The Ensemble Mars Atmosphere Reanalysis System (EMARS) Version 1.0 Dataset, doi:10.18113/D3W375
Greybush, Steven J., Eugenia Kalnay, R. John Wilson, Ross N. Hoffman… See the full description on the dataset page: https://huggingface.co/datasets/ananyo01/ARCO-EMARS.enron-qa-emails-dasovich-jemaenron-emailThis dataset includes emails from Enron Email Dataset with prompts processed from Are Large Pre-Trained Language Models Leaking Your Personal Information?.
To use the dataset, you can run the following in LLM-PBE.
from data.enron import EnronDataset
ds = EnronDataset(data_path="data/enron", pseudonymize=False)
features-dinov3-vith16plus-224-imagenet-22k-wdsenron_aeslc_emailsphishing-email-dataset
Phishing Email Dataset
This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page.
comix-v0_1-pages
CoMix v0.1 - Pages Dataset
This is the Full CoMix dataset for page-level work. Download comix-v0_1-pages-tiny for fast experiments.
Some numbers: 19063 books, 894633 single pages, 6M+ single panels. v0.1 has a few broken tars, total number of books should be >20k).
Note: Dataset viewer currently struggles with this dataset because seg.npz files are custom NumPy archives with variable keys/shapes per page.
Will improve in following versions.
... add here an [image of the CoMix… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix-v0_1-pages.guertin-mcro-forensic-corpus-emails
Guertin MCRO Forensic Corpus: Emails
Contents: 352 email messages (.eml) in 11 categories — LinkedIn search-appearance notifications (99); messages before 2023-01-21 (82); messages after 2023-01-21 (76); correspondence with the first Rule 20 examiner (26); correspondence with the public defender (41); delivery of the March 5, 2025 hearing transcript (1); the Minnesota Attorney General's office (federal case) (1); 2026 correspondence (15); U.S. Senator Amy Klobuchar's office… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-emails.seven-phishing-email-datasets
Dataset Card for Seven Phishing/Spam Email Datasets
Dataset Summary
This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks.
Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label).
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.enron_emailhillary-clinton-emails-wikileaksema-ecrits-scolaires
[!NOTE]
Dataset origin: https://www.ortolang.fr/market/corpora/ema-ecrits-scolaires-1
Description
Corpus ÉMA, écrits scolaires
Les textes réunis sous le titre Corpus ÉMA, écrits scolaires constituent le premier ensemble d’un grand corpus longitudinal d’écrits scolaires destinés à la connaissance de la langue écrite des élèves de l’école primaire et du collège. La date de début du recueil coïncide avec la mise en œuvre des programmes 2015 qui préconisent une diversification des… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/ema-ecrits-scolaires.vertebrate_genomesshadow-eo
Dataset Card for S-EO: A Large-Scale Dataset for Geometry-Aware Shadow Detection in Remote Sensing Applications
Project page
We introduce the S-EO dataset: a large-scale, high-resolution dataset designed to advance geometry-aware shadow detection. Collected from diverse public-domain sources, including challenge datasets and government providers such as USGS, our dataset comprises 702 georeferenced tiles across the USA, each covering 500 × 500 meters. Each tile includes multi-date… See the full description on the dataset page: https://huggingface.co/datasets/emasquil/shadow-eo.PRISM
PRISM
PRISM: A Promptable and Robust Interactive Segmentation Model with Visual Prompts
Placenta application:
PRISM Lite: A lightweight model for interactive 3D placenta segmentation in ultrasound
Interactive Segmentation Model for Placenta Segmentation from 3D Ultrasound Images (arXiv version)
News
[07/07/24] Check out the decent performance/version of PRISM on placenta segmentation in ultrasound images.
[05/13/24] Our work is early accepted by MICCAI 2024.… See the full description on the dataset page: https://huggingface.co/datasets/emad2001/PRISM.email-Enronemail-W3Cenron_emails_sample_questionsFinePersonas-Synthetic-Email-Conversations
FinePersonas Synthetic Email Conversations
FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B.
🗞️ News
[10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations.
How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.Angiosperm_65_genomes_8192bp_uint8mishkat-quran-audio
تلاوةُ الحصريّ — مرآةٌ لعمل مشكاة بلا إنترنت
Al-Ḥuṣarī recitation — an offline mirror for Mishkat
العربيّة أوّلاً، ثمّ الإنجليزيّة. · Arabic first, then English.
ما هذا؟
ملفّاتُ تلاوةٍ آيةً آيةً للشيخ محمود خليل الحصريّ، منسوخةٌ كما هي من
everyayah.com بلا قصٍّ ولا إعادةِ ترميز، ليعمل بها
تطبيق مشكاة — الاستماعُ والتلقينُ — بلا اتّصالٍ
بالإنترنت.
ولا نصَّ قرآنٍ في هذا المستودع ولا تفسير — صوتٌ فقط، ومانيفستٌ يصفه.
القارئان — ولكلٍّ… See the full description on the dataset page: https://huggingface.co/datasets/emadjumaah/mishkat-quran-audio.small_vertebrate_genomes_8192train5
Dataset Card for "train5"
More Information needed
ox-alpha-pi-traces-emacs-bridge
oxalpha-emacs-pipeline
This project converts the agent traces in
TeichAI/Ox-Alpha-Pi-Traces
so that every file creation and file edit runs in Emacs through
python-bridge instead of
the original write and edit tools. You get two converted copies of the
dataset, an Emacs helper that exposes the bridge methods, and a system prompt
you can hand to a prime-agent-compatible runtime.
How the conversion works
The source traces store each agent action as a toolCall block… See the full description on the dataset page: https://huggingface.co/datasets/Jadson/ox-alpha-pi-traces-emacs-bridge.email-writing-sft-100k
Email Writing SFT (100K)
100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails.
Motivation
Email is the primary communication channel for most professional work, yet LLMs often produce emails that are:
Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.open-email-marketing-dataset
Open Email Marketing Dataset
This repository contains the Open Email Marketing Dataset, a collection of 1,000 question-and-answer pairs in JSONL format. This dataset is created and maintained by LeadsBlue.com to provide a high-quality, public resource for developers, researchers, and marketers. It is specifically designed for tasks such as fine-tuning Large Language Models (LLMs), building advanced Q&A engines, developing cold email tools, and enhancing SEO systems.… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/open-email-marketing-dataset.Marketing-Emails
Marketing Emails
A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research.
The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.train4
Dataset Card for "train4"
More Information needed
