CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01corbt /enron-emailstext100K<n<1M13 likes13k downloads1y agoHugging Face02ananyo01 /ARCO-EMARS EMARS — ARCO Format Analysis-Ready and Cloud Optimized Zarr conversion of the EMARS v1.0 reanalysis (Greybush et al., 2019) covering Mars Years 24–33. Citation Bhattacharya, A. (2026). ARCO-EMARS: Analysis-Ready Cloud-Optimized EMARS Reanalysis. https://doi.org/10.57967/hf/8859 Greybush et al., (2018), The Ensemble Mars Atmosphere Reanalysis System (EMARS) Version 1.0 Dataset, doi:10.18113/D3W375 Greybush, Steven J., Eugenia Kalnay, R. John Wilson, Ross N. Hoffman… See the full description on the dataset page: https://huggingface.co/datasets/ananyo01/ARCO-EMARS.0 likes4k downloads1mo agoHugging Face03weaviate /enron-qa-emails-dasovich-jtext1K<n<10K0 likes3.2k downloads1y agoHugging Face04duskchant /ema0 likes2.8k downloads1y agoHugging Face05LLM-PBE /enron-emailThis dataset includes emails from Enron Email Dataset with prompts processed from Are Large Pre-Trained Language Models Leaking Your Personal Information?. To use the dataset, you can run the following in LLM-PBE. from data.enron import EnronDataset ds = EnronDataset(data_path="data/enron", pseudonymize=False) text100K<n<1M6 likes1.9k downloads2y agoHugging Face06Emanresu /features-dinov3-vith16plus-224-imagenet-22k-wdstext1M<n<10M0 likes1.9k downloads11mo agoHugging Face07snoop2head /enron_aeslc_emailstext100K<n<1M13 likes1.6k downloads4y agoHugging Face08zefang-liu /phishing-email-dataset Phishing Email Dataset This dataset on Hugging Face is a direct copy of the 'Phishing Email Detection' dataset from Kaggle, shared under the GNU Lesser General Public License 3.0. The dataset was originally created by the user 'Cyber Cop' on Kaggle. For complete details, including licensing and usage information, please visit the original Kaggle page. texttext-classification10K<n<100K38 likes1.1k downloads3y agoHugging Face09emanuelevivoli /comix-v0_1-pagesgated CoMix v0.1 - Pages Dataset This is the Full CoMix dataset for page-level work. Download comix-v0_1-pages-tiny for fast experiments. Some numbers: 19063 books, 894633 single pages, 6M+ single panels. v0.1 has a few broken tars, total number of books should be >20k). Note: Dataset viewer currently struggles with this dataset because seg.npz files are custom NumPy archives with variable keys/shapes per page. Will improve in following versions. ... add here an [image of the CoMix… See the full description on the dataset page: https://huggingface.co/datasets/emanuelevivoli/comix-v0_1-pages.image-to-text100K<n<1M2 likes976 downloads10mo agoHugging Face10Matt1up /guertin-mcro-forensic-corpus-emails Guertin MCRO Forensic Corpus: Emails Contents: 352 email messages (.eml) in 11 categories — LinkedIn search-appearance notifications (99); messages before 2023-01-21 (82); messages after 2023-01-21 (76); correspondence with the first Rule 20 examiner (26); correspondence with the public defender (41); delivery of the March 5, 2025 hearing transcript (1); the Minnesota Attorney General's office (federal case) (1); 2026 correspondence (15); U.S. Senator Amy Klobuchar's office… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-emails.n<1K0 likes910 downloads8d agoHugging Face11puyang2025 /seven-phishing-email-datasets Dataset Card for Seven Phishing/Spam Email Datasets Dataset Summary This dataset is a unified, row-level email corpus built from seven commonly used public email datasets. It is intended for research on phishing/spam detection and related email-text classification tasks. Each row contains the email body (text), optional header-like fields (e.g., sender, receiver, date), the source dataset name (dataset_name), and a binary label (label). Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/puyang2025/seven-phishing-email-datasets.tabulartext-classification100K<n<1M1 likes595 downloads8mo agoHugging Face12ridalefdali /enron_emaildocumentn<1K0 likes575 downloads1y agoHugging Face13from-our-page /hillary-clinton-emails-wikileakstext0 likes525 downloads1y agoHugging Face14datasets-CNRS /ema-ecrits-scolaires [!NOTE] Dataset origin: https://www.ortolang.fr/market/corpora/ema-ecrits-scolaires-1 Description Corpus ÉMA, écrits scolaires Les textes réunis sous le titre Corpus ÉMA, écrits scolaires constituent le premier ensemble d’un grand corpus longitudinal d’écrits scolaires destinés à la connaissance de la langue écrite des élèves de l’école primaire et du collège. La date de début du recueil coïncide avec la mise en œuvre des programmes 2015 qui préconisent une diversification des… See the full description on the dataset page: https://huggingface.co/datasets/datasets-CNRS/ema-ecrits-scolaires.0 likes521 downloads1y agoHugging Face15emarro /vertebrate_genomestabular10M<n<100M0 likes513 downloads9mo agoHugging Face16emasquil /shadow-eo Dataset Card for S-EO: A Large-Scale Dataset for Geometry-Aware Shadow Detection in Remote Sensing Applications Project page We introduce the S-EO dataset: a large-scale, high-resolution dataset designed to advance geometry-aware shadow detection. Collected from diverse public-domain sources, including challenge datasets and government providers such as USGS, our dataset comprises 702 georeferenced tiles across the USA, each covering 500 × 500 meters. Each tile includes multi-date… See the full description on the dataset page: https://huggingface.co/datasets/emasquil/shadow-eo.image100K<n<1M4 likes504 downloads1y agoHugging Face17emad2001 /PRISM PRISM PRISM: A Promptable and Robust Interactive Segmentation Model with Visual Prompts Placenta application: PRISM Lite: A lightweight model for interactive 3D placenta segmentation in ultrasound Interactive Segmentation Model for Placenta Segmentation from 3D Ultrasound Images (arXiv version) News [07/07/24] Check out the decent performance/version of PRISM on placenta segmentation in ultrasound images. [05/13/24] Our work is early accepted by MICCAI 2024.… See the full description on the dataset page: https://huggingface.co/datasets/emad2001/PRISM.0 likes459 downloads8mo agoHugging Face18HypernetworkRG /email-Enron0 likes457 downloads6mo agoHugging Face19HypernetworkRG /email-W3C0 likes457 downloads6mo agoHugging Face20corbt /enron_emails_sample_questionstabular10K<n<100K11 likes423 downloads10mo agoHugging Face21argilla /FinePersonas-Synthetic-Email-Conversations FinePersonas Synthetic Email Conversations FinePersonas Synthetic Email Conversations is a dataset containing around 115k conversations via email between two personas from the argilla/FinePersonas-v0.1. Conversations were generated using NousResearch/Hermes-3-Llama-3.1-70B. 🗞️ News [10/16/2024] New subsets: added two new subsets unfriendly_email_conversations and unprofessional_email_conversations. How were the conversations generated?… See the full description on the dataset page: https://huggingface.co/datasets/argilla/FinePersonas-Synthetic-Email-Conversations.texttext-generation100K<n<1M8 likes390 downloads2y agoHugging Face22emarro /Angiosperm_65_genomes_8192bp_uint81M<n<10M0 likes292 downloads8d agoHugging Face23emadjumaah /mishkat-quran-audio تلاوةُ الحصريّ — مرآةٌ لعمل مشكاة بلا إنترنت Al-Ḥuṣarī recitation — an offline mirror for Mishkat العربيّة أوّلاً، ثمّ الإنجليزيّة. · Arabic first, then English. ما هذا؟ ملفّاتُ تلاوةٍ آيةً آيةً للشيخ محمود خليل الحصريّ، منسوخةٌ كما هي من everyayah.com بلا قصٍّ ولا إعادةِ ترميز، ليعمل بها تطبيق مشكاة — الاستماعُ والتلقينُ — بلا اتّصالٍ بالإنترنت. ولا نصَّ قرآنٍ في هذا المستودع ولا تفسير — صوتٌ فقط، ومانيفستٌ يصفه. القارئان — ولكلٍّ… See the full description on the dataset page: https://huggingface.co/datasets/emadjumaah/mishkat-quran-audio.audio10K<n<100K0 likes246 downloads1mo agoHugging Face24emarro /small_vertebrate_genomes_8192tabular1M<n<10M0 likes238 downloads6mo agoHugging Face25emaeon /train5 Dataset Card for "train5" More Information needed tabular1M<n<10M0 likes229 downloads2y agoHugging Face26Jadson /ox-alpha-pi-traces-emacs-bridge oxalpha-emacs-pipeline This project converts the agent traces in TeichAI/Ox-Alpha-Pi-Traces so that every file creation and file edit runs in Emacs through python-bridge instead of the original write and edit tools. You get two converted copies of the dataset, an Emacs helper that exposes the bridge methods, and a system prompt you can hand to a prime-agent-compatible runtime. How the conversion works The source traces store each agent action as a toolCall block… See the full description on the dataset page: https://huggingface.co/datasets/Jadson/ox-alpha-pi-traces-emacs-bridge.text-generation100K<n<1M0 likes228 downloads18d agoHugging Face27stindardlogic /email-writing-sft-100k Email Writing SFT (100K) 100,000 ShareGPT conversations demonstrating professional email writing across 22 business contexts. Each example shows how to draft clear, purposeful emails that achieve their communication goal — from cold outreach to salary negotiations to apology emails. Motivation Email is the primary communication channel for most professional work, yet LLMs often produce emails that are: Too long: Including unnecessary preamble, excessive context… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/email-writing-sft-100k.texttext-generation100K<n<1M3 likes213 downloads2mo agoHugging Face28emailmarketingdataset /open-email-marketing-dataset Open Email Marketing Dataset This repository contains the Open Email Marketing Dataset, a collection of 1,000 question-and-answer pairs in JSONL format. This dataset is created and maintained by LeadsBlue.com to provide a high-quality, public resource for developers, researchers, and marketers. It is specifically designed for tasks such as fine-tuning Large Language Models (LLMs), building advanced Q&A engines, developing cold email tools, and enhancing SEO systems.… See the full description on the dataset page: https://huggingface.co/datasets/emailmarketingdataset/open-email-marketing-dataset.1 likes210 downloads1y agoHugging Face29marketeam /Marketing-Emails Marketing Emails A curated corpus of synthetically generated yet realistic marketing email messages designed to support research in Domain Adaptation, Natural Language Processing (NLP), Data Science, Machine Learning, and Communication research. The dataset is appropriate for a wide spectrum of training paradigms—including pre-training, fine-tuning, and domain adaptation—as well as for rigorous evaluation of models targeting domain-specific language understanding and generation… See the full description on the dataset page: https://huggingface.co/datasets/marketeam/Marketing-Emails.texttext-generation10K<n<100K17 likes205 downloads10mo agoHugging Face30emaeon /train4 Dataset Card for "train4" More Information needed tabular1M<n<10M0 likes203 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.