Yale-LILY/aeslc
Dataset Card for "aeslc" Dataset Summary A collection of email messages of employees in the Enron Corporation. There are two features: email_body: email body text. subject_line: email subject text. Supported Tasks and Leaderboards More Information Needed Languages Monolingual English (mainly en-US) with some exceptions. Dataset Structure Data Instances default Size of downloaded dataset… See the full description on the dataset page: https://huggingface.co/datasets/Yale-LILY/aeslc.
Dataset Card for "aeslc"
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage:
- Repository: https://github.com/ryanzhumich/AESLC
- Paper: This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation
- Point of Contact: More Information Needed
- Size of downloaded dataset files: 11.64 MB
- Size of the generated dataset: 14.95 MB
- Total amount of disk used: 26.59 MB
Dataset Summary
A collection of email messages of employees in the Enron Corporation.
There are two features:
- email_body: email body text.
- subject_line: email subject text.
Supported Tasks and Leaderboards
Languages
Monolingual English (mainly en-US) with some exceptions.
Dataset Structure
Data Instances
default
- Size of downloaded dataset files: 11.64 MB
- Size of the generated dataset: 14.95 MB
- Total amount of disk used: 26.59 MB
An example of 'train' looks as follows.
{
"email_body": "B/C\n<<some doc>>\n",
"subject_line": "Service Agreement"
}Data Fields
The data fields are the same among all splits.
default
email_body: astringfeature.subject_line: astringfeature.
Data Splits
Dataset Creation
Curation Rationale
Source Data
Initial Data Collection and Normalization
Who are the source language producers?
Annotations
Annotation process
Who are the annotators?
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
@inproceedings{zhang-tetreault-2019-email,
title = "This Email Could Save Your Life: Introducing the Task of Email Subject Line Generation",
author = "Zhang, Rui and
Tetreault, Joel",
booktitle = "Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics",
month = jul,
year = "2019",
address = "Florence, Italy",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/P19-1043",
doi = "10.18653/v1/P19-1043",
pages = "446--456",
}Contributions
Thanks to @patrickvonplaten, @thomwolf, @lewtun for adding this dataset.
