google-research-datasets/newsgroup
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups. The 20 newsgroups collection has become a popular data set for experiments in text applications of machine learning techniques, such as text classification and text clustering.
Dataset Card for "newsgroup"
Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage: http://qwone.com/~jason/20Newsgroups/
- Repository: More Information Needed
- Paper: NewsWeeder: Learning to Filter Netnews
- Point of Contact: More Information Needed
- Size of downloaded dataset files: 929.27 MB
- Size of the generated dataset: 124.41 MB
- Total amount of disk used: 1.05 GB
Dataset Summary
The 20 Newsgroups data set is a collection of approximately 20,000 newsgroup documents, partitioned (nearly) evenly across 20 different newsgroups. To the best of my knowledge, it was originally collected by Ken Lang, probably for his Newsweeder: Learning to filter netnews paper, though he does not explicitly mention this collection. The 20 newsgroups collection has become a popular data set for experiments in text applications of machine learning techniques, such as text classification and text clustering.
does not include cross-posts and includes only the "From" and "Subject" headers.
Supported Tasks and Leaderboards
Languages
Dataset Structure
Data Instances
18828_alt.atheism
- Size of downloaded dataset files: 14.67 MB
- Size of the generated dataset: 1.67 MB
- Total amount of disk used: 16.34 MB
An example of 'train' looks as follows.
18828_comp.graphics
- Size of downloaded dataset files: 14.67 MB
- Size of the generated dataset: 1.66 MB
- Total amount of disk used: 16.33 MB
An example of 'train' looks as follows.
18828_comp.os.ms-windows.misc
- Size of downloaded dataset files: 14.67 MB
- Size of the generated dataset: 2.38 MB
- Total amount of disk used: 17.05 MB
An example of 'train' looks as follows.
18828_comp.sys.ibm.pc.hardware
- Size of downloaded dataset files: 14.67 MB
- Size of the generated dataset: 1.18 MB
- Total amount of disk used: 15.85 MB
An example of 'train' looks as follows.
18828_comp.sys.mac.hardware
- Size of downloaded dataset files: 14.67 MB
- Size of the generated dataset: 1.06 MB
- Total amount of disk used: 15.73 MB
An example of 'train' looks as follows.
Data Fields
The data fields are the same among all splits.
18828_alt.atheism
text: astringfeature.
18828_comp.graphics
text: astringfeature.
18828_comp.os.ms-windows.misc
text: astringfeature.
18828_comp.sys.ibm.pc.hardware
text: astringfeature.
18828_comp.sys.mac.hardware
text: astringfeature.
Data Splits
Dataset Creation
Curation Rationale
Source Data
Initial Data Collection and Normalization
Who are the source language producers?
Annotations
Annotation process
Who are the annotators?
Personal and Sensitive Information
Considerations for Using the Data
Social Impact of Dataset
Discussion of Biases
Other Known Limitations
Additional Information
Dataset Curators
Licensing Information
Citation Information
@incollection{LANG1995331,
title = {NewsWeeder: Learning to Filter Netnews},
editor = {Armand Prieditis and Stuart Russell},
booktitle = {Machine Learning Proceedings 1995},
publisher = {Morgan Kaufmann},
address = {San Francisco (CA)},
pages = {331-339},
year = {1995},
isbn = {978-1-55860-377-6},
doi = {https://doi.org/10.1016/B978-1-55860-377-6.50048-7},
url = {https://www.sciencedirect.com/science/article/pii/B9781558603776500487},
author = {Ken Lang},
}Contributions
Thanks to @mariamabarham, @thomwolf, @lhoestq for adding this dataset.
