CoolFace
Datasetpublic

TopicNet/Reuters

Reuters The Reuters Corpus contains 10,788 news documents totaling 1.3 million words. The documents have been classified into 90 topics, and grouped into two sets, called "training" and "test"; thus, the text with fileid 'test/14826' is a document drawn from the test set. This split is for training and testing algorithms that automatically detect the topic of a document, as we will see in chap-data-intensive. Language: English Number of topics: 90 Number of… See the full description on the dataset page: https://huggingface.co/datasets/TopicNet/Reuters.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
0likes29downloads
Dataset Card

Reuters

The Reuters Corpus contains 10,788 news documents totaling 1.3 million words. The documents have been classified into 90 topics, and grouped into two sets, called "training" and "test"; thus, the text with fileid 'test/14826' is a document drawn from the test set. This split is for training and testing algorithms that automatically detect the topic of a document, as we will see in chap-data-intensive.

  • Language: English
  • Number of topics: 90
  • Number of articles: ~10.000
  • Year: 2000

References

  • NLTK datasets: https://www.nltk.org/book/ch02.html.
  • Dataset site: https://trec.nist.gov/data/reuters/reuters.html.