CoolFace
Datasetpublic

Mozilla/docornot

The DocOrNot dataset contains 50% of images that are pictures, and 50% that are documents. It was built using 8k images from each one of these sources: RVL CDIP (Small) - https://www.kaggle.com/datasets/uditamin/rvl-cdip-small - license: https://www.industrydocuments.ucsf.edu/help/copyright/ Flickr8k - https://www.kaggle.com/datasets/adityajn105/flickr8k - license: https://creativecommons.org/publicdomain/zero/1.0/ It can be used to train a model and classify an image as being a picture or a… See the full description on the dataset page: https://huggingface.co/datasets/Mozilla/docornot.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
6likes38downloads
Dataset Card

The DocOrNot dataset contains 50% of images that are pictures, and 50% that are documents.

It was built using 8k images from each one of these sources:

  • RVL CDIP (Small) - https://www.kaggle.com/datasets/uditamin/rvl-cdip-small - license: https://www.industrydocuments.ucsf.edu/help/copyright/
  • Flickr8k - https://www.kaggle.com/datasets/adityajn105/flickr8k - license: https://creativecommons.org/publicdomain/zero/1.0/

It can be used to train a model and classify an image as being a picture or a document.

Source code used to generate this dataset : https://github.com/mozilla/docornot