CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /govdocs1-pdf-source govdocs1: source PDF files [!NOTE] Converted versions of other document types (word, txt, etc) are available in this repo This is ~220,000 open-access PDF documents (about 6.6M pages) from the dataset govdocs1. It wants to be OCR'd. Uploaded as tar file pieces of ~10 GiB each due to size/file count limits with an index.csv covering details 5,000 randomly sampled PDFs are available unarchived in sample/. Hugging Face supports previewing these in-browser, for example this one… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-pdf-source.documentimage-text-to-text100K<n<1M6 likes4.5k downloads9mo agoHugging Face02BEE-spoke-data /govdocs1-by-extension govdocs1 Dataset: By File Extension [!NOTE] PDFs from govdocs1 are at this repo in "raw" file form - no simple "mostly correct" way to convert to text Markdown-parsed versions of documents in govdocs1 with light filtering. Usage Load specific file formats (e.g., .doc files) parsed to markdown with pandoc: from datasets import load_dataset # Replace "doc" with desired config name dataset = load_dataset("BEE-spoke-data/govdocs1-by-extension", "doc")… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-by-extension.texttext-generation100K<n<1M2 likes714 downloads9mo agoHugging Face03BEE-spoke-data /govdocs1-image BEE-spoke-data/govdocs1-image This contains .jpg files from govdocs1. Light deduplication was applied (i.e. jdupes on all files) which removed ~500 duplicate images. DatasetDict({ train: Dataset({ features: ['image'], num_rows: 108895 }) }) source Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/ @inproceedings{garfinkel2009bringing, title={Bringing Science to Digital Forensics with Standardized Forensic Corpora}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-image.image100K<n<1M0 likes172 downloads9mo agoHugging Face04BEE-spoke-data /govdocs1-txt-raw Dataset Card for "govdocs1-txt-raw" Somewhere to put the raw txt files before filtering them Source info/page: https://digitalcorpora.org/corpora/file-corpora/files/ @inproceedings{garfinkel2009bringing, title={Bringing Science to Digital Forensics with Standardized Forensic Corpora}, author={Garfinkel, Simson and Farrell, Paul and Roussev, Vassil and Dinolt, George}, booktitle={Digital Forensic Research Workshop (DFRWS) 2009}, year={2009}, address={Montreal, Canada}… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/govdocs1-txt-raw.texttext-generation10K<n<100K0 likes154 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.