CoolFace
Datasetpublic

pixparse/idl-wds

Dataset Card for Industry Documents Library (IDL) Dataset Summary Industry Documents Library (IDL) is a document dataset filtered from UCSF documents library with 19 million pages kept as valid samples. Each document exists as a collection of a pdf, a tiff image with the same contents rendered, a json file containing extensive Textract OCR annotations from the idl_data project, and a .ocr file with the original, older OCR annotation. In each pdf, there may be from… See the full description on the dataset page: https://huggingface.co/datasets/pixparse/idl-wds.

sourceHugging Faceotherupdated 2y agoView on Hugging Face
196likes2.8kdownloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
pixparse/idl-wds · CoolFace