CoolFace
Datasetpublic

byczong/pl-insurance-terms-struct

A dataset for the task of document structuring (parsing) Polish legal documents with nested lists. The dataset contains 2 columns: image: pdf pages converted to 1080x1440 images, gt_json: a json containing a list of detected objects under gt_parse key. The detected objects in the gt_json column are specified as follows: Object Type Keys Description Heading content Represents a heading-like loose block of text. List title, items A collection of items, which can be an element, an… See the full description on the dataset page: https://huggingface.co/datasets/byczong/pl-insurance-terms-struct.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
2likes23downloads
Dataset Card

A dataset for the task of document structuring (parsing) Polish legal documents with nested lists.

The dataset contains 2 columns:

  • —image: pdf pages converted to 1080x1440 images,
  • —gt_json: a json containing a list of detected objects under gt_parse key.

The detected objects in the gt_json column are specified as follows:

Object TypeKeysDescription
HeadingcontentRepresents a heading-like loose block of text.
Listtitle, itemsA collection of items, which can be an element, an overflowing element, or another list.
ElementcontentA basic unit of content within a list (along with the list prefix, e.g. "1.") or a loose block of text.
Overflowing ElementcontentContent that continues from the main list or element, but isn't part of the standard list flow.
Element ContinuationcontentContent that continues from an element on the previous page.
List ContinuationitemsList continuation from the previous page; items can be an element, an overflowing element, or another list.

All images were obtained from Polish insurance companies official websites. All sources are provided below: