CoolFace
Datasetpublic

nanonets/small_sparse_unstructured_table

This dataset is generated syhthetically to create tables with following characteristics: Empty cell percentage in following range [40,70] (Sparse) There is no seperator between rows and columns (un-structured). 4 <= num rows <= 10, 2 <= num columns <= 6 (Small) Load the dataset import io import pandas as pd from PIL import Image def bytes_to_image(self, image_bytes: bytes): return Image.open(io.BytesIO(image_bytes)) def parse_annotations(self, annotations: str) ->… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/small_sparse_unstructured_table.

sourceHugging Facemitupdated 1y agoView on Hugging Face
0likes26downloads
Dataset Card

This dataset is generated syhthetically to create tables with following characteristics:

  1. 1.Empty cell percentage in following range [40,70] (Sparse)
  2. 2.There is no seperator between rows and columns (un-structured).
  3. 3.4 <= num rows <= 10, 2 <= num columns <= 6 (Small)

Load the dataset

python
import io
import pandas as pd
from PIL import Image

def bytes_to_image(self, image_bytes: bytes):
  return Image.open(io.BytesIO(image_bytes))

def parse_annotations(self, annotations: str) -> pd.DataFrame:
  return pd.read_json(StringIO(annotations), orient="records")

test_data = load_dataset('nanonets/small_sparse_unstructured_table', split='test')
data_point = test_data[0]
image, gt_table = (
    bytes_to_image(data_point["images"]),
    parse_annotations(data_point["annotation"]),
)