nanonets/small_sparse_structured_table
This dataset is generated syhthetically to create tables with following characteristics: Empty cell percentage in following range [40,70] (Sparse) There is clear seperator between rows and columns (Structured). 4 <= num rows <= 10, 2 <= num columns <= 6 (Small) Load the dataset import io import pandas as pd from PIL import Image def bytes_to_image(self, image_bytes: bytes): return Image.open(io.BytesIO(image_bytes)) def parse_annotations(self, annotations: str) ->… See the full description on the dataset page: https://huggingface.co/datasets/nanonets/small_sparse_structured_table.
034
This dataset is generated syhthetically to create tables with following characteristics:
- Empty cell percentage in following range [40,70] (Sparse)
- There is clear seperator between rows and columns (Structured).
- 4 <= num rows <= 10, 2 <= num columns <= 6 (Small)
Load the dataset
import io
import pandas as pd
from PIL import Image
def bytes_to_image(self, image_bytes: bytes):
return Image.open(io.BytesIO(image_bytes))
def parse_annotations(self, annotations: str) -> pd.DataFrame:
return pd.read_json(StringIO(annotations), orient="records")
test_data = load_dataset('nanonets/small_sparse_structured_table', split='test')
data_point = test_data[0]
image, gt_table = (
bytes_to_image(data_point["images"]),
parse_annotations(data_point["annotation"]),
)