CoolFace
20 results

Bbox

vikhyatk /openimages-bboxImages and nounding box annotations from the OpenImages dataset. image1M<n<10M10 likes2.5k downloads2y agoHugging Facejiani-huang /chris_stsg_bboxes_v00 likes1.4k downloads8mo agoHugging Facejiani-huang /droid_120_stsg_bboxestabular1M<n<10M0 likes839 downloads8mo agoHugging Facejrzhang /TextVQA_GT_bbox TextVQA validation set with grounding truth bounding box The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns. The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes. We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept. Citation If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.imagequestion-answering1K<n<10K4 likes807 downloads1y agoHugging Facejiani-huang /chris_stsg_bboxes_v1_part0tabular1M<n<10M0 likes671 downloads8mo agoHugging FaceReza2kn /persian-ocr-bench-submitted10-bbox-crops Persian OCR benchmark — selected submitted bbox crops This dataset contains the non-empty OCR bboxes from the ten explicitly selected submitted pages in persian_ocr_bench_bbox_review. Each row is one PNG crop. gold_text is the current editable OCR content from the live Argilla bbox field (content_text). Geometry is stored both as source page pixels and as percentages of the source page. The original record ID, external ID, bbox ID, source URL, and SHA-256 hashes are included for… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-ocr-bench-submitted10-bbox-crops.imageimage-to-textn<1K1 likes654 downloads25d agoHugging Face