CoolFace
15 results

textvqa

lmms-lab-encoder /textvqa Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of TextVQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{singh2019towards, title={Towards vqa models that can read}, author={Singh, Amanpreet and… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/textvqa.image10K<n<100K25 likes48k downloads3y agoHugging Facemaoxx241 /textvqa_subsetimage1K<n<10K2 likes1.1k downloads2y agoHugging Facefacebook /textvqaTextVQA requires models to read and reason about text in images to answer questions about them. Specifically, models need to incorporate a new modality of text present in the images and reason over it to answer TextVQA questions. TextVQA dataset contains 45,336 questions over 28,408 images from the OpenImages dataset.visual-question-answering10K<n<100K38 likes828 downloads3y agoHugging Facejrzhang /TextVQA_GT_bbox TextVQA validation set with grounding truth bounding box The dataset used in the paper MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs for studying MLLMs' attention patterns. The dataset is sourced from TextVQA and annotated manually with ground-truth bounding boxes. We consider questions with a single area of interest in the image so that 4370 out of 5000 samples are kept. Citation If you find our paper and code useful… See the full description on the dataset page: https://huggingface.co/datasets/jrzhang/TextVQA_GT_bbox.imagequestion-answering1K<n<10K4 likes807 downloads1y agoHugging Facenielsr /textvqa-sampleimagen<1K0 likes454 downloads4y agoHugging FaceMultimodal-Fatima /TextVQA_train Dataset Card for "TextVQA_train" More Information needed image10K<n<100K1 likes377 downloads3y agoHugging Face