CoolFace
Modelpublic

YuukiAsuna/Vintern-1B-v2-ViTable-docvqa

sourceHugging Facemitupdated 2y agoView on Hugging Face
2likes21downloads
Model Card

Vintern-1B-v2-ViTable-docvqa

<p align="center"> <a href="https://drive.google.com/file/d/1MU8bgsAwaWWcTl9GN1gXJcSPUSQoyWXy/view?usp=sharing"><b>Report Link</b>๐Ÿ‘๏ธ</a> </p>

<!-- Provide a quick summary of what the model is/does. --> Vintern-1B-v2-ViTable-docvqa is a fine-tuned version of the 5CD-AI/Vintern-1B-v2 multimodal model for the Vietnamese DocVQA (Table data)

Benchmarks

<div align="center">

ModelANLSSemantic SimilarityMLLM-as-judge (Gemini)
Gemini 1.5 Flash0.350.560.40
Vintern-1B-v20.040.450.50
Vintern-1B-v2-ViTable-docvqa0.500.710.59

</div>

<!-- Code benchmark: to be written later -->

Usage

Check out this **๐Ÿค— HF Demo**, or you can open it in Colab: ![Open In Colab](https://colab.research.google.com/drive/1ricMh4BxntoiXIT2CnQvAZjrGZTtx4gj?usp=sharing)

Citation:

bibtex
@misc{doan2024vintern1befficientmultimodallarge,
      title={Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese}, 
      author={Khang T. Doan and Bao G. Huynh and Dung T. Hoang and Thuc D. Pham and Nhat H. Pham and Quan T. M. Nguyen and Bang Q. Vo and Suong N. Hoang},
      year={2024},
      eprint={2408.12480},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2408.12480}, 
}