CoolFace
Modelpublic

ecfirst/360VL_PHI

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes8downloads
README.md126 linesDownload Raw Back to root
1---2license: apache-2.03datasets:4- liuhaotian/LLaVA-CC3M-Pretrain-595K5- liuhaotian/LLaVA-Instruct-150K6- FreedomIntelligence/ALLaVA-4V-Chinese7- shareAI/ShareGPT-Chinese-English-90k8language:9- zh10- en11pipeline_tag: visual-question-answering12---13<br>14<br>15 16# Model Card for 360VL 17<p align="center">18  <img src="https://github.com/360CVGroup/360VL/blob/master/qh360_vl/360vl.PNG?raw=true" width=100%/>19</p>20 21**360VL** is developed based on the LLama3 language model and is also the industry's first open source large multi-modal model based on **LLama3-70B**[[🤗Meta-Llama-3-70B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-70B-Instruct)]. In addition to applying the Llama3 language model, the 360VL model also designs a globally aware multi-branch projector architecture, which enables the model to have more sufficient image understanding capabilities.22 23**Github**:https://github.com/360CVGroup/360VL24## Model Zoo25 26360VL has released the following versions.27 28Model |       Download29|---|---30360VL-8B |  [🤗 Hugging Face](https://huggingface.co/qihoo360/360VL-8B) 31360VL-70B | [🤗 Hugging Face](https://huggingface.co/qihoo360/360VL-70B)  32## Features33 34360VL offers the following features:35 36- Multi-round text-image conversations: 360VL can take both text and images as inputs and produce text outputs. Currently, it supports multi-round visual question answering with one image.37  38- Bilingual text support: 360VL supports conversations in both English and Chinese, including text recognition in images.  39  40- Strong image comprehension: 360VL is adept at analyzing visuals, making it an efficient tool for tasks like extracting, organizing, and summarizing information from images.41  42- Fine-grained image resolution: 360VL supports image understanding at a higher resolution of 672&times;672.43 44## Performance45| Model               | Checkpoints   | MMB<sub>T  | MMB<sub>D|MMB-CN<sub>T  | MMB-CN<sub>D|MMMU<sub>V|MMMU<sub>T| MME |46|:--------------------|:------------:|:----:|:------:|:------:|:-------:|:-------:|:-------:|:-------:|47| QWen-VL-Chat |  [🤗LINK](https://huggingface.co/Qwen/Qwen-VL-Chat) | 61.8 | 60.6 |  56.3  |  56.7  |37| 32.9  | 1860 |48| mPLUG-Owl2 |  [🤖LINK](https://www.modelscope.cn/models/iic/mPLUG-Owl2/summary) | 66.0 | 66.5 |  60.3  |  59.5  |34.7| 32.1  | 1786.4 |49| CogVLM |  [🤗LINK](https://huggingface.co/THUDM/cogvlm-grounding-generalist-hf) | 65.8| 63.7 | 55.9  | 53.8    |37.3| 30.1 | 1736.6|50| Monkey-Chat |  [🤗LINK](https://huggingface.co/echo840/Monkey-Chat) | 72.4| 71 | 67.5  | 65.8    |40.7| - | 1887.4|51| MM1-7B-Chat |  [LINK](https://ar5iv.labs.arxiv.org/html/2403.09611) | -| 72.3 | -  | -    |37.0| 35.6 |  1858.2|52| IDEFICS2-8B |  [🤗LINK](https://huggingface.co/HuggingFaceM4/idefics2-8b) | 75.7 | 75.3 | 68.6  | 67.3    |43.0| 37.7 |1847.6| 53| SVIT-v1.5-13B|  [🤗LINK](https://huggingface.co/Isaachhe/svit-v1.5-13b-full) | 69.1 | - | 63.1  |  -  | 38.0| 33.3|1889| 54| LLaVA-v1.5-13B |  [🤗LINK](https://huggingface.co/liuhaotian/llava-v1.5-13b) | 69.2 | 69.2 | 65  | 63.6    |36.4| 33.6 | 1826.7|  55| LLaVA-v1.6-13B |  [🤗LINK](https://huggingface.co/liuhaotian/llava-v1.6-vicuna-13b) | 70 | 70.7 | 68.5  | 64.3    |36.2| - |1901|56| Honeybee |  [LINK](https://github.com/kakaobrain/honeybee) | 73.6 | 74.3 | -  | -    |36.2|  -|1976.5|57| YI-VL-34B |  [🤗LINK](https://huggingface.co/01-ai/Yi-VL-34B) | 72.4 | 71.1 |  70.7 |   71.4  |45.1| 41.6 |2050.2|58| **360VL-8B** |  [🤗LINK](https://huggingface.co/qihoo360/360VL-8B) | 75.3 | 73.7 | 71.1   | 68.6    |39.7| 37.1 |  1944.6|59| **360VL-70B** |  [🤗LINK](https://huggingface.co/qihoo360/360VL-70B) | 78.1 | 80.4 | 76.9   | 77.7    |50.8| 44.3 |  2012.3|60## Quick Start 🤗61 62```Shell63from transformers import AutoModelForCausalLM, AutoTokenizer64import torch65from PIL import Image66 67checkpoint = "qihoo360/360VL-8B"68 69model = AutoModelForCausalLM.from_pretrained(checkpoint, torch_dtype=torch.float16, device_map='auto', trust_remote_code=True).eval()70tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)71vision_tower = model.get_vision_tower()72vision_tower.load_model()73vision_tower.to(device="cuda", dtype=torch.float16)74image_processor = vision_tower.image_processor75tokenizer.pad_token = tokenizer.eos_token76 77 78image = Image.open("docs/008.jpg").convert('RGB')79query = "Who is this cartoon character?"80terminators = [81    tokenizer.convert_tokens_to_ids("<|eot_id|>",)82]83 84inputs = model.build_conversation_input_ids(tokenizer, query=query, image=image, image_processor=image_processor)85 86input_ids = inputs["input_ids"].to(device='cuda', non_blocking=True)87images = inputs["image"].to(dtype=torch.float16, device='cuda', non_blocking=True)88 89output_ids = model.generate(90    input_ids,91    images=images,92    do_sample=False,93    eos_token_id=terminators,94    num_beams=1,95    max_new_tokens=512,96    use_cache=True)97 98input_token_len = input_ids.shape[1]99outputs = tokenizer.batch_decode(output_ids[:, input_token_len:], skip_special_tokens=True)[0]100outputs = outputs.strip()101print(outputs)102```103 104**Model type:**105360VL-8B is an open-source chatbot trained by fine-tuning LLM on multimodal instruction-following data.106It is an auto-regressive language model, based on the transformer architecture.107Base LLM: [meta-llama/Meta-Llama-3-8B-Instruct](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct)108 109**Model date:**110360VL-8B was trained in April 2024.111 112 113 114## License115This project utilizes certain datasets and checkpoints that are subject to their respective original licenses. Users must comply with all terms and conditions of these original licenses.116The content of this project itself is licensed under the [Apache license 2.0]117 118**Where to send questions or comments about the model:**119https://github.com/360CVGroup/360VL120 121## Related Projects122This work wouldn't be possible without the incredible open-source code of these projects. Huge thanks!123- [Meta Llama 3](https://github.com/meta-llama/llama3)124- [LLaVA: Large Language and Vision Assistant](https://github.com/haotian-liu/LLaVA)125- [Honeybee: Locality-enhanced Projector for Multimodal LLM](https://github.com/kakaobrain/honeybee)126