hakuhodo-tech/japanese-clip-vit-h-14-bert-deeper
064
1---2license: cc-by-nc-sa-4.03language:4 - ja5tags:6 - clip7 - ja8 - japanese9 - japanese-clip10pipeline_tag: feature-extraction11---12 13# Japanese CLIP ViT-H/14 (Deeper)14 15## Table of Contents16 171. [Overview](#overview)181. [Usage](#usage)191. [Model Details](#model-details)201. [Evaluation](#evaluation)211. [Limitations and Biases](#limitations-and-biases)221. [Citation](#citation)231. [See Also](#see-also)241. [Contact Information](#contact-information)25 26## Overview27 28* **Developed by**: [HAKUHODO Technologies Inc.](https://www.hakuhodo-technologies.co.jp/)29* **Model type**: Contrastive Language-Image Pre-trained Model30* **Language(s)**: Japanese31* **LICENSE**: [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)32 33Presented here is a Japanese [CLIP (Contrastive Language-Image Pre-training)](https://arxiv.org/abs/2103.00020) model,34mapping Japanese texts and images to a unified embedding space.35Capable of multimodal tasks including zero-shot image classification,36text-to-image retrieval, and image-to-text retrieval,37this model extends its utility when integrated with other components,38contributing to generative models like image-to-text and text-to-image generation.39 40## Usage41 42### Dependencies43 44```bash45python3 -m pip install pillow sentencepiece torch torchvision transformers46```47 48### Inference49 50The usage is similar to [`CLIPModel`](https://huggingface.co/docs/transformers/model_doc/clip)51and [`VisionTextDualEncoderModel`](https://huggingface.co/docs/transformers/model_doc/vision-text-dual-encoder).52 53```python54import requests55import torch56from PIL import Image57from transformers import AutoModel, AutoProcessor, BatchEncoding58 59# Download60model_name = "hakuhodo-tech/japanese-clip-vit-h-14-bert-deeper"61device = "cuda" if torch.cuda.is_available() else "cpu"62model = AutoModel.from_pretrained(model_name, trust_remote_code=True).to(device)63processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)64 65# Prepare raw inputs66url = "http://images.cocodataset.org/val2017/000000039769.jpg"67image = Image.open(requests.get(url, stream=True).raw)68 69# Process inputs70inputs = processor(71 text=["犬", "猫", "象"],72 images=image,73 return_tensors="pt",74 padding=True,75)76 77# Infer and output78outputs = model(**BatchEncoding(inputs).to(device))79probs = outputs.logits_per_image.softmax(dim=1)80print([f"{x:.2f}" for x in probs.flatten().tolist()]) # ['0.00', '1.00', '0.00']81```82 83## Model Details84 85### Components86 87The model consists of a frozen ViT-H image encoder from88[laion/CLIP-ViT-H-14-laion2B-s32B-b79K](https://huggingface.co/laion/CLIP-ViT-H-14-laion2B-s32B-b79K)89and a 24-layer 12-head BERT text encoder initialized from90[hakuhodo-tech/japanese-clip-vit-h-14-bert-base](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-base)91with [Modified ZerO](https://www.anlp.jp/proceedings/annual_meeting/2024/pdf_dir/B6-5.pdf).92 93### Training94 95Model training is done by Zhi Wang with 8 A100 (80 GB) GPUs.96[Locked-image Tuning (LiT)](https://arxiv.org/abs/2111.07991) is adopted.97See more details in [the paper](https://www.anlp.jp/proceedings/annual_meeting/2024/pdf_dir/B6-5.pdf).98 99### Dataset100 101The Japanese subset of the [laion2B-multi](https://huggingface.co/datasets/laion/laion2B-multi) dataset containing ~120M image-text pairs.102 103## Evaluation104 105### Testing Data106 107The 5K evaluation set (val2017) of [MS-COCO](https://cocodataset.org/)108with [STAIR Captions](http://captions.stair.center/).109 110### Metrics111 112Zero-shot image-to-text and text-to-image recall@1, 5, 10.113 114### Results115 116| | | | | | | |117| :---------------------------------------------------------------------------------------------------------------------- | :------: | :------: | :------: | :------: | :------: | :------: |118| <td colspan=3 align=center>Text Retrieval</td> <td colspan=3 align=center>Image Retrieval</td> |119| | R@1 | R@5 | R@10 | R@1 | R@5 | R@10 |120| [recruit-jp/japanese-clip-vit-b-32-roberta-base](https://huggingface.co/recruit-jp/japanese-clip-vit-b-32-roberta-base) | 23.0 | 46.1 | 57.4 | 16.1 | 35.4 | 46.3 |121| [rinna/japanese-cloob-vit-b-16](https://huggingface.co/rinna/japanese-cloob-vit-b-16) | 37.1 | 63.7 | 74.2 | 25.1 | 48.0 | 58.8 |122| [rinna/japanese-clip-vit-b-16](https://huggingface.co/rinna/japanese-clip-vit-b-16) | 36.9 | 64.3 | 74.3 | 24.8 | 48.8 | 60.0 |123| [**Japanese CLIP ViT-H/14 (Base)**](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-base) | 39.2 | 66.3 | 76.6 | 28.9 | 53.3 | 63.9 |124| [**Japanese CLIP ViT-H/14 (Deeper)**](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-deeper) | **48.7** | 74.0 | 82.4 | 36.5 | 61.5 | 71.8 |125| [**Japanese CLIP ViT-H/14 (Wider)**](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-wider) | 47.9 | **74.2** | **83.2** | **37.3** | **62.8** | **72.7** |126 127\* [Japanese Stable CLIP ViT-L/16](https://huggingface.co/stabilityai/japanese-stable-clip-vit-l-16) is excluded for zero-shot retrieval evaluation as128[the model was partially pre-trained with MS-COCO](https://huggingface.co/stabilityai/japanese-stable-clip-vit-l-16#training-dataset).129 130## Limitations and Biases131 132Despite our data filtering, it is crucial133to acknowledge the possibility of the training dataset134containing offensive or inappropriate content.135Users should be mindful of the potential societal impact136and ethical considerations associated with the outputs137generated by the model when deploying in production systems.138It is recommended not to employ the model for applications139that have the potential to cause harm or distress140to individuals or groups.141 142## Citation143 144If you found this model useful, please consider citing:145 146```bibtex147@article{japanese-clip-vit-h,148 author = {王 直 and 細野 健人 and 石塚 湖太 and 奥田 悠太 and 川上 孝介},149 journal = {言語処理学会年次大会発表論文集},150 month = {Mar},151 pages = {1547--1552},152 title = {日本語特化の視覚と言語を組み合わせた事前学習モデルの開発 Developing Vision-Language Pre-Trained Models for {J}apanese},153 volume = {30},154 year = {2024}155}156```157 158## See Also159 160* [Japanese CLIP ViT-H/14 (Base)](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-base)161* [Japanese CLIP ViT-H/14 (Wider)](https://huggingface.co/hakuhodo-tech/japanese-clip-vit-h-14-bert-wider)162 163## Contact Information164 165Please contact166[hr-koho\@hakuhodo-technologies.co.jp](mailto:hr-koho@hakuhodo-technologies.co.jp?subject=Japanese%20CLIP%20ViT-H/14%20Models)167for questions and comments about the model,168and/or169for business and partnership inquiries.170 171お問い合わせは172[hr-koho\@hakuhodo-technologies.co.jp](mailto:hr-koho@hakuhodo-technologies.co.jp?subject=日本語CLIP%20ViT-H/14モデルについて)173にご連絡ください。174 