CoolFace
Modelpublic

FolSpark/DreamLLM-Qwen2.5-CLIP-SD2.1-CompreOnly

sourceHugging Faceapache-2.0updated 1y agoView on Hugging Face
0likes7downloads
Model Card

https://github.com/FolSpark/DreamLLM-Qwen2.5

Multiple self-trained DreamLLMs.

Model performance

DPO is trained using the dataset of MM-RLHF.

*indicates that only the comprehension data of LLaVA1.5 is used for the model's third-stage training.

Vicuna-CLIP-SD2.1 and Vicuna-CLIP-SD2.1* data comes from paper(https://openreview.net/forum?id=y01KGvd9Bw).

Multimodal Comprehension Assessment

MethodCaptioningVQAComprehensive
COCO12ParagraphVQAv2OKVQAVizWizTextVQAMM-Vet
**Qwen-InternViT-SD3.5**106.410.773.954.249.154.844.0
**Qwen-InternViT-SD3.5***102.110.973.053.648.655.245.7
**Qwen-InternViT-SD3.5-DPO**64.611.674.250.948.955.844.7
Qwen-CLIP-SD3.599.99.772.952.349.044.039.8
Qwen-CLIP-SD3.5*99.110.272.751.149.143.942.1
Qwen-CLIP-SD2.182.89.172.552.449.443.642
Qwen-CLIP-SD2.1*97.310.872.450.449.943.239.0
Vicuna-CLIP-SD2.1115.417.456.644.345.834.935.9
Vicuna-CLIP-SD2.1*103.78.472.952.249.341.836.6

Image Generation Evaluation

MethodMS-COCO
Qwen-InternViT-SD3.5-Stage111.72
Qwen-InternViT-SD3.511.11
Qwen-InternViT-SD3.5-DPO11.33
Qwen-CLIP-SD3.5-Stage111.72
Qwen-CLIP-SD3.511.61
Qwen-CLIP-SD2.1-Stage113.94
Qwen-CLIP-SD2.112.26
Vicuna-CLIP-SD2.1-Stage18.76(+~2)
Vicuna-CLIP-SD2.18.46(+~2)

In the original text of DreamLLm, Vicuna-CLIP-SD2.1 and Vicuna-CLIP-SD2.1 were run 8 times, and the best one among 8 images was selected for each figure. All my models were only tested once, with an approximate error of 2~3.