FolSpark/DreamLLM-Qwen2.5-CLIP-SD2.1-CompreOnly
07
https://github.com/FolSpark/DreamLLM-Qwen2.5
Multiple self-trained DreamLLMs.
Model performance
DPO is trained using the dataset of MM-RLHF.
*indicates that only the comprehension data of LLaVA1.5 is used for the model's third-stage training.
Vicuna-CLIP-SD2.1 and Vicuna-CLIP-SD2.1* data comes from paper(https://openreview.net/forum?id=y01KGvd9Bw).
Multimodal Comprehension Assessment
Image Generation Evaluation
In the original text of DreamLLm, Vicuna-CLIP-SD2.1 and Vicuna-CLIP-SD2.1 were run 8 times, and the best one among 8 images was selected for each figure. All my models were only tested once, with an approximate error of 2~3.
