CoolFace
Datasetpublic

weitao040702/minimind-v_dataset

Ⅰ 数据集 本轮训练用到的图文数据全部来自 ALLaVA-4V 系列。 相比以往从几份 LLaVA 衍生集拼接得到的数据,ALLaVA-4V 的质量更整齐、中英双语原生对照,细粒度描述也更充分。 它由两个子源构成:一份是 LAION 里挑出来的高质量图片(自然图像为主),一份是 VFLAN 指令流里挑出来的图片(文档、图表、合成场景居多)。 Pretrain(pretrain_i2t.parquet,约 127 万条 / ~64 万张唯一图像) ALLaVA-Caption-LAION-4V 英/中:~47万 + ~44万 ALLaVA-Caption-VFLAN-4V 英/中:~19万 + ~17万 任务形式为"请描述这张图片"类的单轮长描述,用于让模型建立视觉 token 到语言 token 的基础对齐。 SFT(sft_i2t.parquet,约 290 万条 / ~65 万张唯一图像) ALLaVA-Instruct-LAION-4V 英/中:~47万 + ~47万… See the full description on the dataset page: https://huggingface.co/datasets/weitao040702/minimind-v_dataset.

sourceHugging Faceapache-2.0updated 3mo agoView on Hugging Face
0likes13downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
weitao040702/minimind-v_dataset · CoolFace