IDEA-CCNL/laion2B-multi-chinese-subset
laion2B-multi-chinese-subset Github: Fengshenbang-LM Docs: Fengshenbang-Docs 简介 Brief Introduction 取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。 A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion2B-multi 下载 Download mkdir laion2b_chinese_release && cd laion2b_chinese_release for i in… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face