CoolFace
Datasetpublic

IDEA-CCNL/laion2B-multi-chinese-subset

laion2B-multi-chinese-subset Github: Fengshenbang-LM Docs: Fengshenbang-Docs 简介 Brief Introduction 取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。 A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion2B-multi 下载 Download mkdir laion2b_chinese_release && cd laion2b_chinese_release for i in… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
42likes348downloads
10 commits on main
58684b73y ago

Update README.md

wanng
c9452e44y ago

Update README.md

wanng
98893dc4y ago

Update README.md

wanng
0a0df2a4y ago

Update dataset_infos.json

wanng
4250e0e4y ago

update infos

wanng
2c61ab94y ago

upload from wanng

wanng
ae380674y ago

Update README.md

wanng
f8ec8704y ago

Update README.md

wanng
c61ff554y ago

Update README.md

wanng
57063014y ago

initial commit

wanng