CoolFace
Datasetpublic

IDEA-CCNL/laion2B-multi-chinese-subset

laion2B-multi-chinese-subset Github: Fengshenbang-LM Docs: Fengshenbang-Docs 简介 Brief Introduction 取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。 A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion2B-multi 下载 Download mkdir laion2b_chinese_release && cd laion2b_chinese_release for i in… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.

sourceHugging Facecc-by-4.0updated 3y agoView on Hugging Face
42likes348downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
IDEA-CCNL/laion2B-multi-chinese-subset · CoolFace