CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AmazarashiEndure /Chinese_Landscape_Painting 数据集名称 Chinese_Landscape_Painting 数据集简介 这是一份为南京大学智能科学与技术专业大二秋季学期课程人工智能导论课程的项目训练而搭建的数据集。 由于目前较大规模、高质量、且适应现代flux模型的高分辨率的山水画数据集稀缺,故我们搭建了这个数据集,以进行flux模型的lora微调训练。 数据集包含了1017张局部图片,79张全景图片,全部采样自中国山水画的十大名画,并全部带有精细的严格结构化标注。 如果你想使用该数据集进行flux模型训练,可以直接下载并自行修改相应的子文件夹名称以改变每张图片的训练次数。 引用说明 如果您在研究或项目中使用了本数据集,请按以下格式引用: BibTeX: @dataset{Chinese_Landscape_Painting, author = {Wei Liangxu}, title = {Chinese_Landscape_Painting}, year = {2025}… See the full description on the dataset page: https://huggingface.co/datasets/AmazarashiEndure/Chinese_Landscape_Painting.imageimage-to-image1K<n<10K4 likes1.2k downloads10mo agoHugging Face02kaupane /chinese-painting-collection Chinese Painting Collection 91,438 images of Chinese paintings with bilingual (Chinese/English) VLM captions, calligraphy OCR transcriptions, view-type classification, a long caption on part of the set, and foreign-object detection with one final crop box per image. Sources Component Source Images Image license npm_tw_c0–npm_tw_c3 National Palace Museum (Taipei) Open Data 86,666 Taiwan Open Government Data License v1 (attribution required) met_china… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/chinese-painting-collection.imageimage-to-text10K<n<100K0 likes1.2k downloads14d agoHugging Face03opencsg /LLaVA-Instruct-600K-Chinese 仿照 LLaVA-Instruct-150K ,使用 Qwen2.5-VL-32B-Instruct 合成的用于微调中文VLM的数据;也可以与英文数据集混合使用,训练多语言VLM 任务类型为基于单张图片的问答和对话,每个样本都对应一张不同的图片,其中大部分图片包含中文字符,更适合中文场景下视觉语言模型的训练。 图片从各类中文网站上爬取 包含3类任务:日常对话、复杂推理、描述图片。日常对话通常是5轮对话,其余任务是1轮对话。 每种任务的数量如下: 任务类型 数量 日常对话 247,431 复杂推理 194,646 描述图片 199,791 用于生成对话数据的prompt如下 日常对话 设计一个你和一个询问这张照片的人之间的对话。答案应该是视觉AI助手看到图像并回答问题的语气。 你需要提出不同的问题并给出相应的答案。问题可以包括询问图像视觉内容的问题,包括对象类型、对象计数、对象动作、对象位置、对象之间的相对位置等。必须是有明确答案的问题,即 (1) 人们可以在图像中明确看到问题所问的内容,并且可以自信地回答; (2)… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/LLaVA-Instruct-600K-Chinese.imagevisual-question-answering100K<n<1M9 likes1k downloads1y agoHugging Face04EpicZhang /ChineseTrafficRegulatorySignDataBase CTRSDB: Chinese Traffic Regulatory Sign DataBase 数据集简介 CTRSDB是聚焦中国道路场景限速、禁行、让行三类核心管制交通标志的目标检测专用数据集,专为边缘端轻量级交通目标识别模型训练优化,覆盖阴天、雨雾、信号干扰等真实复杂道路场景,完美适配YOLO系列等主流检测模型。 核心亮点 场景针对性强:聚焦自动驾驶、辅助驾驶最核心的管制类交通标志,无冗余类别,标注精度高 恶劣场景适配:通过AI生成扩增了雨雾极端天气低能见度场景数据,提升模型在复杂天气下的鲁棒性 开箱即用:原生支持YOLO格式标注,配套训练配置文件,clone后可直接用于模型训练 合规开源:遵循CC BY-NC-SA 4.0协议,仅用于学术学习与非商用场景 数据集详情 项目 详情 总图片数量 3960张 核心类别 限速、禁行通行、禁止驶入、减速让行、停车让行 标注格式 YOLO原生txt格式 数据集划分 训练集:验证集:测试集=7:2:1… See the full description on the dataset page: https://huggingface.co/datasets/EpicZhang/ChineseTrafficRegulatorySignDataBase.imageobject-detection1K<n<10K0 likes1k downloads6mo agoHugging Face05mingyy /chinese_landscape_paintings Dataset Card for "chinese_landscape_paintings" More Information needed image1K<n<10K12 likes950 downloads3y agoHugging Face06svjack /Chinese_Children_Image_Captioning_Dataset_Split0 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.image1K<n<10K0 likes652 downloads1y agoHugging Face07fritz1024 /garbage-classify-with-chineseimage0 likes582 downloads1y agoHugging Face08ReopenAI /English-Chinese-Toxic-Contentimage9 likes570 downloads2y agoHugging Face09wanng /laion-high-resolution-chinese laion-high-resolution-chinese 简介 Brief Introduction 取自Laion5B-high-resolution多语言多模态数据集中的中文部分,一共2.66M个图文对。 A subset from Laion5B-high-resolution (a multimodal dataset), around 2.66M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共2.66M个中文图文对。大约占用381MB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion-high-resolution 下载 Download mkdir release && cd release for i in {00000..00015}; do wget… See the full description on the dataset page: https://huggingface.co/datasets/wanng/laion-high-resolution-chinese.imagefeature-extraction1M<n<10M24 likes480 downloads4y agoHugging Face10seasnake /chinese-traditional-cultureimagen<1K2 likes412 downloads2y agoHugging Face11SWHL /ChineseOCRBench Chinese OCRBench 由于对于多模态LLM的OCR方向的评测集中,缺少专门中文OCR任务的评测,因此考虑专门做一个中文OCR任务的评测。 关注到On the Hidden Mystery of OCR in Large Multimodal Models工作中已经做了两个中文OCR任务的评测,于是,ChineseOCRBench仅仅是将该篇工作中提出的中文评测数据集提了出来,作为专门中文OCR评测基准。 使用方式 建议与MultimodalOCR评测脚本结合使用。 from datasets import load_dataset dataset = load_dataset("SWHL/ChineseOCRBench") test_data = dataset['test'] print(test_data[0]) # {'image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=760x1080 at 0x12544E770>… See the full description on the dataset page: https://huggingface.co/datasets/SWHL/ChineseOCRBench.image1K<n<10K27 likes403 downloads2y agoHugging Face12svjack /Chinese_Children_Image_Captioning_Dataset_Split1 CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions) CODP-1200: An AIGC based benchmark for assisting in child language acquisition 数据集介绍 目前已知最大的儿童图像描述数据集,children image captioning 共有1200张图片 每张图片对应五个中文描述,每两张图片为一组 描述文字600*5=3000 如果使用CODP-1200数据集,请引用以下文章 @article{LENG2024102627, title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition}, journal = {Displays}, volume = {82}, pages = {102627}, year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.image1K<n<10K0 likes367 downloads1y agoHugging Face13IDEA-CCNL /laion2B-multi-chinese-subset laion2B-multi-chinese-subset Github: Fengshenbang-LM Docs: Fengshenbang-Docs 简介 Brief Introduction 取自Laion2B多语言多模态数据集中的中文部分,一共143M个图文对。 A subset from Laion2B (a multimodal dataset), around 143M image-text pairs (only Chinese). 数据集信息 Dataset Information 大约一共143M个中文图文对。大约占用19GB空间(仅仅是url等文本信息,不包含图片)。 Homepage: laion-5b Huggingface: laion/laion2B-multi 下载 Download mkdir laion2b_chinese_release && cd laion2b_chinese_release for i in {00000..00012}; do… See the full description on the dataset page: https://huggingface.co/datasets/IDEA-CCNL/laion2B-multi-chinese-subset.imagefeature-extraction10M<n<100M42 likes363 downloads3y agoHugging Face14poorguys /chinese_fonts_common_128x128 Dataset Card for "chinese_fonts_common_128x128" More Information needed image100K<n<1M1 likes343 downloads2y agoHugging Face15Iess /chinese_modern_poetry 简介 数据集包括了近现代的中国诗人及外国诗人(中译版)作品,所有作品著作权归原作者所有,侵删请联系aa531811820@gmail.com chinese_poems.jsonl为原数据,training_imagery2-5_maxlen256.json 分别是根据2-5个关键意象生成诗歌的相关数据集 数据来源于网络,包括但不限于 https://github.com/sheepzh/poetry https://bedtimepoem.com/ https://poemwiki.org/ baidu、google、zhihu等 一些作品 使用此数据集训练ChatGLM、LLaMA7b模型生成的诗歌,更多诗歌查看poems目录 image100K<n<1M31 likes322 downloads3y agoHugging Face16shuangzhiaishang /Chinese-Char-Wordsimage1K<n<10K0 likes319 downloads1y agoHugging Face17Mxode /Chinese-Multimodal-Instruct 中文(视觉)多模态指令数据集 💻 Github Repo 本项目旨在构建一个高质量、大规模的中文(视觉)多模态指令数据集,目前仍在施工中 🚧💦 [!Important] 本数据集仍处于 WIP (Work in Progress) 状态,目前 Dataset Viewer 展示的是 100 条示例。 初步预计规模大约在 1~2M(不包含其他来源的数据集),均为多轮对话形式。 [!Tip] [2025/05/05] 图片已经上传完毕,后续文字部分正等待上传。 imagevisual-question-answeringn<1K11 likes279 downloads1y agoHugging Face18priyank-m /chinese_text_recognitionSource of data: https://github.com/FudanVI/benchmarking-chinese-text-recognition imageimage-to-text100K<n<1M35 likes269 downloads4y agoHugging Face19shortdramatv /chinesedramaimageimagen<1K0 likes238 downloads4mo agoHugging Face20gavinlaw /chinese-lips-speech-slide-probe Chinese-LiPS Speech + Slide Probe A self-contained probe set for testing whether visual slide context helps simultaneous speech translation — with the input as audio, not transcripts. Why audio matters: feeding a transcript to a text LLM deletes the acoustic ambiguity (homophones, polysemy) that slide context is meant to resolve; the transcript already commits to one reading. Any honest test of "does vision help streaming ST" must consume speech. Contents… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/chinese-lips-speech-slide-probe.audiotranslationn<1K0 likes237 downloads2mo agoHugging Face21poorguys /chinese_fonts_common_512x512 Dataset Card for "chinese_fonts_common_512x512" More Information needed image100K<n<1M0 likes188 downloads3y agoHugging Face22ZihCiLin /traditional-chinese-ocr-synthetic Traditional Chinese OCR Synthetic Dataset A large-scale synthetic dataset containing 4.1 million image-text pairs specifically designed for Traditional Chinese historical document recognition. Dataset Overview Existing large-scale Traditional Chinese OCR datasets (e.g., TCSynth) are primarily designed for scene text recognition, characterized by: Horizontal layouts Short text sequences (2-5 characters on average) Modern commonly-used characters These characteristics… See the full description on the dataset page: https://huggingface.co/datasets/ZihCiLin/traditional-chinese-ocr-synthetic.imageimage-to-text1M<n<10M3 likes184 downloads9mo agoHugging Face23frank-chieng /chinese_architecture_siheyuanimagen<1K3 likes175 downloads3y agoHugging Face24OpenStellarTeam /Chinese-SimpleVQA Overview 🌐 Website • 🤗 Hugging Face • ⏬ Data • 📃 Paper 中文 | English Dataset 1.chinese_simple_vqa.jsonl(the image is in url format) 2.chinese_simplevqa.parquet (the image is in base64 format and can be downloaded) Chinese SimpleVQA is the first factuality-based visual question-answering benchmark in Chinese, aimed at assessing the visual factuality of LVLMs across 8 major topics and 56 subtopics. The key features of this benchmark include a focus… See the full description on the dataset page: https://huggingface.co/datasets/OpenStellarTeam/Chinese-SimpleVQA.image1K<n<10K2 likes150 downloads2y agoHugging Face25nova-dynamics /Chinese_Commercial_Kitchen_Manipulation_Dataset_Preview 🍳 Chinese Commercial Kitchen Manipulation Dataset — Sample Pack v0.1 Asia's first real commercial kitchen manipulation dataset.Professional chef (20 years) · Real restaurant environment · Multi-view RGB-D · Egocentric video 📧 Request evaluation samples or full data: andy@dynamicnova.com Overview This sample pack contains real-world cooking demonstrations collected in an operating Chinese commercial kitchen in Zhongshan, Guangdong, China. The data focuses on… See the full description on the dataset page: https://huggingface.co/datasets/nova-dynamics/Chinese_Commercial_Kitchen_Manipulation_Dataset_Preview.imageroboticsn<1K1 likes125 downloads4mo agoHugging Face26lucasjin /chinese_ocr_llavaimage3 likes123 downloads3y agoHugging Face27AlienKevin /chinese_fontsimage100K<n<1M2 likes117 downloads1y agoHugging Face28chaeso /food_chinese_2017 Dataset Card for "food_chinese_2017" More Information needed image10K<n<100K1 likes104 downloads4y agoHugging Face29FreedomIntelligence /ALLaVA-4V-Chinese ALLaVA-4V for Chinese This is the Chinese version of the ALLaVA-4V data. We have translated the ALLaVA-4V data into Chinese through ChatGPT and instructed ChatGPT not to translate content related to OCR. The original dataset can be found here, and the image data can be downloaded from ALLaVA-4V. Citation If you find our data useful, please consider citing our work! We are FreedomIntelligence from Shenzhen Research Institute of Big Data and The Chinese University of… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ALLaVA-4V-Chinese.imagequestion-answering100K<n<1M16 likes88 downloads2y agoHugging Face30sufyCoder /ChineseBeeimagen<1K0 likes85 downloads3d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.