datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ChildMandarin
ChildMandarin: A Comprehensive Mandarin Speech Dataset for Young Children Aged 3-5
Introduction
ChildMandarin is a comprehensive, open-source Mandarin Chinese speech dataset specifically designed for research on young children aged 3 to 5. This dataset addresses the critical lack of publicly available resources for this age group, enabling advancements in automatic speech recognition (ASR), speaker verification (SV), and other related fields. The dataset is released… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/ChildMandarin.illustrations_for_children
数据集 README
数据集概述
欢迎使用我们的数据集,该数据集主要包含网络收集的儿童插画(儿插)。这些插画旨在为教育和研究目的提供丰富的视觉素材。我们鼓励用户在遵守本README中规定的条款和条件的前提下,充分利用这些资源进行学习和研究。
许可协议
本数据集遵循Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0)许可协议。这意味着您可以:
自由分享:复制和分发数据集中的材料。
自由改编:基于本数据集的材料进行修改和再创作。
但请注意以下限制:
非商业性:您不得将本数据集用于商业目的。
相同方式共享:如果您对数据集进行了修改或衍生,您必须以相同的许可协议分发您的作品。
署名:您必须给出适当的署名,提供许可协议链接,并说明是否进行了更改。您可以以任何合理的方式进行署名,但不得以任何方式暗示许可人认可您或您的使用。
使用限制… See the full description on the dataset page: https://huggingface.co/datasets/shiertier/illustrations_for_children.sd2hd_images_v2ChineseConversationEmoticillustrations_for_children_1024
数据集 README
数据集概述
欢迎使用我们的数据集,该数据集主要包含网络收集的儿童插画(儿插)。这些插画旨在为教育和研究目的提供丰富的视觉素材。我们鼓励用户在遵守本README中规定的条款和条件的前提下,充分利用这些资源进行学习和研究。
图像被处理为接近1024**2的分辨率大小(不放大)。
许可协议
本数据集遵循Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0)许可协议。这意味着您可以:
自由分享:复制和分发数据集中的材料。
自由改编:基于本数据集的材料进行修改和再创作。
但请注意以下限制:
非商业性:您不得将本数据集用于商业目的。
相同方式共享:如果您对数据集进行了修改或衍生,您必须以相同的许可协议分发您的作品。
署名:您必须给出适当的署名,提供许可协议链接,并说明是否进行了更改。您可以以任何合理的方式进行署名,但不得以任何方式暗示许可人认可您或您的使用。
使用限制… See the full description on the dataset page: https://huggingface.co/datasets/shiertier/illustrations_for_children_1024.Chinese-Dialogue-180k-Instruct-AudioDanbooru2024-Webp-4MPixel-NL
📊 Dataset Overview
The Danbooru2024-Webp-4MPixel-NL dataset is an extension of the deepghs/danbooru2024-webp-4Mpixel collection, specifically curated to provide natural language descriptions for approximately 7.8 million high-quality images sourced from the official Danbooru platform. Each image is paired with a detailed textual description generated using the fancyfeast/llama-joycaption-alpha-two-hf-llava model. The dataset features a filtering rule ensuring only images with an… See the full description on the dataset page: https://huggingface.co/datasets/chinoll/Danbooru2024-Webp-4MPixel-NL.chicago-police-scannerhuawei-zhuofan-chineseML2025_HW_Cardiac_Muscle
ML 2025 Cardiac Muscle Dataset
You have to extract the file before using the dataset.
new-childchinese_celeb_dataset
