CoolFace
20 results

chinese

opencsg /Fineweb-Edu-Chinese-V2.1 Chinese Fineweb Edu Dataset V2.1 [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report The Chinese Fineweb Edu Dataset V2.1 is an enhanced version of the V2 dataset, designed specifically for natural language processing (NLP) tasks in the education sector. This version introduces two new data sources, map-cc and opencsg-cc, and retains data with scores ranging from 2 to 3. The dataset entries are organized into different… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.1.text-generation10B<n<100B80 likes55k downloads8mo agoHugging Faceopencsg /chinese-fineweb-edu This version is deprecated. We recommend you to use the newest version Fineweb-edu-chinese-v2.1 ! Chinese Fineweb Edu Dataset [中文] [English] [OpenCSG Community] [👾github] [wechat] [Twitter] 📖Technical Report Chinese Fineweb Edu dataset is a meticulously constructed high-quality Chinese pre-training corpus, specifically designed for natural language processing tasks in the education domain. This dataset undergoes a rigorous selection and… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/chinese-fineweb-edu.texttext-generation10M<n<100M117 likes17k downloads10mo agoHugging Faceopencsg /Fineweb-Edu-Chinese-V2.2 Chinese Fineweb Edu Dataset V2.2 (Instruct & Pre-train) [[中文]] | [[English]] OpenCSG Community | 👾 GitHub | 📖 Technical Report Dataset Introduction: Filling the Data Puzzle for Chinese Education LLMs Chinese Fineweb Edu Dataset V2.2is a rare high-quality dataset in the open-source community that covers the full process from Pre-training to Supervised Fine-Tuning (SFT) for the Chinese education domain. This project aims to solve the core pain point of… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/Fineweb-Edu-Chinese-V2.2.text-generation10B<n<100B83 likes17k downloads8mo agoHugging FaceFreedomIntelligence /MMLU_ChineseChinese version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. 2 likes15k downloads3y agoHugging Face86Cao /google-landmark-v2-chinese-filtered Google Landmark V2 Chinese Filtered Dataset This dataset contains landmark images and metadata for training landmark retrieval models, with Chinese translations of landmark names to facilitate Chinese multimodal retrieval tasks. Dataset Source This dataset is based on the Google Landmarks V2 dataset from Kaggle. The original data has been filtered and processed to create a high-quality training dataset for landmark retrieval. Key Features Filtered… See the full description on the dataset page: https://huggingface.co/datasets/86Cao/google-landmark-v2-chinese-filtered.text-to-image100K<n<1M0 likes6.5k downloads9mo agoHugging Facesilk-road /alpaca-data-gpt4-chinesetexttext-generation10K<n<100K104 likes6.1k downloads3y agoHugging Face