CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lorinma /PetrochemicalCorpora_CPTtest_200bks_zhChinese Corpora in the field of petrochemical, for the purpose of LLM continue-pretrain. 用于垂域(化工)LLM的增量预训练使用的语料,测试版。 200本书,仅经过了OCR,没有进行任何数据清理,所以质量不高。尤其是涉及到复杂的表格和公式,以及这批书的扫描质量偏低。 仅用于测试使用。 样例1: i所有安全泄压设施:如安全阀、爆破片、呼吸阀都应编号,并表示清楚设计要求; j异径管需注明其形式及规格;对改、扩建装置,版表示与已有设备或管道的连接点 (3) 仪表 a所有在线仪表,包括测量、记录、调节、分析仪表等,所有仪表均需编号; b所有调节阀; e联锁关系; d 随机仪表应在PID上注明。 (4) PID注释 h设备注释主要注明设备布置的特殊要求和催化剂、化学品和填料装卸处的空间要求 等内容; b管道注释主要注明工艺、配管方面的一些特殊要求; c仪表注释主要注明仪表安装方面的特殊要求。 3.0.10公用系统管道和仪表流程图应表示下列内容: (1)… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/PetrochemicalCorpora_CPTtest_200bks_zh.texttext-generation10K<n<100K1 likes196 downloads3y agoHugging Face02lorinma /Slim-Wildchat-zhA big shout out to AllenAI, you guys rock! 从WildChat中抽出中文对话,但是因为发现了很多重复对话,有的人会反复的用一个prompt进行提问,有的人会换3.5或4去问同样的问题,所以进行了简单的去重。 去重方法大致为,使用bert-base-chinese将第一个问题转换为embedding,使用类knn的方法抽取了1万条。并转换成了sharegpt格式。 注意!在对话中发现了NSFW的内容,并没有进行过滤,使用请注意甄别。 你会找到三个jsonl文件: wildchat-seed-multi-200.json 是使用每一个单独的Dialogue的首个HumanQuestion为基础,采样的200个种子任务,用于EvolInsturction。 Subsample_10K.jsonl 原始版本,是使用每一个单独的Dialogue的首个HumanQuestion为基础,采样的1万个对话。 1213_Wildchat_zh_Sharegpt_ConcatSubsample_20k.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/Slim-Wildchat-zh.texttext-generation10K<n<100K12 likes85 downloads3y agoHugging Face03LoRID-Math /GSM8K LoRID: A Reasoning Distillation Method via Multi-LoRA Interaction 📃 Paper • 💻 Code • 🤗 HF Repo Abstract The datasets for "Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction" [IJCAI 2025]. Key Contributions We focus on the mathematical reasoning distillation task and propose a novel method LoRID, which draws inspiration from the human beings teaching and learning… See the full description on the dataset page: https://huggingface.co/datasets/LoRID-Math/GSM8K.texttext-generation100K<n<1M1 likes78 downloads1y agoHugging Face04LoRID-Math /MATH LoRID: A Reasoning Distillation Method via Multi-LoRA Interaction 📃 Paper • 💻 Code • 🤗 HF Repo Abstract The datasets for "Can Large Models Teach Student Models to Solve Mathematical Problems Like Human Beings? A Reasoning Distillation Method via Multi-LoRA Interaction" [IJCAI 2025]. Key Contributions We focus on the mathematical reasoning distillation task and propose a novel method LoRID, which draws inspiration from the human beings teaching and learning… See the full description on the dataset page: https://huggingface.co/datasets/LoRID-Math/MATH.texttext-generation100K<n<1M1 likes49 downloads1y agoHugging Face05lorinma /Slim-LCCC-zhgated在LLM横行的今天,大家都在讲究SFT数据质量。相比于各种一板一眼的AI回复,又是step by step又是detailed reasoning,这种非常casual的对话显得那么的独特,更适合用作情感陪伴闲聊机器人的目的。 本项目提供了一个大规模中文对话数据集,原始数据来自于清华大学的LCCC(Large-scale Cleaned Chinese Conversation)数据集 基于LCCC-large,但因为有1200万。故使用bert-base-chinese转换为embedding,且使用类knn的方法抽取了1万条。并转换成了sharegpt格式。 从实用的角度来说,因为对话都只有两句,需要通过GPT进行续写。但是实测发现openai系列的太严肃了,失去了casual的味道。浅测了一下文心一言可以续写这种闲聊对话。只是测试了一下,并没有放在这个数据集中。 当然了,最好的还是收集真实世界的对话。 texttext-generation10K<n<100K12 likes36 downloads3y agoHugging Face06lorinma /Slim-Moss003sft-zh因为原生的Moss003数量太大,所以进行了简单的去重。 去重方法大致为,只选择中文的对话,使用bert-base-chinese将第一个问题转换为embedding,使用类knn的方法抽取了1万条。并转换成了sharegpt格式。 texttext-generation10K<n<100K1 likes20 downloads3y agoHugging Face07lorinma /ChemTextbookCorporaTestgated用于测试的化工训练语料,1本。来自于教科书的OCR。 样例: 第一章化工安全生产概述 化工行业是国民经济的基础行业目前中国的石油和化学工业从石油天然气等矿 产资源勘探开发到化工天然气化工煤化工盐化工国防化工化肥纯碱氯碱 电石无机盐基本有机原料农药染料涂料新领域精细化工橡胶工业新材料 等已经形成具有20多个行业可生产4万多种产品门类比较齐全品种大体配套完 整的全产业链的石化产业体系并具有一定国际竞争力 近十多年来我国化工企业发展迅速区域化工产业带已初步形成据不完全统计 截至2019年底全国重点化工园区或以石油和化工为主导的产业园区共有676家其中 国家级57家省级351家如依托长江水系形成长江经济带和长江三角洲地区上游有 重庆长寿化工园四川西部化工城下游有南京无锡常州镇江南通泰兴常 熟扬子江和苏州工业园以及上海化学工业园区依托珠江水系的珠江经济带和泛珠三 角地区主要有广东湛江茂名广州惠州深圳珠海等沿海地区的化工园区如 环杭州湾地区形成的精细化工园区山东半岛和环渤海地区的青岛齐鲁天津沧州 大连和福州湄洲湾的泉港厦门莆田等均建立了化工园区一批具有特色的内陆地区化… See the full description on the dataset page: https://huggingface.co/datasets/lorinma/ChemTextbookCorporaTest.texttext-generationn<1K1 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.