CoolFace
20 results

firefly

YeungNLP /firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4 如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。 我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示: 每条数据的格式如下,包含任务类型、输入、目标输出: { "kind": "ClassicalChinese", "input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。", "target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。" } 训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600: text1M<n<10M346 likes8.6k downloads3y agoHugging FaceYeungNLP /firefly-pretrain-dataset Firefly中文Llama2增量预训练数据 欢迎加入Firefly大模型技术交流群,关注我们的公众号。 数据简介 技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践 该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。 模型列表 & 数据列表 我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩充中文词表后增量预训练得到的模型,Chat模型是在Base模型的基础上进行多轮对话指令微调。 为了探究基座模型对指令微调的影响,我们也微调了baichuan2-base模型,获得firefly-baichuan2-13b,具有不错的效果。更多中文微调,可查看Firefly项目。 模型 类型 训练任务 训练长度 🤗Firefly-LLaMA2-7B-Base 基座模型… See the full description on the dataset page: https://huggingface.co/datasets/YeungNLP/firefly-pretrain-dataset.text1M<n<10M42 likes563 downloads3y agoHugging Faceopen-llm-leaderboard-old /details_YeungNLP__firefly-mixtral-8x7b-v0.1 Dataset Card for Evaluation run of YeungNLP/firefly-mixtral-8x7b-v0.1 Dataset automatically created during the evaluation run of model YeungNLP/firefly-mixtral-8x7b-v0.1 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-mixtral-8x7b-v0.1.0 likes158 downloads3y agoHugging Faceopen-llm-leaderboard-old /details_YeungNLP__firefly-llama2-13b-chat Dataset Card for Evaluation run of YeungNLP/firefly-llama2-13b-chat Dataset Summary Dataset automatically created during the evaluation run of model YeungNLP/firefly-llama2-13b-chat on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-llama2-13b-chat.0 likes156 downloads3y agoHugging Faceerhwenkuo /firefly-train-chinese-zhtw Dataset Card for "firefly-train-chinese-zhtw" 資料集摘要 本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。 [Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。 資料量為115萬 。數據分佈如下圖所示: 訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600: 原始資料來源: YeungNLP/firefly-train-1.1M Firefly(流萤): 中文对话式大语言模型 資料下載清理 下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo 使用 OpenCC 來進行簡繁轉換 使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.texttext-generation1M<n<10M2 likes149 downloads3y agoHugging Faceopen-llm-leaderboard-old /details_YeungNLP__firefly-zephyr-6x7b Dataset Card for Evaluation run of YeungNLP/firefly-zephyr-6x7b Dataset automatically created during the evaluation run of model YeungNLP/firefly-zephyr-6x7b on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-zephyr-6x7b.0 likes143 downloads3y agoHugging Face