firefly
Datasets
All datasets matching “firefly”firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4
如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。
我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示:
每条数据的格式如下,包含任务类型、输入、目标输出:
{
"kind": "ClassicalChinese",
"input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。",
"target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。"
}
训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600:
firefly-pretrain-dataset
Firefly中文Llama2增量预训练数据
欢迎加入Firefly大模型技术交流群,关注我们的公众号。
数据简介
技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践
该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。
模型列表 & 数据列表
我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩充中文词表后增量预训练得到的模型,Chat模型是在Base模型的基础上进行多轮对话指令微调。
为了探究基座模型对指令微调的影响,我们也微调了baichuan2-base模型,获得firefly-baichuan2-13b,具有不错的效果。更多中文微调,可查看Firefly项目。
模型
类型
训练任务
训练长度
🤗Firefly-LLaMA2-7B-Base
基座模型… See the full description on the dataset page: https://huggingface.co/datasets/YeungNLP/firefly-pretrain-dataset.details_YeungNLP__firefly-mixtral-8x7b-v0.1
Dataset Card for Evaluation run of YeungNLP/firefly-mixtral-8x7b-v0.1
Dataset automatically created during the evaluation run of model YeungNLP/firefly-mixtral-8x7b-v0.1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-mixtral-8x7b-v0.1.details_YeungNLP__firefly-llama2-13b-chat
Dataset Card for Evaluation run of YeungNLP/firefly-llama2-13b-chat
Dataset Summary
Dataset automatically created during the evaluation run of model YeungNLP/firefly-llama2-13b-chat on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-llama2-13b-chat.firefly-train-chinese-zhtw
Dataset Card for "firefly-train-chinese-zhtw"
資料集摘要
本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。
[Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。
資料量為115萬 。數據分佈如下圖所示:
訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600:
原始資料來源:
YeungNLP/firefly-train-1.1M
Firefly(流萤): 中文对话式大语言模型
資料下載清理
下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo
使用 OpenCC 來進行簡繁轉換
使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.details_YeungNLP__firefly-zephyr-6x7b
Dataset Card for Evaluation run of YeungNLP/firefly-zephyr-6x7b
Dataset automatically created during the evaluation run of model YeungNLP/firefly-zephyr-6x7b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_YeungNLP__firefly-zephyr-6x7b.
