datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
firefly-train-chinese-zhtw
Dataset Card for "firefly-train-chinese-zhtw"
資料集摘要
本資料集主要是應用於專案:Firefly(流螢): 中文對話式大語言模型 ,經過訓練後得到的模型 firefly-1b4。
[Firefly(流螢): 中文對話式大語言模型]專案(https://github.com/yangjianxin1/Firefly)收集了 23 個常見的中文資料集,并且對於每種不同的 NLP 任務,由人工書寫若干種指令模板來保證資料的高品質與豐富度。
資料量為115萬 。數據分佈如下圖所示:
訓練資料集的 token 長度分佈如下圖所示,絕大部分資料的長度都小於 600:
原始資料來源:
YeungNLP/firefly-train-1.1M
Firefly(流萤): 中文对话式大语言模型
資料下載清理
下載 chinese-poetry: 最全中文诗歌古典文集数据库 的 Repo
使用 OpenCC 來進行簡繁轉換
使用 Huggingface Datasets 來上傳至… See the full description on the dataset page: https://huggingface.co/datasets/erhwenkuo/firefly-train-chinese-zhtw.Firefly-1.1M-Rephrased
