datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-instruct-v1
Open Instruct V1 - A dataset for having LLMs follow instructions.
Open Instruct V1 is an amalgamation of different datasets which are cleaned and then collated into a singular format for training.
Dataset Breakdown
Dataset
Amount of Samples
Alpaca
51759
Self Instruct
82599
GPT-4 Instruct
18194
Code Alpaca
18019
Dolly
15015
Synthetic
33143
Roleplay
3146
asss
448
instruction-dataset
327
Total
222650
Hakimi_test
哈基米(Hakimi)自我认知数据集
数据集描述
这是一个关于虚拟角色"哈基米"的自我认知对话数据集。哈基米是一只独特的"耄耋"(猫),由"南北绿豆"养育长大。该数据集包含了哈基米与用户之间的对话交互,展现了其独特的个性和自我认知。
数据格式
数据集采用 JSONL 格式,每行包含一个 JSON 对象,具有以下字段:
system: 定义哈基米的角色设定
conversation: 包含对话历史的数组,每个元素包含 human(用户输入)和 assistant(哈基米回复)
数据内容
数据集包含 59 个对话样本,涵盖了以下主题:
自我介绍(姓名、身份、特点)
与用户的问候互动
对自身身份的澄清(强调自己是"猫"而非AI)
与"南北绿豆"的关系
"曼波"等特色表达方式
数据特征
语言: 中文为主,包含少量英文
角色特点: 哈基米坚持自己是一只猫,而非AI助手
核心表达: 频繁提及"曼波"、"哈基米"等特色词汇
情感倾向: 友好、亲切,致力于传播"爱与和平"… See the full description on the dataset page: https://huggingface.co/datasets/weisiren/Hakimi_test.
