datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Gath_baizedetails_project-baize__baize-v2-13b
Dataset Card for Evaluation run of project-baize/baize-v2-13b
Dataset Summary
Dataset automatically created during the evaluation run of model project-baize/baize-v2-13b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_project-baize__baize-v2-13b.Baize-TCM-Corpus-for-Large-Language-Models-V3
白泽中医药大模型语料库
版本:3.0语料数量:157,438 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 157,438 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V3.baize-metadataMedQuaAD-Italian-Fauno-Baize
MedQuaAD-Italian-Fauno-Baize
This dataset is an Italian translation of the MedQuaAD dataset presented by Baize's authors.
Languages
Italian
Dataset Structure
Data Instances
Sentences 46,867
average number of turns 3.8
response lengths of each turn 35.8
Data Fields
topic, input
Data Splits
Train
Dataset Creation
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/andreabac3/MedQuaAD-Italian-Fauno-Baize.alpaca_baizebaize_chatbot(https://github.com/project-baize/baize-chatbot/tree/main/data)
StackOverflow-Italian-Fauno-Baize
StackOverflow-Italian-Fauno-Baize
This dataset is an Italian translation of the StackOverflow dataset presented by Baize's authors.
Languages
Italian
Dataset Structure
Data Instances
Sentences 57,046
average number of turns 3.6
response lengths of each turn 36.0
Data Fields
topic, input
Data Splits
Train
Dataset Creation
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/andreabac3/StackOverflow-Italian-Fauno-Baize.details_TheBloke__Project-Baize-v2-7B-GPTQ
Dataset Card for Evaluation run of TheBloke/Project-Baize-v2-7B-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Project-Baize-v2-7B-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Project-Baize-v2-7B-GPTQ.Quora-Italian-Fauno-Baize
Quora-Italian-Fauno-Baize
This dataset is an Italian translation of the Quora dataset presented by Baize's authors.
Languages
Italian
Dataset Structure
Data Instances
Sentences 54,456
average number of turns 3.9
response lengths of each turn 35.9
Data Fields
topic, input
Data Splits
Train
Dataset Creation
Source Data
Initial Data Collection and Normalization… See the full description on the dataset page: https://huggingface.co/datasets/andreabac3/Quora-Italian-Fauno-Baize.Haruhi-Baize-Role-Playing-Conversation
Haruhi-Zero的Conversation训练数据
我们计划拓展ChatHaruhi,从Few-shot到Zero-shot,这个数据集记录使用各个(中文)角色扮演api进行Baize式相互聊天后得到的数据结果
ids代表聊天的时候两张bot的角色卡片, 角色卡片的信息可以在https://huggingface.co/datasets/silk-road/Haruhi-Zero-RolePlaying-movie-PIPPA 中找到
并且对于第一次出现的id0,也会在prompt字段中进行记录。
聊天的时候id和ids的卡片进行对应
openai 代表两个聊天的bot都使用openai
GLM 代表两个聊天的bot都使用CharacterGLM
Claude 代表两个聊天的bot都使用Claude
Claude_openai 代表id0的使用Claude, id1的使用openai
Baichuan 代表两个聊天的bot都使用Character-Baichuan-Turbo… See the full description on the dataset page: https://huggingface.co/datasets/silk-road/Haruhi-Baize-Role-Playing-Conversation.details_project-baize__baize-healthcare-lora-7B
Dataset Card for Evaluation run of project-baize/baize-healthcare-lora-7B
Dataset Summary
Dataset automatically created during the evaluation run of model project-baize/baize-healthcare-lora-7B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_project-baize__baize-healthcare-lora-7B.Baize-TCM-Corpus-for-Large-Language-Models-V2
白泽中医药大模型语料库
版本:2.0语料数量:10.578 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 10,578 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V2.tulu_baizedetails_TheBloke__Project-Baize-v2-13B-GPTQ
Dataset Card for Evaluation run of TheBloke/Project-Baize-v2-13B-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Project-Baize-v2-13B-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 60 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Project-Baize-v2-13B-GPTQ.baizedetails_project-baize__baize-v2-7b
Dataset Card for Evaluation run of project-baize/baize-v2-7b
Dataset Summary
Dataset automatically created during the evaluation run of model project-baize/baize-v2-7b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_project-baize__baize-v2-7b.Baize-TCM-Corpus-for-Large-Language-Models-V1
白泽中医药大模型语料库
版本:1.0语料数量:4,735 条语言:中文领域:中医药(Traditional Chinese Medicine, TCM)格式:问答对(QA Pair)用途:中医药大模型训练、知识问答系统、语义理解研究
📚 简介
“白泽中医药大模型语料库”是一个专注于中医药领域的高质量问答语料集合,旨在支持中医药知识的数字化、智能化应用。语料库共包含 4,735 条 经过整理与校对的问答对,涵盖中医基础理论、中药学、方剂学、诊断学、针灸推拿、经典医籍、临床实践等多个子领域。
本语料库可广泛应用于:
中医药大语言模型的预训练与微调
智能问答系统开发
医学自然语言处理任务(如实体识别、关系抽取)
中医药知识图谱构建
🧩 数据内容
每条语料为一个标准的问答对,格式如下:
{
"instruction": "广义转录组和狭义转录组在定义上的主要区别是什么?",
"input": "",
"output":… See the full description on the dataset page: https://huggingface.co/datasets/DigitalIntelligenceCenter-of-ICMM/Baize-TCM-Corpus-for-Large-Language-Models-V1.baize
Baize dataset
This is an unofficial version of the data used to train the Baize chatbot.
now in ShareGPT-like format
with basic string cleaning and turn ordering
According to the original authors:
Baize is an open-source chat model trained with LoRA. It uses 100k dialogs generated by letting ChatGPT chat with itself. We also use Alpaca's data to improve its performance. We have released 7B, 13B and 30B models. Please refer to the paper for more details.
Licence
GNU… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/baize.baize-quoraBaize has scrapped questions from Quora. The dialogs are generated by letting ChatGPT chat with itself.
This dataset is in alpaca format.
baize-chat-data
Dataset Description
Original Repository: https://github.com/project-baize/baize-chatbot/tree/main/data
This is a dataset of the training data used to train the Baize family of models. This dataset is used for instruction fine-tuning of LLMs, particularly in "chat" format. Human and AI messages are marked by [|Human|] and [|AI|] tags respectively. The data from the orignial repo consists of 4 datasets (alpaca, medical, quora, stackoverflow), and this dataset combines all four into… See the full description on the dataset page: https://huggingface.co/datasets/linkanjarad/baize-chat-data.chitchat_baizeGameClip-TimeBenchaurora-mix-data-baize-formatGath_baize_mod
