datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
claude_multiround_chat_30kThis dataset is the result of 50k instruction/response pairs generated by Claude and two additional follow-up instructions for each base instruction (for a total of 150k instructions), with instances of blatant alignment removed.
32170 (96510) instructions remain.
The instructions were generated synethically using a method that can be tenatively described as "multi-instruct." These instructions consist of numerous discrete tasks that the AI has to work its way through, thereby hopefully… See the full description on the dataset page: https://huggingface.co/datasets/Norquinal/claude_multiround_chat_30k.multiround-programming-convo
Multi-Round Programming Conversations
Based on previous evol-codealpaca-v1 dataset with added sampled questions from stackoverflow, crossvalidated and make it multiround!
It should be more suited to train a code assistant which works side by side.
Tasks included in here:
Data science, statistic, programming questions
Code translation : translate a short function from Python, Golang, C++, Java, Javascript
Code fixing : Fix randomly corrupts characters with no tab… See the full description on the dataset page: https://huggingface.co/datasets/theblackcat102/multiround-programming-convo.MultiRoundConvos-Code-JS-HTML-CSS-PythonAI-AgentsTools-GPT-Multiround-Conversationclaude_multiround_chat_1kThis dataset is ~1k random samplings from my claude_multiround_chat_30k dataset.
The instructions were generated synethically using a method that can be tenatively described as "multi-instruct." These instructions consist of numerous discrete tasks that the AI has to work its way through, thereby hopefully increasing its comprehension and awareness of complex instructions.
The topics of the instruction ranged from STEM, Arts & Humanities, Social Knowledge, and General Knowledge.
Norquinal_claude_multiround_chat_30k-SlopOnly-KTOSloPreferenceShareGPTbelle-multiround
Dataset Card for belle-multiround
本資料集為以 BELLE 系列指令資料為基礎,整理/轉寫之多輪對話(multi-round) 繁體中文版本,可作為繁中對話模型在多輪互動上的補強資料。
Dataset Details
Dataset Description
BELLE 是早期具規模的中文指令資料集系列。其原始版本以簡體中文為主、且多為單輪指令對話。本資料集做了兩件事:
多輪化:以 BELLE 的單輪指令為起點,請 LLM 生成自然延伸的後續輪次(追問、澄清、延伸要求),形成 multi-round 結構。
繁中化:將內容轉寫為繁體中文,並調整在地用語。
可用於補強模型在多輪對話中的脈絡保持能力。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese
License: cc-by-nc-sa-4.0
Dataset Sources
Repository:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/belle-multiround.
