CoolFace
Datasetpublicgated

NLTF-mock/tw-instruct-500k-QA-R1

tw-instruct-500k-QA-R1 台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。 Dataset Details Dataset Description 這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成 reference-free 時,則是以常見的種子提示(seed… See the full description on the dataset page: https://huggingface.co/datasets/NLTF-mock/tw-instruct-500k-QA-R1.

sourceHugging Facecc-by-nc-sa-4.0updated 2y agoView on Hugging Face
0likes4downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
NLTF-mock/tw-instruct-500k-QA-R1 · CoolFace