datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TwinRouterBench
TwinRouterBench Static
This dataset contains the static track for TwinRouterBench: Fast Static and Live Dynamic Evaluation for Realistic Agentic LLM Routing.
Paper: arXiv:2605.18859
Contents
data/train.parquet: Hugging Face viewer-friendly table with 970 rows. Nested fields such as messages and optional tool/function schemas are stored as JSON strings so all benchmark sources share a stable schema.
question_bank.jsonl: the original static question bank exported by… See the full description on the dataset page: https://huggingface.co/datasets/Amorph/TwinRouterBench.tw-instruct-500k-Q-R1
tw-instruct-500k-Q-R1
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。
Dataset Details
Dataset Description
這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成 reference-free 時,則是以常見的種子提示(seed prompts)作為參考,讓 LLM… See the full description on the dataset page: https://huggingface.co/datasets/NLTF-mock/tw-instruct-500k-Q-R1.tw-instruct-500k-QA-R1
tw-instruct-500k-QA-R1
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。
Dataset Details
Dataset Description
這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成 reference-free 時,則是以常見的種子提示(seed prompts)作為參考,讓… See the full description on the dataset page: https://huggingface.co/datasets/NLTF-mock/tw-instruct-500k-QA-R1.tw-instruct-500k-Q-R1
tw-instruct-500k-Q-R1
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。
Dataset Details
Dataset Description
這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成… See the full description on the dataset page: https://huggingface.co/datasets/c00cjz00/tw-instruct-500k-Q-R1.tw-instruct-500k-QA-R1
tw-instruct-500k-QA-R1
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為台灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本,進行理解力(Reasoning)資料補充生成。
Dataset Details
Dataset Description
這個資料集為合成資料集(synthetic datasets),內容由 a. reference-based 和 b. reference-free 的子資料集組合而成。生成 reference-based 資料集時,會先以我們收集用來訓練 lianghsun/Llama-3.2-Taiwan-3B 時的繁體中文文本作為參考文本,透過 LLM 去生成指令對話集,如果參考文本有特別領域的問法,我們將會特別設計該領域或者是適合該文本的問題;生成… See the full description on the dataset page: https://huggingface.co/datasets/c00cjz00/tw-instruct-500k-QA-R1.
