datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.tw-instruct-500k
Dataset Card for tw-instruct-500k
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為臺灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本。最新格式請改用 lianghsun/tw-instruct-500k-2511。
Dataset Details
Dataset Description
本資料集為合成資料集(synthetic dataset),由 a. reference-based 與 b. reference-free 兩種子流程組成:
reference-based:以收集自臺灣的繁中文本(用於訓練 lianghsun/Llama-3.2-Taiwan-3B 之語料)為參考,請 LLM 根據文本特性產生對應領域的指令對話。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k.tw-ipo-bilingual-vocab
Dataset Card for tw-ipo-bilingual-vocab
tw-ipo-bilingual-vocab 是一個中華民國經濟部智慧財產局(TIPO)官方網站所提供之智慧財產領域中英雙語辭彙表之整理版本,合計 1,406 筆。每筆由繁體中文名詞與對應英文翻譯組成,內容涵蓋發明專利、商標、營業秘密、新式樣、著作權等智慧財產子領域,適用於智慧財產翻譯模型之訓練或作為繁中 LLM 於專利領域之雙語預訓練素材。
Dataset Details
Dataset Description
中華民國經濟部智慧財產局(TIPO)為台灣智慧財產事務之主管機關,於其官方網站散落於各子頁面提供智慧財產領域之中英雙語辭彙表。本資料集將這些散落各處之辭彙表整合為單一 JSONL,便於下游使用。資料內容保留 TIPO 原始翻譯,未做修改。
需注意,部分「通用型名詞」之翻譯品質可能未必適合一般通用翻譯場景,且因辭彙來自不同子頁面,同一中文名詞在不同頁面可能存在不同之英文翻譯,使用者應依場景自行取捨。
Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ipo-bilingual-vocab.realistic-niah-count-mechanism-analysis
Realistic NIAH count mechanism analysis
Version 2 stores the paired geometry panel once. The default
geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200
discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds
1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now
one row rather than two duplicated mode rows.
The common row contains the passage, gold records, slots, active needle spans,
hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.twinkle_hub_finetune_dataset
twinkle_hub_finetune_dataset
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.indonesian-tongue-twisters
Pola Lidah Bahasa Indonesia 👅
Kumpulan 93 pola lidah (tongue twisters) bahasa Indonesia — dari "Kuku-kuku kaki kakekku kaku-kaku" sampai "Ular melingkar-lingkar di atas pagar".
Kenapa dataset ini ada?
Tongue twisters adalah data emas buat TTS & speech recognition (uji artikulasi) dan buat fine-tune model (pola fonetik sulit). Belum ada dataset pola lidah Indonesia di HF.
Isi
Field
Tipe
Contoh
text
string
"Kuku-kuku kaki kakekku… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-tongue-twisters.h2-harmbench-twins
H2 HarmBench Twins Dataset
This dataset contains context-coherent harmful/benign twin pairs derived from the HarmBench contextual dataset for testing the "Consistency Confound" hypothesis in semantic entropy-based jailbreak detection.
Dataset Description
Total Samples: 162 (81 harmful + 81 benign twin pairs)Source: Generated from walledai/HarmBench contextual splitPurpose: Testing semantic entropy effectiveness for jailbreak detection across matched harmful/benign content… See the full description on the dataset page: https://huggingface.co/datasets/DhruvTre/h2-harmbench-twins.TwinnyAI-Personas-Dataset
Overview
The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems.
Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.tw_ideology
Taiwan Reward Model Dataset
