CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01twinkle-ai /tw-leetcode Dataset Card for tw-leetcode A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily. Dataset Details Dataset Description tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。 本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.texttext-generationn<1K18 likes691 downloads10h agoHugging Face02lianghsun /tw-instruct-500k Dataset Card for tw-instruct-500k [👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉] 台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為臺灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本。最新格式請改用 lianghsun/tw-instruct-500k-2511。 Dataset Details Dataset Description 本資料集為合成資料集(synthetic dataset),由 a. reference-based 與 b. reference-free 兩種子流程組成: reference-based:以收集自臺灣的繁中文本(用於訓練 lianghsun/Llama-3.2-Taiwan-3B 之語料)為參考,請 LLM 根據文本特性產生對應領域的指令對話。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k.texttext-generation100K<n<1M27 likes69 downloads5mo agoHugging Face03lianghsun /tw-ipo-bilingual-vocab Dataset Card for tw-ipo-bilingual-vocab tw-ipo-bilingual-vocab 是一個中華民國經濟部智慧財產局(TIPO)官方網站所提供之智慧財產領域中英雙語辭彙表之整理版本,合計 1,406 筆。每筆由繁體中文名詞與對應英文翻譯組成,內容涵蓋發明專利、商標、營業秘密、新式樣、著作權等智慧財產子領域,適用於智慧財產翻譯模型之訓練或作為繁中 LLM 於專利領域之雙語預訓練素材。 Dataset Details Dataset Description 中華民國經濟部智慧財產局(TIPO)為台灣智慧財產事務之主管機關,於其官方網站散落於各子頁面提供智慧財產領域之中英雙語辭彙表。本資料集將這些散落各處之辭彙表整合為單一 JSONL,便於下游使用。資料內容保留 TIPO 原始翻譯,未做修改。 需注意,部分「通用型名詞」之翻譯品質可能未必適合一般通用翻譯場景,且因辭彙來自不同子頁面,同一中文名詞在不同頁面可能存在不同之英文翻譯,使用者應依場景自行取捨。 Curated by:… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ipo-bilingual-vocab.texttranslation1K<n<10K0 likes36 downloads5mo agoHugging Face04twistshan /realistic-niah-count-mechanism-analysis Realistic NIAH count mechanism analysis Version 2 stores the paired geometry panel once. The default geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200 discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds 1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now one row rather than two duplicated mode rows. The common row contains the passage, gold records, slots, active needle spans, hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.tabulartext-generationn<1K0 likes28 downloads1mo agoHugging Face05Simon-Liu /twinkle_hub_finetune_dataset twinkle_hub_finetune_dataset MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。 語言:繁體中文 工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.texttext-generation1K<n<10K1 likes23 downloads2mo agoHugging Face06LorthGyu /indonesian-tongue-twisters Pola Lidah Bahasa Indonesia 👅 Kumpulan 93 pola lidah (tongue twisters) bahasa Indonesia — dari "Kuku-kuku kaki kakekku kaku-kaku" sampai "Ular melingkar-lingkar di atas pagar". Kenapa dataset ini ada? Tongue twisters adalah data emas buat TTS & speech recognition (uji artikulasi) dan buat fine-tune model (pola fonetik sulit). Belum ada dataset pola lidah Indonesia di HF. Isi Field Tipe Contoh text string "Kuku-kuku kaki kakekku… See the full description on the dataset page: https://huggingface.co/datasets/LorthGyu/indonesian-tongue-twisters.texttext-generationn<1K0 likes17 downloads2mo agoHugging Face07DhruvTre /h2-harmbench-twins H2 HarmBench Twins Dataset This dataset contains context-coherent harmful/benign twin pairs derived from the HarmBench contextual dataset for testing the "Consistency Confound" hypothesis in semantic entropy-based jailbreak detection. Dataset Description Total Samples: 162 (81 harmful + 81 benign twin pairs)Source: Generated from walledai/HarmBench contextual splitPurpose: Testing semantic entropy effectiveness for jailbreak detection across matched harmful/benign content… See the full description on the dataset page: https://huggingface.co/datasets/DhruvTre/h2-harmbench-twins.texttext-generationn<1K0 likes8 downloads1y agoHugging Face08Mostafa190 /TwinnyAI-Personas-Datasetgated Overview The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems. Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.texttext-generation10K<n<100K1 likes7 downloads6mo agoHugging Face09kartd /tw_ideologygated Taiwan Reward Model Dataset texttext-generation10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.