datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.gpt-oss-eval-logs-and-scores
This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
llama-4-eval-logs-and-scores
Dataset Card for llama-4-eval-logs-and-scores
This repository contains the detailed evaluation results of Llama 4 models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
Dataset Details
Dataset Description
This dataset provides the complete evaluation logs and per-question scores of various Llama 4 models, including Scout and… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/llama-4-eval-logs-and-scores.gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgemma-3-taide-12b-chat-eval-logs-and-scoresNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresministral-14b-eval-logs-and-scoresgemma-3-4b-it-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresLlama-Breeze2-8B-Instruct-eval-logs-and-scoresnemotron-nano-eval-logs-and-scoresgemma-3-4B-T1-it-eval-logs-and-scoresdevstral-eval-logs-and-scoresllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.3-70B-Instruct-eval-logs-and-scoresGemma-3-12b-it-eval-logs-and-scoresphi-4-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresDevstral-Small-2505-eval-logs-and-scoresgemma-3-27b-it-eval-logs-and-scoresmistral-675b-eval-logs-and-scorestw-instruct-500k
Dataset Card for tw-instruct-500k
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為臺灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本。最新格式請改用 lianghsun/tw-instruct-500k-2511。
Dataset Details
Dataset Description
本資料集為合成資料集(synthetic dataset),由 a. reference-based 與 b. reference-free 兩種子流程組成:
reference-based:以收集自臺灣的繁中文本(用於訓練 lianghsun/Llama-3.2-Taiwan-3B 之語料)為參考,請 LLM 根據文本特性產生對應領域的指令對話。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k.twin-uniref50-faiss
Twin-Model UniRef50 FAISS Index
FAISS index over Twin model mean-pooled embeddings of all UniRef50
representative proteins (~49.8M). The Twin model is a two-tower contrastive
encoder fine-tuned on Resnik GO similarity:
Custom tower: AA-vocab Transformer → padding-masked mean pool → MLP → 512-dim
ESM tower: facebook/esm2_t33_650M_UR50D (frozen) → masked mean pool → MLP → 512-dim
Output: concat(custom, esm) → 1024-dim
Checkpoint:… See the full description on the dataset page: https://huggingface.co/datasets/genomenet/twin-uniref50-faiss.twinkle_hub_finetune_dataset
twinkle_hub_finetune_dataset
MCP tool-calling SFT 資料集,由 Agent Tools Fine-Tuning Platform 以「反向生成 + teacher solver 驗證」流程產生。
語言:繁體中文
工具(來自 MCP server):search_datasets, get_dataset, query_rows, materialize_dataset, search_patents, get_patent_body, search_exam, search_exam_questions, get_exam_paper, search_teacher_exam, search_teacher_exam_questions, get_teacher_exam_paper, search_teacher_recruit, search_teacher_recruit_questions, get_teacher_recruit_paper, search_taiwan_md… See the full description on the dataset page: https://huggingface.co/datasets/Simon-Liu/twinkle_hub_finetune_dataset.F1-identity
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/F1-identity.h2-harmbench-twins
H2 HarmBench Twins Dataset
This dataset contains context-coherent harmful/benign twin pairs derived from the HarmBench contextual dataset for testing the "Consistency Confound" hypothesis in semantic entropy-based jailbreak detection.
Dataset Description
Total Samples: 162 (81 harmful + 81 benign twin pairs)Source: Generated from walledai/HarmBench contextual splitPurpose: Testing semantic entropy effectiveness for jailbreak detection across matched harmful/benign content… See the full description on the dataset page: https://huggingface.co/datasets/DhruvTre/h2-harmbench-twins.TwinnyAI-Personas-Dataset
Overview
The TWINNY.AI Personas Dataset is a synthetic collection of 400 richly structured professional personas, engineered to power behavioral AI twins, persona-driven language model fine-tuning, and professional simulation systems.
Each persona is built from 14 attributes spanning demographics, professional context, behavioral psychology, and communication style sampled with realistic non-uniform distributions that mirror actual workforce demographics rather than uniform… See the full description on the dataset page: https://huggingface.co/datasets/Mostafa190/TwinnyAI-Personas-Dataset.
