datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-leetcode
Dataset Card for tw-leetcode
A curated Traditional Chinese LeetCode solution dataset with high-efficiency answers (Beats 100%), structured explanation in "Top Concept → Step Implement → Complexity Analysis" style, updated daily.
Dataset Details
Dataset Description
tw-leetcode 是一個針對 LeetCode 題目的繁體中文資料集,內容包含高效能程式解法、完整的解題思路,以及時間與空間複雜度分析。每份題解都經由人工清洗與優化,並依循「Top Concept → Step Implement → Complexity Explanation」的結構撰寫,方便機器學習模型或人類讀者理解程式邏輯的推理過程。
本資料集適合作為:… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-leetcode.gpt-oss-eval-logs-and-scores
This repository contains the detailed evaluation results of gpt-oss models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
llama-4-eval-logs-and-scores
Dataset Card for llama-4-eval-logs-and-scores
This repository contains the detailed evaluation results of Llama 4 models, tested using Twinkle Eval, a robust and efficient AI evaluation tool developed by Twinkle AI. Each entry includes per-question scores across multiple benchmark suites.
Dataset Details
Dataset Description
This dataset provides the complete evaluation logs and per-question scores of various Llama 4 models, including Scout and… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/llama-4-eval-logs-and-scores.gpt-oss-120b-mandarin-thinking-eval-logs-and-scoresgemma-3-taide-12b-chat-eval-logs-and-scoresNVIDIA-Nemotron-3-Super-120B-A12B-FP8-eval-logs-and-scoresministral-14b-eval-logs-and-scoresgemma-3-4b-it-eval-logs-and-scoresgpt-oss-20b-mandarin-thinking-eval-logs-and-scoresLlama-Breeze2-8B-Instruct-eval-logs-and-scoresnemotron-nano-eval-logs-and-scoresgemma-3-4B-T1-it-eval-logs-and-scoresdevstral-eval-logs-and-scoresllama-3.2-3B-f1-instruct-eval-logs-and-scoresLlama-3.1-8B-Instruct-eval-logs-and-scoresLlama-3.3-70B-Instruct-eval-logs-and-scoresGemma-3-12b-it-eval-logs-and-scoresphi-4-eval-logs-and-scoresLlama-3-Taiwan-70B-Instruct-eval-logs-and-scoresLlama-3.2-3B-Instruct-eval-logs-and-scoresLlama-3.1-Taiwan-8B-Instruct-eval-logs-and-scoresDevstral-Small-2505-eval-logs-and-scoresgemma-3-27b-it-eval-logs-and-scoresmistral-675b-eval-logs-and-scorescyber_twist_the_tube_v0.1
CyberOrigin Dataset
Our data includes information from home services, the logistics industry, and laboratory scenarios.
For more details, please refer to our Offical Data Website
contents of dataset:
cyber_twist_the_tube # dataset root path
└── data/
├── metadata_ID1_240808.json
├── segment_ids_ID1_240808.bin # for each frame segment_ids uniquely points to the segment index that frame i came from. You may want to use this to separate non-contiguous frames from… See the full description on the dataset page: https://huggingface.co/datasets/cyberorigin/cyber_twist_the_tube_v0.1.twitter_suicidal_risk
Twitter Suicide Risk Level Dataset
Short English tweets paired with a 0–4 suicide risk label, used for fine-tuning and
evaluating risk-level classification. This directory holds the final splits:
train.jsonl / val.jsonl / test.jsonl.
Files and size
File
Rows
Share
train.jsonl
7006
80%
val.jsonl
875
10%
test.jsonl
875
10%
Total
8756
100%
Fields
JSONL, one sample per line, three fields only:
Field
Type
Description
id… See the full description on the dataset page: https://huggingface.co/datasets/AdamLeung/twitter_suicidal_risk.twitch-top-live-streams-metadata
Top Twitch Channels Live Viewership & Broadcast Metadata
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/twitch-live-streams-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/twitch-live-fortnite-streams… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/twitch-top-live-streams-metadata.tw-instruct-500k
Dataset Card for tw-instruct-500k
[👋歡迎加入 Discord 討論,我們正在找人一塊擴充這個對話集🎉]
台灣常見任務對話集(Common Task-Oriented Dialogues in Taiwan) 為臺灣社會裡常見的任務對話,從 lianghsun/tw-instruct 截取出 50 萬筆的子集合版本。最新格式請改用 lianghsun/tw-instruct-500k-2511。
Dataset Details
Dataset Description
本資料集為合成資料集(synthetic dataset),由 a. reference-based 與 b. reference-free 兩種子流程組成:
reference-based:以收集自臺灣的繁中文本(用於訓練 lianghsun/Llama-3.2-Taiwan-3B 之語料)為參考,請 LLM 根據文本特性產生對應領域的指令對話。… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-instruct-500k.nick-routledge-x-twitter-archiveNick Routledge X/Twitter Archive
A growing machine-readable collection of posts and replies authored by Nick Routledge (@nick_routledge), drawn from his public X/Twitter archive.
The complete searchable archive is available at https://archive.nickroutledge.org.
Files
posts.jsonl: one structured post per line, suitable for data pipelines and language-model research.
posts.csv: the same authored posts in spreadsheet-compatible form.
Licence
Nick Routledge’s authored post text is made available… See the full description on the dataset page: https://huggingface.co/datasets/fellowservant/nick-routledge-x-twitter-archive.fool-me-twicehttps://github.com/google-research/fool-me-twice
@inproceedings{eisenschlos-etal-2021-fool,
title = "Fool Me Twice: Entailment from {W}ikipedia Gamification",
author = {Eisenschlos, Julian Martin and
Dhingra, Bhuwan and
Bulian, Jannis and
B{\"o}rschinger, Benjamin and
Boyd-Graber, Jordan},
booktitle = "Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies",
month… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/fool-me-twice.
