datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Hypa-Keyboard-v1
Hypa Keyboard v1 is a 409,598-example instruction dataset for training on-device smart keyboard models — next-word prediction, word completion, autocorrect, and grammatical error correction — across 27 languages, weighted heavily toward African languages that no mainstream keyboard supports.
Every example is a three-turn chat (system → user → assistant), so the dataset can be fed directly to any chat-template SFT pipeline. Errors in the input text are synthetically injected from a… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Keyboard-v1.Hypa-Keyboard-v2
Hypa Keyboard v2 is a 409,598-example instruction dataset for training on-device smart keyboard models — next-word prediction, word completion, autocorrect, and grammatical error correction — across 27 languages, weighted heavily toward African languages that no mainstream keyboard supports.
Every example is a three-turn chat (system → user → assistant), so the dataset can be fed directly to any chat-template SFT pipeline. Errors in the input text are synthetically injected from a… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Keyboard-v2.German-KeyboardLM-Corpus
German KeyboardLM Training Corpus
Dieser Datensatz wurde für das Training eines ultrakompakten 33M-Parameter-Sprachmodells für mobile On-Device-Tastaturen (FUTO Keyboard) zusammengestellt.
Datenquellen & Herkunft
Der Korpus ist eine kuratierte Zusammenstellung aus folgenden Open-Source-Datensätzen:
German Wikipedia Dumps (CC BY-SA 4.0)
OpenAssistant Conversations (OASST) (Apache 2.0)
Leipzig Corpora Collection / News (CC BY)
Durchgeführte… See the full description on the dataset page: https://huggingface.co/datasets/VerbalJungle/German-KeyboardLM-Corpus.tw-ptt-keyboard-warrior-chat
Dataset Card for tw-ptt-keyboard-warrior-chat
本資料集模擬臺灣網路論壇「鍵盤戰士」風格,產生具備臺灣鄉民語氣(含特定流行語、嘴砲、引戰/護航式回應)之對話。可作為 persona / style 微調資料集,用於賦予模型臺灣鄉民風格的對話能力。
Dataset Details
Dataset Description
資料集以 OpenAI messages 格式儲存,每筆樣本包含一段對話與對應的生成模型名稱(model)。對話設計取材自臺灣網路論壇常見的提問與爭論模式,並要求 LLM 以鄉民/戰文風格作答(例如使用「==」、「(?)」、「樓上正解」、「在野黨表示」等典型語氣)。目前包含 gossiping 一個 config,共 3,100 筆樣本。
⚠️ 本資料集刻意保留嘴砲/戲謔風格,使用前請評估目標應用是否適合此語氣。
Curated by: Huang Liang Hsun
Language(s) (NLP): Traditional Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ptt-keyboard-warrior-chat.
