CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hypaai /Hypa-Keyboard-v1 Hypa Keyboard v1 is a 409,598-example instruction dataset for training on-device smart keyboard models — next-word prediction, word completion, autocorrect, and grammatical error correction — across 27 languages, weighted heavily toward African languages that no mainstream keyboard supports. Every example is a three-turn chat (system → user → assistant), so the dataset can be fed directly to any chat-template SFT pipeline. Errors in the input text are synthetically injected from a… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Keyboard-v1.tabulartext-generation100K<n<1M0 likes55 downloads5d agoHugging Face02hypaai /Hypa-Keyboard-v2 Hypa Keyboard v2 is a 409,598-example instruction dataset for training on-device smart keyboard models — next-word prediction, word completion, autocorrect, and grammatical error correction — across 27 languages, weighted heavily toward African languages that no mainstream keyboard supports. Every example is a three-turn chat (system → user → assistant), so the dataset can be fed directly to any chat-template SFT pipeline. Errors in the input text are synthetically injected from a… See the full description on the dataset page: https://huggingface.co/datasets/hypaai/Hypa-Keyboard-v2.tabulartext-generation100K<n<1M0 likes47 downloads5d agoHugging Face03VerbalJungle /German-KeyboardLM-Corpus German KeyboardLM Training Corpus Dieser Datensatz wurde für das Training eines ultrakompakten 33M-Parameter-Sprachmodells für mobile On-Device-Tastaturen (FUTO Keyboard) zusammengestellt. Datenquellen & Herkunft Der Korpus ist eine kuratierte Zusammenstellung aus folgenden Open-Source-Datensätzen: German Wikipedia Dumps (CC BY-SA 4.0) OpenAssistant Conversations (OASST) (Apache 2.0) Leipzig Corpora Collection / News (CC BY) Durchgeführte… See the full description on the dataset page: https://huggingface.co/datasets/VerbalJungle/German-KeyboardLM-Corpus.texttext-generation1M<n<10M1 likes26 downloads1mo agoHugging Face04lianghsun /tw-ptt-keyboard-warrior-chatgated Dataset Card for tw-ptt-keyboard-warrior-chat 本資料集模擬臺灣網路論壇「鍵盤戰士」風格,產生具備臺灣鄉民語氣(含特定流行語、嘴砲、引戰/護航式回應)之對話。可作為 persona / style 微調資料集,用於賦予模型臺灣鄉民風格的對話能力。 Dataset Details Dataset Description 資料集以 OpenAI messages 格式儲存,每筆樣本包含一段對話與對應的生成模型名稱(model)。對話設計取材自臺灣網路論壇常見的提問與爭論模式,並要求 LLM 以鄉民/戰文風格作答(例如使用「==」、「(?)」、「樓上正解」、「在野黨表示」等典型語氣)。目前包含 gossiping 一個 config,共 3,100 筆樣本。 ⚠️ 本資料集刻意保留嘴砲/戲謔風格,使用前請評估目標應用是否適合此語氣。 Curated by: Huang Liang Hsun Language(s) (NLP): Traditional Chinese… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-ptt-keyboard-warrior-chat.texttext-generation1K<n<10K2 likes8 downloads5mo agoHugging Face05mary742 /Russian-keyboard-online Description The Russian Keyboard Online dataset is a curated collection of text inputs, phonetic mappings, and keyboard interaction data generated through a virtual Russian keyboard interface. It supports research and development in multilingual text input, phonetic transcription, and natural language processing for the Russian language. Intended Uses Russian language text input training Cyrillic character recognition Virtual keyboard interaction modeling NLP… See the full description on the dataset page: https://huggingface.co/datasets/mary742/Russian-keyboard-online.text-generation1K<n<10K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.