CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01normster /SystemCheck Dataset Card for SystemCheck Dataset Summary [Project Repo] [🏁 Checkpoints] This repository contains data for our paper, SystemCheck: A Closer Look at System Prompt Reliability, which studies the reliability of system prompts in large language models. SystemCheck is a collection of LLM training and evaluation datasets designed to study the robustness of LLM guardrails. It contains a set of 3000+ system prompts scraped from the ChatGPT store and HuggingChat, SFT/DPO… See the full description on the dataset page: https://huggingface.co/datasets/normster/SystemCheck.texttext-generation100K<n<1M6 likes1.4k downloads1y agoHugging Face02Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes207 downloads20d agoHugging Face03adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes102 downloads15d agoHugging Face04skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes54 downloads1mo agoHugging Face05francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes50 downloads1mo agoHugging Face06Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes32 downloads4mo agoHugging Face07atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes21 downloads1mo agoHugging Face08NormanRey /nuclei-template-generation-dataset-2.3K nuclei-template-generation-dataset-2.3K Description: A specialized instruction-tuning dataset of 2350 examples for training large language models to generate Nuclei YAML templates. Each example consists of a fixed instruction, a structured JSON input describing a vulnerability (CVE, product, HTTP details, detection logic), and the corresponding valid Nuclei template as output. The dataset was constructed from the official Nuclei Templates repository (HTTP… See the full description on the dataset page: https://huggingface.co/datasets/NormanRey/nuclei-template-generation-dataset-2.3K.texttext-generation1K<n<10K0 likes18 downloads4mo agoHugging Face09saarus72 /pikabu_text_normTexts inverse normalized obtained from pikabu dataset. Normalized using these notebooks for a personal russian normalization model (avaliable on HF Space as well). All put into single jsonl file with lines like (beautified): { "tn": "\\- Ну как так то? У нас в Норильске при минус сорока градусах в буран люди не замерзают, а у вас при минус десяти без ветра человек насмерть замёрз?", "itn": "\\- Ну как так то? У нас в Норильске при минус 40 градусах в буран люди не замерзают, а у вас при… See the full description on the dataset page: https://huggingface.co/datasets/saarus72/pikabu_text_norm.tabulartext-generation1M<n<10M1 likes17 downloads3y agoHugging Face10deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes12 downloads5mo agoHugging Face11saarus72 /ficbook_text_normTexts inverse normalized obtained from ficbook dataset. Normalized using these notebooks for a personal russian normalization model (avaliable on HF Space as well). All put into single jsonl file with lines like (beautified): { "replaces": [ { "text_from": "Боль во всем теле...Боже...я так и знала.... ", "text_to": "Боль во всем теле...Боже...я так и знала.... " }, { "text_from": "5", "text_to": "Пятая" }, { "text_from": " точка буквально… See the full description on the dataset page: https://huggingface.co/datasets/saarus72/ficbook_text_norm.texttext-generation1M<n<10M0 likes9 downloads3y agoHugging Face12sraivante /Financial-Form-Normalization-Instructions Financial Form Normalization Instructions 3,657 English instruction examples across 38 tasks, associated with sraivante/TinyLlama-1.1B-Financial-Form-Normalizer-LoRA. The examples teach short user replies to map to predefined application field values: dates, amounts, ZIP codes, yes/no or boolean values, category labels, navigation intents and small stage/state JSON objects. The intended task is supplied by a system prompt. Copyright (c) 2026 sraivante, for original dataset… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/Financial-Form-Normalization-Instructions.texttext-generation1K<n<10K0 likes18h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.