datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SystemCheck
Dataset Card for SystemCheck
Dataset Summary
[Project Repo] [🏁 Checkpoints]
This repository contains data for our paper, SystemCheck: A Closer Look at System Prompt Reliability, which studies the reliability of system prompts in large language models.
SystemCheck is a collection of LLM training and evaluation datasets designed to study the robustness of LLM guardrails. It contains a set of 3000+ system prompts scraped from the ChatGPT store and HuggingChat, SFT/DPO… See the full description on the dataset page: https://huggingface.co/datasets/normster/SystemCheck.Multilingual-Normalizer
Multilingual TTS text normalizer (written → spoken)
Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence
the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the
exact spoken form, in the same language, with nothing left that a TTS model cannot say.
52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is
digit-free on the spoken side.
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.yoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.uk-text-normalization
Український TTS-нормалізатор — датасет
Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час,
гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські
цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом
мовлення.
{"task_id": 0,
"combo_names": ["Кількісні числівники (написані цифрами)",
"Порядкові числівники (написані цифрами з закінченням)"],
"original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.kyrgyz-text-normalization
Kyrgyz Text Normalization Dataset
A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026).
What is in this release
This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper.
Split
Examples
Source
Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.nuclei-template-generation-dataset-2.3K
nuclei-template-generation-dataset-2.3K
Description:
A specialized instruction-tuning dataset of 2350 examples for training large language models to generate Nuclei YAML templates. Each example consists of a fixed instruction, a structured JSON input describing a vulnerability (CVE, product, HTTP details, detection logic), and the corresponding valid Nuclei template as output. The dataset was constructed from the official Nuclei Templates repository (HTTP… See the full description on the dataset page: https://huggingface.co/datasets/NormanRey/nuclei-template-generation-dataset-2.3K.pikabu_text_normTexts inverse normalized obtained from pikabu dataset.
Normalized using these notebooks for a personal russian normalization model (avaliable on HF Space as well).
All put into single jsonl file with lines like (beautified):
{
"tn": "\\- Ну как так то? У нас в Норильске при минус сорока градусах в буран люди не замерзают, а у вас при минус десяти без ветра человек насмерть замёрз?",
"itn": "\\- Ну как так то? У нас в Норильске при минус 40 градусах в буран люди не замерзают, а у вас при… See the full description on the dataset page: https://huggingface.co/datasets/saarus72/pikabu_text_norm.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.ficbook_text_normTexts inverse normalized obtained from ficbook dataset.
Normalized using these notebooks for a personal russian normalization model (avaliable on HF Space as well).
All put into single jsonl file with lines like (beautified):
{
"replaces": [
{
"text_from": "Боль во всем теле...Боже...я так и знала.... ",
"text_to": "Боль во всем теле...Боже...я так и знала.... "
},
{
"text_from": "5",
"text_to": "Пятая"
},
{
"text_from": " точка буквально… See the full description on the dataset page: https://huggingface.co/datasets/saarus72/ficbook_text_norm.Financial-Form-Normalization-Instructions
Financial Form Normalization Instructions
3,657 English instruction examples across 38 tasks, associated with sraivante/TinyLlama-1.1B-Financial-Form-Normalizer-LoRA.
The examples teach short user replies to map to predefined application field values: dates, amounts, ZIP codes, yes/no or boolean values, category labels, navigation intents and small stage/state JSON objects. The intended task is supplied by a system prompt.
Copyright (c) 2026 sraivante, for original dataset… See the full description on the dataset page: https://huggingface.co/datasets/sraivante/Financial-Form-Normalization-Instructions.
