datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
normalized-datasets-for-koreanLLM
Normalized Datasets for Korean LLM (Stage 0.5)
This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains.
Quick Stats
Metric
Value
Total Documents
~944,545,628
Source Datasets
42
Format
JSONL (one JSON object per line)
Languages
English, Korean
Domains
English, Korean, Code, Science
Pipeline Stage
Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.SystemCheck
Dataset Card for SystemCheck
Dataset Summary
[Project Repo] [🏁 Checkpoints]
This repository contains data for our paper, SystemCheck: A Closer Look at System Prompt Reliability, which studies the reliability of system prompts in large language models.
SystemCheck is a collection of LLM training and evaluation datasets designed to study the robustness of LLM guardrails. It contains a set of 3000+ system prompts scraped from the ChatGPT store and HuggingChat, SFT/DPO… See the full description on the dataset page: https://huggingface.co/datasets/normster/SystemCheck.Multilingual-Normalizer
Multilingual TTS text normalizer (written → spoken)
Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence
the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the
exact spoken form, in the same language, with nothing left that a TTS model cannot say.
52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is
digit-free on the spoken side.
from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.normas-tcuNormasTCU is a Brazilian Portuguese legal IR test collection composed of normative documents from the Brazilian Federal Court of Accounts (TCU). These normative acts may have internal effects (e.g., rules governing internal procedures) or external effects (e.g., rules regulating how the court interacts with other public institutions) and differ from jurisprudential documents in both purpose and structure. Jurisprudential documents typically describe specific cases and present the legal… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/normas-tcu.normistral-11b-thinking-evaluationpusht_96_norm2_remap
pusht_96_norm2
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2_remap.pusht_96_norm4
pusht_96_norm4
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
200
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4.pusht_96_norm4_10hz
pusht_96_norm4_10hz
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Each episode is success-only and capped at 150 environment steps.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz.yoruba-normalization-pairs
Normalization pairs dataset
What this is
24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code.
The library
This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot
PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT
This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker.
Each row contains one full successful trajectory from the first move through the final stop action.
Main files:
training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows.
testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows.
metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.bangla-dialect-normalization
Bangla Dialect Normalization Dataset
A parallel corpus mapping standard Bangla to five regional Bangla dialects,
built from the Vashantor dataset. Each row contains the same sentence in
standard Bangla and Banglish (romanized), alongside its dialect Bangla and
dialect Banglish equivalent, plus an English gloss.
Regions covered
Barishal, Chittagong, Mymensingh, Noakhali, Sylhet
Schema
Field
Description
standard_bangla
Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages
Citation
@misc{samuel2025fluentalignmentdisfluentjudges,
title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages},
author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov},
year={2025},
eprint={2512.08777},
archivePrefix={arXiv},
primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.pusht_96_norm4_visual_nomarker
pusht_96_norm4_visual_nomarker
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 90.
Each episode is success-only and capped at 30 environment steps.
Visual marker mode: none. The pusher is rendered with radius 11.0 in 512-space; physics still uses the environment's collision radius.
Prompt mode:… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker.norma
Norma Syllabarum Graecarum - A Benchmark for grc Syllabification and Vowel Length Annotation
We introduce Norma as a common benchmark for the evaluation and comparison of NLP tools concerning markup of two tasks for Ancient Greek (grc): (1) vowel length of dichronic vowels (alpha, iota, ypsilon) in open syllables (where they impact syllable weight) and (2) syllabification, both boundaries and weight. This means that the benchmark also indirectly tests handling of sandhi… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/norma.uk-text-normalization
Український TTS-нормалізатор — датасет
Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час,
гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські
цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом
мовлення.
{"task_id": 0,
"combo_names": ["Кількісні числівники (написані цифрами)",
"Порядкові числівники (написані цифрами з закінченням)"],
"original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.pusht_96_norm4_remap
pusht_96_norm4
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
200
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_remap.nynorsk_norm_200eval
Nynorsk Norm 200eval
nynorsk_norm_200eval is a high-quality, small-scale parallel corpus comprising 200 Norwegian Bokmål–Nynorsk sentence pairs collected from official sources and public institutions. Each example includes:
nb: Original sentence in Bokmål
nn_original: Original Nynorsk sentence (typically an official translation)
nn_alt_original: Original Nynorsk sentence (typically an official translation) - alt version
nn_husnorm: Sentence rewritten in Nynorsk following an… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_norm_200eval.DistillDetect-normalized-traces
DistillDetect — format-normalized teacher traces
Teacher responses from Reference-Based Distillation Detection in LLMs
(arXiv:2607.09692), rewritten so that every
teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set
pairs.
Why this exists
In the released data each teacher emits a structurally different response, so a
student trained on it — and any detector trained to attribute it — can key on
surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA
Github: https://github.com/normal-computing/extended-mind-transformers/
ArXiv: https://arxiv.org/abs/2406.02332
Original dataset by Abacus AI.
Developed by: Normal Computing, Adapted from Abacus AI
License: Apache 2.0
Long-range Counterfactual Retrieval Benchmark
This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.OpenThoughts-114k-Normalizedprefixes = [
"Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.",
"Return your final response within \\boxed{}. ",
"Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.",
]
// -1 if None
pusht_96_norm4_10hz_remap
pusht_96_norm4_10hz
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Each episode is success-only and capped at 150 environment steps.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz_remap.pusht_96_norm2
pusht_96_norm2
96px PushT PPO successful trajectory dataset.
The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85.
Splits
split
records
format
train
500,000
gzip-compressed JSONL
test
1,000
gzip-compressed JSONL
Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/.
Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2.repro-siamesenorm-breaking-the-barrier-to-reconciling-pre-post-norm-traces
Agent traces
Agent sessions published from a Trackio Logbook.
kyrgyz-text-normalization
Kyrgyz Text Normalization Dataset
A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026).
What is in this release
This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper.
Split
Examples
Source
Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.halueval-spans-normalized
HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts)
🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility.
Quick Start
from datasets import load_dataset
dataset = load_dataset("llm-semantic-router/halueval-spans-normalized")
Why Normalized Prompts?
Training on mixed datasets with different prompt formats causes distribution shift:
Original Format
Normalized Format… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/halueval-spans-normalized.pusht_96_norm4_cot_chunk_k3_20260622_perseg
pusht_96_norm4_cot_chunk_k3_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k3_20260622_perseg.norma
norma
Norma Syllabarum Graecarum: hand-annotated macronization and syllabification benchmark.
Mirrored under its original GPL-3.0 licence; the annotation is the creators' work, not ours.
Part of the Stoicheia release.
prompts_for_tables_normalization_and_new_sqlspusht_96_norm4_cot_chunk_k10_20260622_perseg
pusht_96_norm4_cot_chunk_k10_20260622_perseg
PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the
BAGEL-7B-MoT feedback-interval study.
Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment
layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a
loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k10_20260622_perseg.PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalized
