CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mkd-chanwoo /normalized-datasets-for-koreanLLM Normalized Datasets for Korean LLM (Stage 0.5) This is a comprehensive pretraining dataset for Korean LLMs containing ~944M normalized documents from 42 source datasets across English, Korean, Code, and Science domains. Quick Stats Metric Value Total Documents ~944,545,628 Source Datasets 42 Format JSONL (one JSON object per line) Languages English, Korean Domains English, Korean, Code, Science Pipeline Stage Stage 0.5 (after download, before… See the full description on the dataset page: https://huggingface.co/datasets/mkd-chanwoo/normalized-datasets-for-koreanLLM.text1M<n<10M1 likes2.7k downloads4mo agoHugging Face02normster /SystemCheck Dataset Card for SystemCheck Dataset Summary [Project Repo] [🏁 Checkpoints] This repository contains data for our paper, SystemCheck: A Closer Look at System Prompt Reliability, which studies the reliability of system prompts in large language models. SystemCheck is a collection of LLM training and evaluation datasets designed to study the robustness of LLM guardrails. It contains a set of 3000+ system prompts scraped from the ChatGPT store and HuggingChat, SFT/DPO… See the full description on the dataset page: https://huggingface.co/datasets/normster/SystemCheck.texttext-generation100K<n<1M6 likes1.4k downloads1y agoHugging Face03Scicom-intl /Multilingual-Normalizer Multilingual TTS text normalizer (written → spoken) Training pairs for fine-tuning a small LLM as a text-to-speech normalizer: text is a sentence the way people type it (digits, currency symbols, dates, phone numbers, …) and normalized is the exact spoken form, in the same language, with nothing left that a TTS model cannot say. 52,698 rows — 16 monolingual locales and 6 Malaysian code-switched pairs. Every row is digit-free on the spoken side. from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Multilingual-Normalizer.texttext-generation100K<n<1M0 likes207 downloads19d agoHugging Face04ufca-llms /normas-tcuNormasTCU is a Brazilian Portuguese legal IR test collection composed of normative documents from the Brazilian Federal Court of Accounts (TCU). These normative acts may have internal effects (e.g., rules governing internal procedures) or external effects (e.g., rules regulating how the court interacts with other public institutions) and differ from jurisprudential documents in both purpose and structure. Jurisprudential documents typically describe specific cases and present the legal… See the full description on the dataset page: https://huggingface.co/datasets/ufca-llms/normas-tcu.texttext-retrieval10K<n<100K0 likes199 downloads6mo agoHugging Face05ltg /normistral-11b-thinking-evaluationtext10K<n<100K1 likes139 downloads10mo agoHugging Face06novastar112 /pusht_96_norm2_remap pusht_96_norm2 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Splits split records format train 500,000 gzip-compressed JSONL test 1,000 gzip-compressed JSONL Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/. Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2_remap.tabular100K<n<1M0 likes131 downloads5mo agoHugging Face07novastar112 /pusht_96_norm4 pusht_96_norm4 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Splits split records format train 500,000 gzip-compressed JSONL test 200 gzip-compressed JSONL Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/. Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4.tabular100K<n<1M0 likes125 downloads5mo agoHugging Face08novastar112 /pusht_96_norm4_10hz pusht_96_norm4_10hz 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Each episode is success-only and capped at 150 environment steps. Splits split records format train 500,000 gzip-compressed JSONL test 1,000 gzip-compressed JSONL Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz.tabular100K<n<1M0 likes113 downloads5mo agoHugging Face09adedejimakinde /yoruba-normalization-pairs Normalization pairs dataset What this is 24,475 pairs of Yorùbá text, each a corrupted form next to its canonical form, labelled by corruption type. I built it for testing orthographic normalization code. The library This dataset was built alongside yotext, a Python library for Yorùbá orthographic normalization and diacritic restoration. The library is on PyPI at https://pypi.org/project/yotext/ and the source is at… See the full description on the dataset page: https://huggingface.co/datasets/adedejimakinde/yoruba-normalization-pairs.texttext-generation10K<n<100K1 likes102 downloads14d agoHugging Face10novastar112 /pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot PushT norm4 Visual Nomarker All-Step Thinking Trickiness COT This dataset is derived from successful PushT visual-nomarker trajectories in novastar112/pusht_96_norm4_visual_nomarker. Each row contains one full successful trajectory from the first move through the final stop action. Main files: training/pusht_allstep_thinking_cot.jsonl.gz: 500,000 train rows. testing/pusht_allstep_thinking_cot.jsonl.gz: 200 test rows. metadata/final_scan_validation.json: full local scan after repair… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker_allstep_thinking_trickiness_cot.imageimage-to-text100K<n<1M0 likes80 downloads4mo agoHugging Face11zmsali /bangla-dialect-normalization Bangla Dialect Normalization Dataset A parallel corpus mapping standard Bangla to five regional Bangla dialects, built from the Vashantor dataset. Each row contains the same sentence in standard Bangla and Banglish (romanized), alongside its dialect Bangla and dialect Banglish equivalent, plus an English gloss. Regions covered Barishal, Chittagong, Mymensingh, Noakhali, Sylhet Schema Field Description standard_bangla Sentence in standard… See the full description on the dataset page: https://huggingface.co/datasets/zmsali/bangla-dialect-normalization.texttranslation10K<n<100K0 likes74 downloads25d agoHugging Face12ltg /normistral-fluency-annotationManual fluency annotations for Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages Citation @misc{samuel2025fluentalignmentdisfluentjudges, title={Fluent Alignment with Disfluent Judges: Post-training for Lower-resource Languages}, author={David Samuel and Lilja Øvrelid and Erik Velldal and Andrey Kutuzov}, year={2025}, eprint={2512.08777}, archivePrefix={arXiv}, primaryClass={cs.CL}… See the full description on the dataset page: https://huggingface.co/datasets/ltg/normistral-fluency-annotation.textn<1K0 likes71 downloads10mo agoHugging Face13novastar112 /pusht_96_norm4_visual_nomarker pusht_96_norm4_visual_nomarker 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 90. Each episode is success-only and capped at 30 environment steps. Visual marker mode: none. The pusher is rendered with radius 11.0 in 512-space; physics still uses the environment's collision radius. Prompt mode:… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_visual_nomarker.tabular100K<n<1M0 likes69 downloads5mo agoHugging Face14Urdatorn /norma Norma Syllabarum Graecarum - A Benchmark for grc Syllabification and Vowel Length Annotation We introduce Norma as a common benchmark for the evaluation and comparison of NLP tools concerning markup of two tasks for Ancient Greek (grc): (1) vowel length of dichronic vowels (alpha, iota, ypsilon) in open syllables (where they impact syllable weight) and (2) syllabification, both boundaries and weight. This means that the benchmark also indirectly tests handling of sandhi… See the full description on the dataset page: https://huggingface.co/datasets/Urdatorn/norma.text1K<n<10K0 likes60 downloads2mo agoHugging Face15skypro1111 /uk-text-normalization Український TTS-нормалізатор — датасет Пари «письмовий текст → як його вимовляють» для української. Числа, дати, час, гроші, одиниці, скорочення, коди, телефони, IBAN, домени, пошта, римські цифри, латинські вкраплення — те, що треба розгорнути словами перед синтезом мовлення. {"task_id": 0, "combo_names": ["Кількісні числівники (написані цифрами)", "Порядкові числівники (написані цифрами з закінченням)"], "original": "На 1-й полиці стоять 4 книги."… See the full description on the dataset page: https://huggingface.co/datasets/skypro1111/uk-text-normalization.texttext-generation1K<n<10K2 likes54 downloads1mo agoHugging Face16novastar112 /pusht_96_norm4_remap pusht_96_norm4 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 1Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Splits split records format train 500,000 gzip-compressed JSONL test 200 gzip-compressed JSONL Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/. Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_remap.tabular100K<n<1M0 likes53 downloads5mo agoHugging Face17NbAiLab /nynorsk_norm_200eval Nynorsk Norm 200eval nynorsk_norm_200eval is a high-quality, small-scale parallel corpus comprising 200 Norwegian Bokmål–Nynorsk sentence pairs collected from official sources and public institutions. Each example includes: nb: Original sentence in Bokmål nn_original: Original Nynorsk sentence (typically an official translation) nn_alt_original: Original Nynorsk sentence (typically an official translation) - alt version nn_husnorm: Sentence rewritten in Nynorsk following an… See the full description on the dataset page: https://huggingface.co/datasets/NbAiLab/nynorsk_norm_200eval.texttranslationn<1K1 likes50 downloads1y agoHugging Face18francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes50 downloads1mo agoHugging Face19normalcomputing /wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA Github: https://github.com/normal-computing/extended-mind-transformers/ ArXiv: https://arxiv.org/abs/2406.02332 Original dataset by Abacus AI. Developed by: Normal Computing, Adapted from Abacus AI License: Apache 2.0 Long-range Counterfactual Retrieval Benchmark This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.textn<1K1 likes47 downloads2y agoHugging Face20lemon-mint /OpenThoughts-114k-Normalizedprefixes = [ "Generate an executable Python function generated from the given prompt. The function should take stdin as input and print the output. Simply call the function after the definition.", "Return your final response within \\boxed{}. ", "Generate an executable Python function generated from the given prompt. Return the function body without invoking it at the final solution.", ] // -1 if None text100K<n<1M1 likes47 downloads2y agoHugging Face21novastar112 /pusht_96_norm4_10hz_remap pusht_96_norm4_10hz 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 10Hz PPO PushT solver with action codec norm4, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Each episode is success-only and capped at 150 environment steps. Splits split records format train 500,000 gzip-compressed JSONL test 1,000 gzip-compressed JSONL Train/test initial states are filtered to be… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm4_10hz_remap.tabular100K<n<1M0 likes44 downloads5mo agoHugging Face22novastar112 /pusht_96_norm2 pusht_96_norm2 96px PushT PPO successful trajectory dataset. The trajectories are generated by a 1Hz PPO PushT solver with action codec norm2, rendered images in history[*].image, image_prev, and image_next at 96x96 and JPEG quality 85. Splits split records format train 500,000 gzip-compressed JSONL test 1,000 gzip-compressed JSONL Train/test initial states are filtered to be disjoint by init_state_hash; see metadata/. Coordinates in move actions use… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/pusht_96_norm2.tabular100K<n<1M0 likes41 downloads5mo agoHugging Face23riteshhf /repro-siamesenorm-breaking-the-barrier-to-reconciling-pre-post-norm-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes37 downloads2mo agoHugging Face24Zarinaaa /kyrgyz-text-normalization Kyrgyz Text Normalization Dataset A dataset for training and evaluating Kyrgyz text normalization systems. Released subset accompanying "Kyrgyz Text Normalization: A Comparative Study of Neural and Rule-Based Approaches" (MeLLM Workshop @ ACL 2026). What is in this release This is a representative 20,000-pair subset of a larger 1.67M-pair training corpus, plus the full 1,000-example human-verified test set used in the paper. Split Examples Source Verification… See the full description on the dataset page: https://huggingface.co/datasets/Zarinaaa/kyrgyz-text-normalization.texttext-generation10K<n<100K0 likes32 downloads4mo agoHugging Face25llm-semantic-router /halueval-spans-normalized HaluEval Span-Level Dataset (RAGTruth-Normalized Prompts) 🔍 Span-level hallucination detection dataset with prompts normalized to match RAGTruth format for improved cross-dataset compatibility. Quick Start from datasets import load_dataset dataset = load_dataset("llm-semantic-router/halueval-spans-normalized") Why Normalized Prompts? Training on mixed datasets with different prompt formats causes distribution shift: Original Format Normalized Format… See the full description on the dataset page: https://huggingface.co/datasets/llm-semantic-router/halueval-spans-normalized.texttoken-classification10K<n<100K0 likes31 downloads9mo agoHugging Face26ultrastar111 /pusht_96_norm4_cot_chunk_k3_20260622_perseg pusht_96_norm4_cot_chunk_k3_20260622_perseg PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k3_20260622_perseg.tabularreinforcement-learning10K<n<100K0 likes31 downloads3mo agoHugging Face27Ericu950 /norma norma Norma Syllabarum Graecarum: hand-annotated macronization and syllabification benchmark. Mirrored under its original GPL-3.0 licence; the annotation is the creators' work, not ours. Part of the Stoicheia release. text1K<n<10K0 likes31 downloads2mo agoHugging Face28llmsql-bench /prompts_for_tables_normalization_and_new_sqlstext100K<n<1M0 likes28 downloads4mo agoHugging Face29ultrastar111 /pusht_96_norm4_cot_chunk_k10_20260622_perseg pusht_96_norm4_cot_chunk_k10_20260622_perseg PushT (96px, norm4, JPEG q90; coverage task, no hard split — in-dist claims only) — action-conditioned visual world-model SFT data (CoT self-rollout) for the BAGEL-7B-MoT feedback-interval study. Format: gzipped JSONL shards under training/, 1 row = 1 packed episode. CoT rows: per-segment layout — <think> per-step imagined frame (MSE target) </think> + committed action chunk, with a loss-0 "Action executed." + real frame re-grounding… See the full description on the dataset page: https://huggingface.co/datasets/ultrastar111/pusht_96_norm4_cot_chunk_k10_20260622_perseg.tabularreinforcement-learning10K<n<100K0 likes23 downloads3mo agoHugging Face30PhdDz /PubmedQA_5_WITH_RELATION_vsimilarity_primekg_normalizedtext1K<n<10K0 likes22 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.