datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
script_normalization_ckb
Script normalization CKB — noisy → standard Sorani (script_normalization_ckb)
Nearly 6M sentence pairs: the text column holds Central Kurdish written with
non-standard or distorted characters, summary holds the same sentence in standard
Sorani orthography. Use it for text normalisation, spell correction, or to learn the
character-level mapping rules.
At a glance
Rows
5,997,025 — train 5,330,689 / test 666,336
Columns
text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.date_string_normalizationturkish-text-normalization-1m
Turkish Text Normalization 1M v2
Kontrollü altı gürültü türüyle Türkçe metin normalizasyon çiftleri.
Doğrulanmış boyut
Train: 980,000
Validation: 10,000
Test: 10,000
Toplam: 1,000,000
Ana görev sütunları: id, noisy_text, normalized_text, noise_type
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-text-normalization-1m.normalization_faces
Dataset Card for "normalization_faces"
More Information needed
gemma-russian-normalization-dataseteval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_evalThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 2536,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_eval.eval_groot-vials_rackleft_real_fix_normalization_0202_8This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 1,
"total_frames": 506,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_groot-vials_rackleft_real_fix_normalization_0202_8.Marco-train-K-16-alpha-2-k-8-type-log-clipping-False-normalization-Falsedeita-no-normalization
Dataset Card for deita-no-normalization
This dataset has been created with Distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization.turkish-datetime-normalization-500k
Turkish Datetime Normalization 500K v2
Türkçe tarih-saat ifadelerini ISO-8601 ve Europe/Istanbul saat dilimine eşler.
Doğrulanmış boyut
Train: 490,000
Validation: 5,000
Test: 5,000
Toplam: 500,000
Ana görev sütunları: id, text, normalized_datetime, timezone
Provenance
Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı
depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type,
provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-datetime-normalization-500k.bm-text-normalization
bm-text-normalization
Bambara (Bamanankan) orthographic normalisation: map a non-standard spelling to its
standard form. 4,877 short phrase-level pairs in a single config, bamadaba.
Load
from datasets import load_dataset
train = load_dataset("djelia/bm-text-normalization", "bamadaba", split="train")
dev = load_dataset("djelia/bm-text-normalization", "bamadaba", split="dev")
test = load_dataset("djelia/bm-text-normalization", "bamadaba", split="test")
# rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization.task-normalization-chip2020
Dataset Card for "task-normalization-chip2020"
More Information needed
amharic_summarization_mLongT5_predictions_no_normalization_appliedukrainian-dialect-normalizationamharic_summarization_mLongT5_predictions_normalization_appliedamharic_summarization_gemma_predictions_no_normalization_appliedtext-normalization-benchmark
text-normalization-benchmark
The raw Argilla 2.8.0 export of a Bambara
(Bamanankan) text-normalization project: 160 records from four in-house corpora, each with
the annotator's standard-orthography rewrite. 96 carry a submitted response; 64 were
discarded. For a ready-to-score evaluation set, use
djelia/bm-text-normalization-benchmark, the cleaned export of the 96 finished annotations.
The repo is gated: request access on the Hub and run hf auth login.
Load
from… See the full description on the dataset page: https://huggingface.co/datasets/djelia/text-normalization-benchmark.Tool_Output_Interpretation_Normalization
🇰🇿 Kazakh Tool Output Interpretation and Financial Action Dataset
Dataset Summary
Kazakh Tool Output Interpretation and Financial Action Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in tool-augmented agentic workflows that require interpreting structured tool outputs and generating grounded final responses.
The dataset focuses on scenarios where an assistant must understand a Kazakh user request, call the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Output_Interpretation_Normalization.ml_ta_text_normalizationamharic_summarization_walia_II_predictions_normalization_appliedamharic_summarization_llama_3_1_predictions_no_normalization_post_processedtamil_ml_text_normalizationamharic_summarization_gemma_predictions_normalization_appliedamharic_summarization_llama_3_1_predictions_normalization_post_processedbm-text-normalization-benchmark
bm-text-normalization-benchmark
A small human-annotated evaluation set for Bambara (Bamanankan) orthographic normalisation:
96 real-world Bambara strings, each paired with a hand-written standard-orthography rewrite.
It is the cleaned export of the finished annotations from
djelia/text-normalization-benchmark.
Load
from datasets import load_dataset
# the current, whitespace-clean evaluation set
bench = load_dataset("djelia/bm-text-normalization-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization-benchmark.ta_ml_text_normalizationMarco-train-K-16-alpha-2-k-8-type-log-clipping-True-normalization-Falseamharic_summarization_walia_II_predictions_no_normalization_appliedInverse_Text_Normalization_SinhalaMarco-train-K-16-alpha-2-k-8-type-linear-clipping-False-normalization-False
