CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01razhan /script_normalization_ckb Script normalization CKB — noisy → standard Sorani (script_normalization_ckb) Nearly 6M sentence pairs: the text column holds Central Kurdish written with non-standard or distorted characters, summary holds the same sentence in standard Sorani orthography. Use it for text normalisation, spell correction, or to learn the character-level mapping rules. At a glance Rows 5,997,025 — train 5,330,689 / test 666,336 Columns text (noisy input), summary… See the full description on the dataset page: https://huggingface.co/datasets/razhan/script_normalization_ckb.text1M<n<10M0 likes52 downloads2d agoHugging Face02LegionIntel /date_string_normalizationtext10K<n<100K0 likes31 downloads2y agoHugging Face03GoktugD /turkish-text-normalization-1m Turkish Text Normalization 1M v2 Kontrollü altı gürültü türüyle Türkçe metin normalizasyon çiftleri. Doğrulanmış boyut Train: 980,000 Validation: 10,000 Test: 10,000 Toplam: 1,000,000 Ana görev sütunları: id, noisy_text, normalized_text, noise_type Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-text-normalization-1m.texttext-generation1M<n<10M0 likes29 downloads1mo agoHugging Face04dataautogpt3 /normalization_faces Dataset Card for "normalization_faces" More Information needed imagen<1K4 likes27 downloads3y agoHugging Face05kenenbek /gemma-russian-normalization-datasettext100K<n<1M0 likes26 downloads11mo agoHugging Face06sreetz-nv /eval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_evalThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 2, "total_frames": 2536, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:2" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_gr00t_n1d6-vials_rackleft_real_fix_normalization_0202_eval.tabularrobotics1K<n<10K0 likes26 downloads8mo agoHugging Face07sreetz-nv /eval_groot-vials_rackleft_real_fix_normalization_0202_8This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 1, "total_frames": 506, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 30, "splits": { "train": "0:1" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/sreetz-nv/eval_groot-vials_rackleft_real_fix_normalization_0202_8.tabularroboticsn<1K0 likes23 downloads8mo agoHugging Face08mengdili /Marco-train-K-16-alpha-2-k-8-type-log-clipping-False-normalization-Falsetext10K<n<100K0 likes21 downloads1y agoHugging Face09distilabel-internal-testing /deita-no-normalization Dataset Card for deita-no-normalization This dataset has been created with Distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/distilabel-internal-testing/deita-no-normalization.tabular1K<n<10K0 likes20 downloads2y agoHugging Face10GoktugD /turkish-datetime-normalization-500k Turkish Datetime Normalization 500K v2 Türkçe tarih-saat ifadelerini ISO-8601 ve Europe/Istanbul saat dilimine eşler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, normalized_datetime, timezone Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-datetime-normalization-500k.texttext-generation100K<n<1M0 likes17 downloads1mo agoHugging Face11djelia /bm-text-normalizationgated bm-text-normalization Bambara (Bamanankan) orthographic normalisation: map a non-standard spelling to its standard form. 4,877 short phrase-level pairs in a single config, bamadaba. Load from datasets import load_dataset train = load_dataset("djelia/bm-text-normalization", "bamadaba", split="train") dev = load_dataset("djelia/bm-text-normalization", "bamadaba", split="dev") test = load_dataset("djelia/bm-text-normalization", "bamadaba", split="test") # rows… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization.texttext-generation1K<n<10K0 likes13 downloads2mo agoHugging Face12LRAI /task-normalization-chip2020 Dataset Card for "task-normalization-chip2020" More Information needed text10K<n<100K0 likes12 downloads3y agoHugging Face13userdavek /amharic_summarization_mLongT5_predictions_no_normalization_appliedtext1K<n<10K0 likes12 downloads9mo agoHugging Face14DenysKovalML /ukrainian-dialect-normalizationtext100K<n<1M0 likes12 downloads5mo agoHugging Face15userdavek /amharic_summarization_mLongT5_predictions_normalization_appliedtext1K<n<10K0 likes9 downloads9mo agoHugging Face16userdavek /amharic_summarization_gemma_predictions_no_normalization_appliedtext1K<n<10K0 likes9 downloads9mo agoHugging Face17djelia /text-normalization-benchmarkgated text-normalization-benchmark The raw Argilla 2.8.0 export of a Bambara (Bamanankan) text-normalization project: 160 records from four in-house corpora, each with the annotator's standard-orthography rewrite. 96 carry a submitted response; 64 were discarded. For a ready-to-score evaluation set, use djelia/bm-text-normalization-benchmark, the cleaned export of the 96 finished annotations. The repo is gated: request access on the Hub and run hf auth login. Load from… See the full description on the dataset page: https://huggingface.co/datasets/djelia/text-normalization-benchmark.texttext-generationn<1K0 likes9 downloads2mo agoHugging Face18farabi-lab /Tool_Output_Interpretation_Normalizationgated 🇰🇿 Kazakh Tool Output Interpretation and Financial Action Dataset Dataset Summary Kazakh Tool Output Interpretation and Financial Action Dataset is a Kazakh-language dataset designed for training and evaluating Large Language Models (LLMs) in tool-augmented agentic workflows that require interpreting structured tool outputs and generating grounded final responses. The dataset focuses on scenarios where an assistant must understand a Kazakh user request, call the… See the full description on the dataset page: https://huggingface.co/datasets/farabi-lab/Tool_Output_Interpretation_Normalization.texttext-generation1K<n<10K0 likes9 downloads2mo agoHugging Face19Saikrishna2403 /ml_ta_text_normalizationtext10K<n<100K0 likes8 downloads2y agoHugging Face20userdavek /amharic_summarization_walia_II_predictions_normalization_appliedtext1K<n<10K0 likes8 downloads9mo agoHugging Face21userdavek /amharic_summarization_llama_3_1_predictions_no_normalization_post_processedtext1K<n<10K0 likes8 downloads9mo agoHugging Face22Saikrishna2403 /tamil_ml_text_normalizationtext10K<n<100K0 likes7 downloads2y agoHugging Face23userdavek /amharic_summarization_gemma_predictions_normalization_appliedtext1K<n<10K0 likes7 downloads9mo agoHugging Face24userdavek /amharic_summarization_llama_3_1_predictions_normalization_post_processedtext1K<n<10K0 likes7 downloads9mo agoHugging Face25djelia /bm-text-normalization-benchmarkgated bm-text-normalization-benchmark A small human-annotated evaluation set for Bambara (Bamanankan) orthographic normalisation: 96 real-world Bambara strings, each paired with a hand-written standard-orthography rewrite. It is the cleaned export of the finished annotations from djelia/text-normalization-benchmark. Load from datasets import load_dataset # the current, whitespace-clean evaluation set bench = load_dataset("djelia/bm-text-normalization-benchmark"… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bm-text-normalization-benchmark.texttext-generationn<1K0 likes7 downloads2mo agoHugging Face26Saikrishna2403 /ta_ml_text_normalizationtext10K<n<100K0 likes6 downloads2y agoHugging Face27mengdili /Marco-train-K-16-alpha-2-k-8-type-log-clipping-True-normalization-Falsetext10K<n<100K0 likes6 downloads1y agoHugging Face28userdavek /amharic_summarization_walia_II_predictions_no_normalization_appliedtext1K<n<10K0 likes6 downloads9mo agoHugging Face29SPEAK-PP /Inverse_Text_Normalization_Sinhalatext1K<n<10K0 likes6 downloads6mo agoHugging Face30mengdili /Marco-train-K-16-alpha-2-k-8-type-linear-clipping-False-normalization-Falsetext10K<n<100K0 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.