CoolFace
20 results

text_correction

Cyberfish /text_error_correction文本纠错的相关数据 1 likes3.7k downloads5y agoHugging Faceshibing624 /chinese_text_correction Dataset Card 中文真实场景文本纠错数据集,包括拼写纠错、语法纠错、校对数据。 Repository: shibing624/pycorrector Dataset Summary 拼写纠错数据 lemon_*.tsv:各领域拼写纠错数据集,包括汽车、医疗、新闻、游戏等领域,来自 https://github.com/gingasan/lemon/tree/main/lemon_v2 ec_*.tsv:法律、医学、政府领域拼写纠错数据集,来自 https://github.com/aopolin-lv/ECSpell/tree/main/Data/domains_data medical_csc.tsv :医学领域拼写纠错数据集,来自 https://github.com/yzhihao/MCSCSet/tree/main/data/mcsc_benchmark_dataset… See the full description on the dataset page: https://huggingface.co/datasets/shibing624/chinese_text_correction.text100K<n<1M15 likes337 downloads2y agoHugging Face8uBob /text_error_correction文本纠错的相关数据 0 likes259 downloads2mo agoHugging FaceGenSEC-LLM /SLT-Task1-Post-ASR-Text-Correction Dataset Name: Pilot dataset for Multi-domain ASR corrections Description This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains. It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0 Structure Data Split The dataset is divided into training and test splits: Training Data: 281,082 entries Approximately 6,255… See the full description on the dataset page: https://huggingface.co/datasets/GenSEC-LLM/SLT-Task1-Post-ASR-Text-Correction.text100K<n<1M2 likes114 downloads2y agoHugging FaceFrancophonIA /ICDAR_2019_Competition_Post-OCR_Text_Correction [!NOTE] Dataset origin: https://zenodo.org/records/3515403 Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019) => Website: http://l3i.univ-larochelle.fr/ICDAR2019PostOCR Description: The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/ICDAR_2019_Competition_Post-OCR_Text_Correction.0 likes77 downloads1y agoHugging FaceLakoreAI /vn-text-correction-0001 Dataset Card for Vietnamese Text Correction Dataset Dataset Description This dataset contains Vietnamese text pairs for training and evaluating text correction models. Each example consists of an erroneous text and its corrected version, making it ideal for: Grammar correction Spelling correction Text normalization Language model fine-tuning Dataset Summary Language: Vietnamese (vi) Format: Text correction pairs Size: ~4.0M examples across… See the full description on the dataset page: https://huggingface.co/datasets/LakoreAI/vn-text-correction-0001.textfill-mask1M<n<10M0 likes61 downloads11mo agoHugging Face