CoolFace
Datasetpublicgated

lianghsun/zhcn-zhtw-localization-pairs

zh-CN → zh-TW Localization Pairs 簡體中文 → 台灣繁體中文在地化平行語料:49,857 組句對,每組包含真實簡體中文 段落與對應的台灣繁中改寫(字形轉換 + 中國大陸用語 → 台灣慣用詞)。 輸出由內部的大型在地化模型生成(synthetic),非人工翻譯 每列附七道自動品質檢查的判定(c_pass),整體通過率 87.8% 用途:訓練/蒸餾輕量的簡轉繁在地化模型、簡繁轉換評測 欄位 欄位 型別 說明 id str sha1(src) 前 16 碼,全集唯一 src str 簡體中文原文(真實網路文本,非合成簡中) out str 台灣繁中在地化輸出 simp_ratio float 原文的簡體專屬字率(經驗字頻法判定) c_pass bool 七道檢查全數通過 c_reasons str 未通過時的原因(分號分隔) 來源與生成 輸入:取自公開網路語料的真實簡體中文段落,以兩個判別器篩選: 簡體專屬字率… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/zhcn-zhtw-localization-pairs.

sourceHugging Faceodc-byupdated 28d agoView on Hugging Face
0likes23downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
lianghsun/zhcn-zhtw-localization-pairs · CoolFace