CoolFace
2 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Variable65536 /immersive_translate_en-zh MiniCPM5-1B Immersive Translation SFT Dataset 英译中微调数据集,专为沉浸式翻译插件场景设计。用于微调 MiniCPM5-1B-Base,使其在插件运行时稳定遵循翻译规则、保留代码与 HTML 格式、正确处理多段 %% 分隔。 数据集描述 本数据集主要训练以下能力: 严格遵循沉浸式翻译 system prompt 中的 5 条翻译规则 多段输入的 %% 段落分隔,输入输出段落数严格一致 代码块、行内代码、HTML 标签、URL、专有名词的原样保留 技术文档(GitHub README、Hugging Face 文档)与学术摘要(arXiv)的英译中 单段输入直接输出译文,无"翻译:"等额外前缀 数据格式为 ShareGPT 对话格式,每条样本包含 system / user / assistant 三角色。 数据来源 来源 说明 原始规模 本数据集采样量 License Mxode/BiST… See the full description on the dataset page: https://huggingface.co/datasets/Variable65536/immersive_translate_en-zh.texttranslation10K<n<100K0 likes91 downloads14d agoHugging Face02Orion-zhen /immersive-translate-en2zh Immersive Translate 本数据集是适用于沉浸式翻译中调用大模型翻译时的prompt模板的数据集. 受本地大模型性能限制, 我没有采用多段文字的prompt模板, 仅使用单段文字的默认prompt模板. system prompt也采用通用翻译专家的默认system prompt 本数据集由Garsa3112/ChineseEnglishTranslationDataset和bfsujason@github/mac生成 texttext-generation100K<n<1M1 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.