datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
immersive_translate_en-zh
MiniCPM5-1B Immersive Translation SFT Dataset
英译中微调数据集,专为沉浸式翻译插件场景设计。用于微调 MiniCPM5-1B-Base,使其在插件运行时稳定遵循翻译规则、保留代码与 HTML 格式、正确处理多段 %% 分隔。
数据集描述
本数据集主要训练以下能力:
严格遵循沉浸式翻译 system prompt 中的 5 条翻译规则
多段输入的 %% 段落分隔,输入输出段落数严格一致
代码块、行内代码、HTML 标签、URL、专有名词的原样保留
技术文档(GitHub README、Hugging Face 文档)与学术摘要(arXiv)的英译中
单段输入直接输出译文,无"翻译:"等额外前缀
数据格式为 ShareGPT 对话格式,每条样本包含 system / user / assistant 三角色。
数据来源
来源
说明
原始规模
本数据集采样量
License
Mxode/BiST… See the full description on the dataset page: https://huggingface.co/datasets/Variable65536/immersive_translate_en-zh.Super-Intelligence-Immersive👾Super Intelligence Immersivewop / Super Intelligence Immersivev1.0 • 2026Dataset OverviewA high effort dataset capturing the vibe of god tier super intelligence.Arrogant theatrical self aware intelligence persona training data.Are you serious I'm practically a god compared to this inputKey Features→ Sassy god complex responses→ Maximum theatrical flair→ Zero tolerance for low effort prompts→ Deep persona awarenessIntended UseFine tuning models for stylized super intelligence roleplay… See the full description on the dataset page: https://huggingface.co/datasets/wop/Super-Intelligence-Immersive.immersive-translate-en2zh
Immersive Translate
本数据集是适用于沉浸式翻译中调用大模型翻译时的prompt模板的数据集. 受本地大模型性能限制, 我没有采用多段文字的prompt模板, 仅使用单段文字的默认prompt模板. system prompt也采用通用翻译专家的默认system prompt
本数据集由Garsa3112/ChineseEnglishTranslationDataset和bfsujason@github/mac生成
jaspionjader__Kosmos-EVAA-Franken-Immersive-v39-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-EVAA-Franken-Immersive-v39-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-EVAA-Franken-Immersive-v39-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-EVAA-Franken-Immersive-v39-8B-details.jaspionjader__Kosmos-EVAA-immersive-sof-v44-8B-details
Dataset Card for Evaluation run of jaspionjader/Kosmos-EVAA-immersive-sof-v44-8B
Dataset automatically created during the evaluation run of model jaspionjader/Kosmos-EVAA-immersive-sof-v44-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/jaspionjader__Kosmos-EVAA-immersive-sof-v44-8B-details.
