Wulinjuan/CULTURE-MT
CULTURE-MT: Beyond Literal Translation β Evaluating Cultural Effectiveness in Social Media UGC CULTURE-MT is a benchmark for evaluating CULtural Transmission and UGC-specific emotion REsonance in Chinese-to-English social media translation. It consists of 1,002 user-generated notes (UGC) spanning 14 content domains, presented at ICML 2026. π Paper: Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC π Why CULTURE-MT? Standardβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Wulinjuan/CULTURE-MT.
CULTURE-MT: Beyond Literal Translation β Evaluating Cultural Effectiveness in Social Media UGC
 
CULTURE-MT is a benchmark for evaluating CULtural Transmission and UGC-specific emotion REsonance in Chinese-to-English social media translation. It consists of 1,002 user-generated notes (UGC) spanning 14 content domains, presented at ICML 2026.
π Paper: Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC
π Why CULTURE-MT?
Standard machine translation metrics (BLEU, ChrF, COMET) fail to capture whether a translation truly resonates with target-language users. CULTURE-MT introduces cultural effectiveness as a new evaluation criterion, covering:
- Expressive accuracy β semantic fidelity, emotional tone, proper noun handling, and unit/measurement accuracy
- Cultural adaptability β culture-loaded term handling, overall cultural fluency, and addressing/politeness adaptation
π Dataset
Content Domains
Pets Β· Travel Β· Food Β· Crafts Β· Painting Β· Home Decoration Β· Outdoor Β· Sports Β· Fitness & Weight Loss Β· Technology & Gadgets Β· Cars Β· Games Β· Movies & TV Β· Celebrity News
Note Types
π Evaluation
Translations are evaluated by JUDGER, a fine-tuned Qwen3-32B model trained on 30K expert- and LLM-annotated samples. It achieves 86.03% accuracy and Cohen's ΞΊ = 0.72 against human expert judgments.
Scoring Rubric (0β3 scale)
Scores 0β1 are treated as culturally ineffective; scores 2β3 as culturally effective.
π Leaderboard
Submit your translations and get evaluated automatically by our trained JUDGER model at: π https://huggingface.co/spaces/Wulinjuan/CULTURE-MT
π Note on Leaderboard vs. Paper Results Scores reported in the ICML 2026 paper were produced with a single JUDGER inference pass (temperature = 0.6, top-p = 0.95). Since stochastic decoding introduces minor variation across runs, leaderboard scores are computed as the average of four independent inference passes under identical settings to ensure fairness and reproducibility. As a result, leaderboard scores may differ slightly from those reported in the paper.
Top Results (as of 2026-05-27)
π€ How to Submit
Submissions are made directly through the leaderboard interface at: π https://huggingface.co/spaces/Wulinjuan/CULTURE-MT
Step 1 β Upload your translations
Upload a submission.jsonl file. Each line should be a JSON object with the id of the source note and your translation:
{"id": "0001", "translation": "This place is amazing. I'll definitely come back!"}
{"id": "0002", "translation": "This is way too ridiculous."}Step 2 β Fill in model information
Complete the submission form with your model details (Model Name, Organization / Team, Base Model, Method, and a brief description). No additional files are required.
Step 3 β Get your results
Aggregated scores will appear on the leaderboard automatically after evaluation. If you need detailed per-sample evaluation results, please contact us by email at wulinjuan525@zju.edu.cn.
π Citation
If you use CULTURE-MT in your research, please cite:
@misc{wu2026literaltranslationevaluatingcultural,
title={Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC},
author={Linjuan Wu and Ruiqi Zhang and Xinze Lyu and Ye Guo and Daoxin Zhang and Zhe Xu and Yao Hu and Yixin Cao and Yongliang Shen and Weiming Lu},
year={2026},
eprint={2605.25626},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.25626},
}π¬ Contact
- Linjuan Wu: wulinjuan525@zju.edu.cn
This benchmark was developed with support from Zhejiang University, Fudan University, and Xiaohongshu Inc.
