CoolFace
Datasetpublic

Wulinjuan/CULTURE-MT

CULTURE-MT: Beyond Literal Translation β€” Evaluating Cultural Effectiveness in Social Media UGC CULTURE-MT is a benchmark for evaluating CULtural Transmission and UGC-specific emotion REsonance in Chinese-to-English social media translation. It consists of 1,002 user-generated notes (UGC) spanning 14 content domains, presented at ICML 2026. πŸ“„ Paper: Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC 🌟 Why CULTURE-MT? Standard… See the full description on the dataset page: https://huggingface.co/datasets/Wulinjuan/CULTURE-MT.

sourceHugging Faceupdated 4mo agoView on Hugging Face
6likes36downloads
Dataset Card

CULTURE-MT: Beyond Literal Translation β€” Evaluating Cultural Effectiveness in Social Media UGC

![Leaderboard](https://huggingface.co/spaces/Wulinjuan/CULTURE-MT) ![License](https://creativecommons.org/licenses/by/4.0/)

CULTURE-MT is a benchmark for evaluating CULtural Transmission and UGC-specific emotion REsonance in Chinese-to-English social media translation. It consists of 1,002 user-generated notes (UGC) spanning 14 content domains, presented at ICML 2026.

πŸ“„ Paper: Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC


🌟 Why CULTURE-MT?

Standard machine translation metrics (BLEU, ChrF, COMET) fail to capture whether a translation truly resonates with target-language users. CULTURE-MT introduces cultural effectiveness as a new evaluation criterion, covering:

  • β€”Expressive accuracy β€” semantic fidelity, emotional tone, proper noun handling, and unit/measurement accuracy
  • β€”Cultural adaptability β€” culture-loaded term handling, overall cultural fluency, and addressing/politeness adaptation

πŸ“Š Dataset

SplitNotesDomainsNote Types
Benchmark1,002144 (General, Express, Symbol, Hybrid)

Content Domains

Pets Β· Travel Β· Food Β· Crafts Β· Painting Β· Home Decoration Β· Outdoor Β· Sports Β· Fitness & Weight Loss Β· Technology & Gadgets Β· Cars Β· Games Β· Movies & TV Β· Celebrity News

Note Types

TypeDescription
GeneralInformal UGC with few culture-loaded symbols or distinctive styles
ExpressStrong rhetorical/expressive style (e.g., "planting grass", hyperbole, rhetorical questions)
SymbolHigh density of internet cultural symbols (slang, memes, platform-specific jargon)
HybridBoth rich cultural symbols and distinctive expressive style

πŸ“ Evaluation

Translations are evaluated by JUDGER, a fine-tuned Qwen3-32B model trained on 30K expert- and LLM-annotated samples. It achieves 86.03% accuracy and Cohen's ΞΊ = 0.72 against human expert judgments.

Scoring Rubric (0–3 scale)

ScoreMeaning
0Severe meaning loss or distortion; target readers cannot grasp the original intent or emotion
1Main idea barely understandable; critical cultural errors, poor adaptation
2Main information conveyed accurately; reasonable emotional/contextual expression
3Precise, natural, culturally fluent; fully conveys all information and emotion for English social media readers

Scores 0–1 are treated as culturally ineffective; scores 2–3 as culturally effective.


πŸ† Leaderboard

Submit your translations and get evaluated automatically by our trained JUDGER model at: πŸ‘‰ https://huggingface.co/spaces/Wulinjuan/CULTURE-MT

πŸ“Œ Note on Leaderboard vs. Paper Results Scores reported in the ICML 2026 paper were produced with a single JUDGER inference pass (temperature = 0.6, top-p = 0.95). Since stochastic decoding introduces minor variation across runs, leaderboard scores are computed as the average of four independent inference passes under identical settings to ensure fairness and reproducibility. As a result, leaderboard scores may differ slightly from those reported in the paper.

Top Results (as of 2026-05-27)

RankModelIneff. ↓Eff. ↑0 ↓1 ↓23Avg Score
1GPT-58.08%91.92%0.90%7.19%51.56%40.35%2.31
2Gemini 3 Pro8.95%91.05%0.20%8.75%51.97%39.09%2.30
3CULTURE-MT-baseline-32B12.08%87.92%0.73%11.35%56.56%31.37%2.19
4CULTURE-MT-baseline-8B14.34%85.66%1.27%13.07%57.42%28.24%2.13
5GLM4.615.94%84.06%3.12%12.81%57.39%26.68%2.08
6DeepSeek V3.217.73%82.27%2.59%15.14%56.82%25.45%2.05
7Qwen3-235B-A22B29.27%70.73%12.31%16.97%52.86%17.87%1.76
8Seed-X-PPO30.58%69.42%3.43%27.25%62.23%7.19%1.73

πŸ“€ How to Submit

Submissions are made directly through the leaderboard interface at: πŸ‘‰ https://huggingface.co/spaces/Wulinjuan/CULTURE-MT

Step 1 β€” Upload your translations

Upload a submission.jsonl file. Each line should be a JSON object with the id of the source note and your translation:

json
{"id": "0001", "translation": "This place is amazing. I'll definitely come back!"}
{"id": "0002", "translation": "This is way too ridiculous."}

Step 2 β€” Fill in model information

Complete the submission form with your model details (Model Name, Organization / Team, Base Model, Method, and a brief description). No additional files are required.

Step 3 β€” Get your results

Aggregated scores will appear on the leaderboard automatically after evaluation. If you need detailed per-sample evaluation results, please contact us by email at wulinjuan525@zju.edu.cn.


πŸ“– Citation

If you use CULTURE-MT in your research, please cite:

bibtex
@misc{wu2026literaltranslationevaluatingcultural,
      title={Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC}, 
      author={Linjuan Wu and Ruiqi Zhang and Xinze Lyu and Ye Guo and Daoxin Zhang and Zhe Xu and Yao Hu and Yixin Cao and Yongliang Shen and Weiming Lu},
      year={2026},
      eprint={2605.25626},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2605.25626}, 
}

πŸ“¬ Contact

  • β€”Linjuan Wu: wulinjuan525@zju.edu.cn

This benchmark was developed with support from Zhejiang University, Fudan University, and Xiaohongshu Inc.