CoolFace
Datasetpublic

xzx34/cross-lingual-pitfalls

Cross-Lingual Pitfalls Cross-Lingual Pitfalls is a fixed, failure-focused dataset from the ACL 2025 paper "Cross-Lingual Pitfalls: Automatic Probing Cross-Lingual Weakness of Multilingual Large Language Models." It contains 6,713 bilingual English-to-target-language question pairs across 16 target languages. The paper's search-based multilingual LLM evaluation method uses beam search and LLM-based simulation to discover cases where a model answers correctly in English but fails… See the full description on the dataset page: https://huggingface.co/datasets/xzx34/cross-lingual-pitfalls.

sourceHugging Facemitupdated 5d agoView on Hugging Face
0likes142downloads

xzx34/cross-lingual-pitfalls · main · files are served by the source, never re-hosted here