balmussebastian/cross_lingual_wsd_en_ro
English-Romanian Cross-Lingual Word Sense Disambiguation Dataset This dataset accompanies the paper "Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models", accepted at KES 2026. It contains a sense-aligned English-Romanian benchmark for evaluating monolingual and cross-lingual word sense disambiguation (WSD). Each row pairs an English sentence and a Romanian sentence that instantiate the same WordNet/RoWordNet synset. The dataset is intended as… See the full description on the dataset page: https://huggingface.co/datasets/balmussebastian/cross_lingual_wsd_en_ro.
English-Romanian Cross-Lingual Word Sense Disambiguation Dataset
This dataset accompanies the paper "Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models", accepted at KES 2026.
It contains a sense-aligned English-Romanian benchmark for evaluating monolingual and cross-lingual word sense disambiguation (WSD). Each row pairs an English sentence and a Romanian sentence that instantiate the same WordNet/RoWordNet synset.
The dataset is intended as an evaluation benchmark, not as a training corpus.
Dataset Summary
- Examples: 2,771
- Languages: English, Romanian
- Unique English lemmas: 1,197
- Unique Romanian lemmas: 1,203
- Distinct senses/synsets: 2,771
- Parts of speech: nouns and verbs
- License: CC BY-SA 4.0
Files
text data/parallel_sentences.csv
Columns
Intended Use
This dataset is designed for evaluating:
- monolingual WSD in English and Romanian;
- cross-lingual WSD in EN→RO and RO→EN directions;
- definition-based sense selection;
- multilingual LLM lexical-semantic alignment;
- robustness to sense inventory size and candidate definition ordering.
Data Creation
The dataset was constructed from aligned WordNet and RoWordNet synsets. English and Romanian lemma pairs were selected through shared synset identifiers, filtered for polysemy and alignment consistency, and expanded into sense inventories.
For each aligned sense, an English sentence was generated to instantiate the target meaning. Romanian sentences were produced through lexically constrained translation, requiring the Romanian target lemma to appear in an inflected form compatible with the intended sense. UDPipe was used during validation to check Romanian lemma realization.
The final dataset was filtered to remove malformed cases and examples where the target lemma or intended sense could not be reliably verified.
Splits
The dataset is released as a single evaluation split.
No train/dev/test split is provided because the dataset was designed for zero-shot diagnostic evaluation rather than supervised training.
Limitations
The dataset contains automatically generated and translated sentences, so some artifacts may remain. It covers only the English-Romanian language pair and should be interpreted as a controlled diagnostic benchmark rather than a fully naturalistic WSD corpus.
Because the dataset is grounded in WordNet/RoWordNet synsets, it only covers senses represented in those lexical resources.
Licensing
This dataset is released under CC BY-SA 4.0.
It contains material derived from Princeton WordNet and RoWordNet. Users should comply with the attribution and licensing requirements of the original lexical resources.
Citation
If you use this dataset, please cite the forthcoming companion paper:
@inproceedings{balmus2026crosslingualwsd,
title = {Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models},
author = {Balmus, Sebastian and Dura, Bogdan and Uban, Ana-Sabina and Nisioi, Sergiu},
booktitle = {Proceedings of the 30th International Conference on Knowledge-Based and Intelligent Information and Engineering Systems (KES 2026)},
year = {2026},
note = {Forthcoming / Accepted for publication}
}