CoolFace
Datasetpublic

balmussebastian/cross_lingual_wsd_en_ro

English-Romanian Cross-Lingual Word Sense Disambiguation Dataset This dataset accompanies the paper "Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models", accepted at KES 2026. It contains a sense-aligned English-Romanian benchmark for evaluating monolingual and cross-lingual word sense disambiguation (WSD). Each row pairs an English sentence and a Romanian sentence that instantiate the same WordNet/RoWordNet synset. The dataset is intended as… See the full description on the dataset page: https://huggingface.co/datasets/balmussebastian/cross_lingual_wsd_en_ro.

sourceHugging Facecc-by-sa-4.0updated 4mo agoView on Hugging Face
0likes22downloads
Dataset Card

English-Romanian Cross-Lingual Word Sense Disambiguation Dataset

This dataset accompanies the paper "Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models", accepted at KES 2026.

It contains a sense-aligned English-Romanian benchmark for evaluating monolingual and cross-lingual word sense disambiguation (WSD). Each row pairs an English sentence and a Romanian sentence that instantiate the same WordNet/RoWordNet synset.

The dataset is intended as an evaluation benchmark, not as a training corpus.

Dataset Summary

  • —Examples: 2,771
  • —Languages: English, Romanian
  • —Unique English lemmas: 1,197
  • —Unique Romanian lemmas: 1,203
  • —Distinct senses/synsets: 2,771
  • —Parts of speech: nouns and verbs
  • —License: CC BY-SA 4.0

Files

text data/parallel_sentences.csv

Columns

ColumnDescription
en_lemmaEnglish target lemma
pos_enEnglish part of speech
ensynsetnamePrinceton WordNet synset name
en_offsetWordNet synset offset
en_glossEnglish WordNet gloss
ro_lemmaRomanian target lemma
pos_roRomanian part of speech
rosynsetidRoWordNet synset identifier aligned to WordNet
ro_glossRomanian RoWordNet gloss
en_sentenceEnglish sentence instantiating the target sense
ro_sentenceRomanian sentence instantiating the same target sense

Intended Use

This dataset is designed for evaluating:

  • —monolingual WSD in English and Romanian;
  • —cross-lingual WSD in EN→RO and RO→EN directions;
  • —definition-based sense selection;
  • —multilingual LLM lexical-semantic alignment;
  • —robustness to sense inventory size and candidate definition ordering.

Data Creation

The dataset was constructed from aligned WordNet and RoWordNet synsets. English and Romanian lemma pairs were selected through shared synset identifiers, filtered for polysemy and alignment consistency, and expanded into sense inventories.

For each aligned sense, an English sentence was generated to instantiate the target meaning. Romanian sentences were produced through lexically constrained translation, requiring the Romanian target lemma to appear in an inflected form compatible with the intended sense. UDPipe was used during validation to check Romanian lemma realization.

The final dataset was filtered to remove malformed cases and examples where the target lemma or intended sense could not be reliably verified.

Splits

The dataset is released as a single evaluation split.

No train/dev/test split is provided because the dataset was designed for zero-shot diagnostic evaluation rather than supervised training.

Limitations

The dataset contains automatically generated and translated sentences, so some artifacts may remain. It covers only the English-Romanian language pair and should be interpreted as a controlled diagnostic benchmark rather than a fully naturalistic WSD corpus.

Because the dataset is grounded in WordNet/RoWordNet synsets, it only covers senses represented in those lexical resources.

Licensing

This dataset is released under CC BY-SA 4.0.

It contains material derived from Princeton WordNet and RoWordNet. Users should comply with the attribution and licensing requirements of the original lexical resources.

Citation

If you use this dataset, please cite the forthcoming companion paper:

@inproceedings{balmus2026crosslingualwsd,   
  title = {Cross-Lingual Word Sense Disambiguation Remains Challenging for Large Language Models},   
  author = {Balmus, Sebastian and Dura, Bogdan and Uban, Ana-Sabina and Nisioi, Sergiu},   
  booktitle = {Proceedings of the 30th International Conference on Knowledge-Based and Intelligent Information and Engineering Systems (KES 2026)},   
  year = {2026},
  note = {Forthcoming / Accepted for publication}
}