mikaberidze/sib200-xlmr-tokenized
SIB-200 Tokenized by XLM-R Large This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352. The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper, ensuring consistent and reproducible tokenization across all languages and experimental settings. The dataset is organized as a multi-config… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.
SIB-200 Tokenized by XLM-R Large
This repository provides pre-tokenized versions of SIB-200 used in the paper: Cross-Prompt Encoder for Low-Performing Languages Findings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352.
The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper, ensuring consistent and reproducible tokenization across all languages and experimental settings.
The dataset is organized as a multi-config Hugging Face dataset, where each config corresponds either to (i) a single-language subset (e.g., eng_Latn, kat_Geor, bel_Cyrl, …), following the original SIB-200 structure, or (ii) a grouped source subset (e.g., source_xlmr_joshi5) used for prompt pretraining.
Experimental scope
This dataset release reflects the experimental setup used in the paper:
- Source grouped subsets (3) Used for prompt pretraining:
source_xlmr_enarzho(3 langs)source_xlmr_joshi5(7 langs)source_xlmr_seen(92 langs)
- Target single-language subsets (198) Used for zero-shot and fully supervised cross-lingual evaluation, excluding the 7 highest-resource languages (level 5) as defined by Joshi et al.
