CoolFace
Datasetpublic

mikaberidze/sib200-xlmr-tokenized

SIB-200 Tokenized by XLM-R Large This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352. The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper, ensuring consistent and reproducible tokenization across all languages and experimental settings. The dataset is organized as a multi-config… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.

sourceHugging Facecc-by-sa-4.0updated 9mo agoView on Hugging Face
0likes495downloads
discussions and pull requests

Conversations for this repository live on Hugging Face.

CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.

Open discussions on Hugging Face
mikaberidze/sib200-xlmr-tokenized · CoolFace