datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
bybit-linear-perps-xlmusdtsib200-xlmr-tokenized
SIB-200 Tokenized by XLM-R Large
This repository provides pre-tokenized versions of SIB-200 used in the paper:Cross-Prompt Encoder for Low-Performing LanguagesFindings of IJCNLP–AACL 2025; preprint at arXiv:2508.10352.
The dataset is released to support zero-shot and fully supervised cross-lingual experiments presented in our paper,
ensuring consistent and reproducible tokenization across all languages and experimental settings.
The dataset is organized as a multi-config Hugging… See the full description on the dataset page: https://huggingface.co/datasets/mikaberidze/sib200-xlmr-tokenized.msmarco_token_score_colbertx_xlmr_large_zs_en_en
Dataset Card for "msmarco_token_score_colbertx_xlmr_large_zs_en_en"
More Information needed
cond_ft_none_on_reddit__prcnt_100__test_run_False__xlm-roberta-basecond_ft_subreddit_on_reddit__prcnt_100__test_run_False__xlm-roberta-basexlm-r-bertic-dataData used to train XLM-Roberta-Bertić.wmt14-de-en-xlm-r-tokenized-128cond_ft_none_on_reddit__prcnt_na__test_run_True__xlm-roberta-basetwitter_ae_xlm_roberta_sentiment_stratifiedcond_ft_subreddit_on_reddit__prcnt_20__test_run_False__xlm-roberta-baserapidapi-example-responses-tokenized-xlm-roberta
Dataset Card for "rapidapi-example-responses-tokenized-xlm-roberta"
More Information needed
cond_ft_none_on_reddit__prcnt_20__test_run_False__xlm-roberta-baseparsed-dataset-xlm-robertawsd_myriade_synth_data_gpt4turbo_xlm
Dataset Card for "wsd_myriade_synth_data_gpt4turbo_xlm"
More Information needed
tool-calls-multiturn-salesforce-xlmwsd_myriade_synth_data_multilabel_xlm
Dataset Card for "wsd_myriade_synth_data_multilabel_xlm"
More Information needed
xnli-encoded-xlmrxnli-cl_trans-xlmrKMC-XLMR-KGDataset Summary
KMC-XLMR-KG is a Kurdish medical knowledge graph dataset automatically generated from the Kurdish Medical Corpus (KMC) using the XLM-RoBERTa Large (xlm-roberta-large) multilingual transformer model. The dataset is designed to support information extraction, knowledge graph construction, scientific knowledge discovery, and multilingual AI research for low-resource languages.
The dataset consists of structured semantic triples represented as:
(head, head_type, relation, tail… See the full description on the dataset page: https://huggingface.co/datasets/shkomq/KMC-XLMR-KG.sindhi-gold-tokenized-512-for-xlm-roberta-bertkgz_dataset_chunked_medium_xlm_robertaxlmr_int_hard_curr_trn
Dataset Card for "xlmr_int_hard_curr_trn"
More Information needed
wsd_myriade_synth_data_gpt4turbo_val_xlm
Dataset Card for "wsd_myriade_synth_data_gpt4turbo_val_xlm"
More Information needed
xlmr_eval2
Dataset Card for "xlmr_eval2"
More Information needed
xlmr_int_pr_sw_trn_ep4
Dataset Card for "xlmr_int_pr_sw_trn_ep4"
More Information needed
xlmr_int_hard_curr_trn_ep2_corr
Dataset Card for "xlmr_int_hard_curr_trn_ep2_corr"
More Information needed
xlmr_hard_curr_uda_ep3_corr
Dataset Card for "xlmr_hard_curr_uda_ep3_corr"
More Information needed
ner_locations_dataset_pretokenized_xlm_robertatest_da_xlmr
Dataset Card for "test_da_xlmr"
More Information needed
xlmr_test_10shot
Dataset Card for "xlmr_test_10shot"
More Information needed
