CoolFace
Datasetpublic

permitt/superglue-sr

BalkanBench SuperGLUE - Serbian Part of BalkanBench - the open, reproducible benchmark for language models across Serbian, Croatian, Montenegrin, and Bosnian (BCMS). Live leaderboard at https://balkanbench.com/leaderboard. Background and motivation: Release of BalkanBench - the vision behind it (Medium, 2026-04-27). This is the Serbian SuperGLUE track of BalkanBench v0.1. Serbian is the official frozen track: the leaderboard's ranked average is computed over 6 ranked tasks… See the full description on the dataset page: https://huggingface.co/datasets/permitt/superglue-sr.

sourceHugging Facecc-by-4.0updated 2mo agoView on Hugging Face
1likes48downloads
Dataset Card

BalkanBench SuperGLUE - Serbian

Part of [BalkanBench](https://balkanbench.com) - the open, reproducible benchmark for language models across Serbian, Croatian, Montenegrin, and Bosnian (BCMS). Live leaderboard at <https://balkanbench.com/leaderboard>. Background and motivation: Release of BalkanBench - the vision behind it (Medium, 2026-04-27).

This is the Serbian SuperGLUE track of BalkanBench v0.1. Serbian is the official frozen track: the leaderboard's ranked average is computed over 6 ranked tasks, with 2 diagnostic tasks reported separately. Croatian and Montenegrin previews live in sibling repos (superglue-hr / superglue-mne).

What's inside

7 task configs are published here in their full train + validation + public test form (train + validation labelled here; the test split lives only in the gated sibling repo, see below):

ConfigWhat it testsTrainValidationTest
boolqYes/no question answering over passages9,4273,2703,245
cb3-way textual entailment (entail / contradict / neutral)25056250
copaCausal reasoning between two alternatives400100500
multircMulti-sentence reading comprehension (multiple correct)27,2434,8489,693
recordCommonsense cloze over news articles5,6071,8691,869
rteBinary textual entailment2,4902773,000
wscCoreference resolution requiring world knowledge554104146
TOTAL45,97110,52418,703

For the 6 ranked v0.1 tasks (BoolQ, CB, COPA, RTE, MultiRC, WSC) the total is 65,853 items across all splits; ReCoRD ships in this repo for community use but is not part of the v0.1 ranked average. The diagnostic tasks AX-b (1,104) and AX-g (356) are Serbian-only and ship in the gated superglue-sr-private sibling repo with their test labels.

COPA source attribution: the copa config is adapted from classla/COPA-SR_lat (the Latin-script Serbian COPA, CLARIN.SI, CC-BY-SA-4.0), created by the CLASSLA / ReLDI Centre Belgrade team as part of the CLARIN.SI South Slavic COPA family alongside COPA-HR and COPA-MK. See the Citation section at the end of this card for the full reference.

Test split is gated

The test split for every config lives only in the gated sibling repo (both inputs and labels): `permitt/superglue-sr-private`, and are accessed only by the official scoring pipeline. This preserves leaderboard integrity: nobody (including model authors) can train against test labels, and train + validation remain a fully open and labeled playground.

To get a number on the leaderboard you generate predictions on the public test inputs with balkanbench predict, then submit the predictions.jsonl through the submission flow.

Loading

python
from datasets import load_dataset

# A specific task
boolq = load_dataset("permitt/superglue-sr", "boolq")
print(boolq["train"][0])

# All ranked tasks at once
for cfg in ["boolq", "cb", "copa", "rte", "multirc", "wsc"]:
    ds = load_dataset("permitt/superglue-sr", cfg)
    print(cfg, {sp: len(ds[sp]) for sp in ds})

Pin a revision for reproducible runs:

python
load_dataset("permitt/superglue-sr", "boolq", revision="v0.1.0-data")

Schema

Each example carries an integer row id idx and the task-specific input fields plus a label column. For example, cb:

python
{
    "idx": 0,
    "premise": "Bio je to složen jezik. ...",
    "hypothesis": "Jezik je ogoljen.",
    "label": 0,  # 0=entailment, 1=contradiction, 2=neutral
}

Per-task field lists are documented in the BalkanBench source repo.

Methodology

  • Translation: original SuperGLUE English -> Serbian via translation + human verification by native speakers.
  • Frozen splits: train/validation/test row counts and IDs are pinned at tag v0.1.0-data; reruns of an evaluation against this tag will see the exact same data.
  • Test labels stay private: see the gated sibling repo above.

License

CC-BY-4.0 (matches the original SuperGLUE source license). Translations and human verification are released under the same terms. Note: the copa config is adapted from a CC-BY-SA-4.0 source (classla/COPA-SR_lat, see Citation below); redistribution of that config should retain the ShareAlike attribution.

Sponsor

Compute for the official v0.1 evaluation is sponsored by [Recrewty](https://recrewty.com).

Links

Citation

If you use this dataset, please cite the Serbian SuperGLUE paper (Perović & Mihajlov, LoResLM 2026), the original SuperGLUE paper, and the BalkanBench release. If you use the copa config specifically, please also cite the classla COPA-SR source dataset it is adapted from:

bibtex
@inproceedings{perovic-mihajlov-2026-serbian,
    title = "{S}erbian {S}uper{GLUE}: Towards an Evaluation Benchmark for {S}outh {S}lavic Language Models",
    author = "Perovic, Mitar and Mihajlov, Teodora",
    editor = "Hettiarachchi, Hansi and Ranasinghe, Tharindu and Plum, Alistair and
              Rayson, Paul and Mitkov, Ruslan and Gaber, Mohamed and Premasiri, Damith and
              Tan, Fiona Anting and Uyangodage, Lasitha",
    booktitle = "Proceedings of the Second Workshop on Language Models for Low-Resource Languages (LoResLM 2026)",
    month = mar,
    year = "2026",
    address = "Rabat, Morocco",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.loreslm-1.30/",
    doi = "10.18653/v1/2026.loreslm-1.30",
    pages = "347--361",
    ISBN = "979-8-89176-377-7"
}

@inproceedings{wang2019superglue,
  title={SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems},
  author={Wang, Alex and Pruksachatkun, Yada and Nangia, Nikita and Singh, Amanpreet and
          Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel R.},
  booktitle={NeurIPS},
  year={2019}
}

@misc{balkanbench2026superglue_sr,
  title={BalkanBench SuperGLUE-SR: Serbian SuperGLUE for the BCMS evaluation suite},
  author={Perović, Mitar and contributors},
  year={2026},
  howpublished={\url{https://huggingface.co/datasets/permitt/superglue-sr}},
  note={Part of BalkanBench v0.1, \url{https://balkanbench.com}}
}

@misc{11356/1708,
 title = {Choice of plausible alternatives dataset in Serbian {COPA}-{SR}},
 author = {Ljube{\v s}i{\'c}, Nikola and Starovi{\'c}, Mirjana and Kuzman, Taja and Samard{\v z}i{\'c}, Tanja},
 url = {http://hdl.handle.net/11356/1708},
 note = {Slovenian language resource repository {CLARIN}.{SI}. Distributed on
   Hugging Face as classla/COPA-SR_lat (Latin-script variant used by
   BalkanBench).},
 copyright = {Creative Commons - Attribution-{ShareAlike} 4.0 International ({CC} {BY}-{SA} 4.0)},
 issn = {2820-4042},
 year = {2022}
}