MTEB-BR/leaderboard
121
MTEB-BR — Brazilian Portuguese Massive Text Embedding Benchmark
A Massive Text Embedding Benchmark for Brazilian Portuguese — native, not translated.
- 93 models (73 open-weight + 20 closed commercial APIs) evaluated on 22 native Brazilian-Portuguese tasks across 7 categories
- Admits only data created or found in Portuguese; machine-translated benchmarks (e.g. mMARCO, mkqa) are excluded by construction
- Headline metric: mean_22 (average across all 22 tasks), reported with per-task bootstrap confidence intervals, paired-bootstrap significance, and IRT task discrimination
- Website: mteb-br.org · Paper: arXiv:2607.04581
- All raw results: MTEB-BR/mteb-pt-results · Source code: github.com/tardellirs/mteb-br
Tasks (22)
How to add a model
Submit via GitHub Issues at tardellirs/mteb-br/issues with the model ID, evaluation JSONs, and a reproduction script.
Citation
@article{stekel2026mtebbr,
title = {MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese},
author = {Stekel, Tardelli Ronan Coelho},
journal = {arXiv preprint arXiv:2607.04581},
year = {2026}
}