datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
farsick-sts
Dataset Summary
FarSick STS is a Persian (Farsi) dataset designed for the Semantic Textual Similarity (STS) task. It is a part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was developed by translating and adapting the English SICK (Sentences Involving Compositional Knowledge) dataset, and it features Persian sentence pairs annotated for their degree of semantic relatedness.
Language(s): Persian (Farsi)
Task(s): Semantic Textual Similarity (STS)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/farsick-sts.tajik-farsi-transliteration-benchmark
🇹🇯🇮🇷 Tajik-Farsi Transliteration Benchmark
Официальный бенчмарк машинной транслитерации между таджикским (кириллица) и фарси (персо-арабская графика).Результаты получены на 40k параллельных предложениях с оценкой по 3 случайным сидам, bootstrap 95% CI и парными статистическими тестами.
📊 Ключевые результаты (Top-5)
Модель
Направление
chrF++
BLEU
CER
byt5-small
Tj→Fa
87.35 ± 0.10
73.58
0.054
byt5-small
Fa→Tj
80.07 ± 0.23
56.61
0.090
G2PTransformer
Tj→Fa… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-farsi-transliteration-benchmark.
