datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStories-Farsi
Tiny Stories Farsi
The Tiny Stories Farsi project is a continuous effort to translate the Tiny Stories dataset into the Persian (Farsi) language. The primary goal is to produce a high-quality Farsi dataset, maintaining equivalency with the original English version, and subsequently to utilize it for training language models in Farsi. This seeks to affirm that the advancements and trends observed in English language models are replicable and applicable in other languages. Thus far… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/TinyStories-Farsi.farsick-sts
Dataset Summary
FarSick STS is a Persian (Farsi) dataset designed for the Semantic Textual Similarity (STS) task. It is a part of the FaMTEB (Farsi Massive Text Embedding Benchmark). The dataset was developed by translating and adapting the English SICK (Sentences Involving Compositional Knowledge) dataset, and it features Persian sentence pairs annotated for their degree of semantic relatedness.
Language(s): Persian (Farsi)
Task(s): Semantic Textual Similarity (STS)
Source:… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/farsick-sts.tajik-farsi-transliteration-benchmark
🇹🇯🇮🇷 Tajik-Farsi Transliteration Benchmark
Официальный бенчмарк машинной транслитерации между таджикским (кириллица) и фарси (персо-арабская графика).Результаты получены на 40k параллельных предложениях с оценкой по 3 случайным сидам, bootstrap 95% CI и парными статистическими тестами.
📊 Ключевые результаты (Top-5)
Модель
Направление
chrF++
BLEU
CER
byt5-small
Tj→Fa
87.35 ± 0.10
73.58
0.054
byt5-small
Fa→Tj
80.07 ± 0.23
56.61
0.090
G2PTransformer
Tj→Fa… See the full description on the dataset page: https://huggingface.co/datasets/TajikNLPWorld/tajik-farsi-transliteration-benchmark.Farsi-Greek-parallel-data
Farsi–Greek Parallel Corpus
Dataset Summary
This dataset contains 15.6k sentence pairs aligned between Farsi (fa) and Greek (el).The data was collected from multiple sources and processed for use in machine translation and cross-lingual NLP research.
Languages: Farsi (fa), Greek (el)
Tasks: Translation (Farsi ↔ Greek)
Size: 15,600 parallel pairs
Splits: train (80%), validation (10%), test (10%)
License and Attribution
This dataset is shared under… See the full description on the dataset page: https://huggingface.co/datasets/81aikmastrofo/Farsi-Greek-parallel-data.Data-farsi
DENLI_AI - Professional Synthetic Dataset
In the name of God
📌 Overview
This dataset has been meticulously crafted for training an Offline AI Model that is planned to be developed. It contains rich metadata including names, creators, and all essential details that require minimal editing — ready to use out of the box.
The data is built to be rich, professional, and high-quality. Simply use it, and witness a miracle in your model’s performance.
Thank you for your… See the full description on the dataset page: https://huggingface.co/datasets/ALIASGHARQADIRI/Data-farsi.Intent-Sub_intent-farsi
