Reza2kn/RaahNaameh-1-textual-corpus
RaahNaameh-1 Textual Corpus A large-scale Persian text corpus assembled for training the RaahNaameh-1 embedding model. Sources Source Sentences Description Jomleh 1,002,221 Formal Persian web text LSCP 10,257,866 Iranian tweets — colloquial, slang, emoji Persian Wikipedia 1,107,618 Encyclopedic articles Total 12,367,705 Processing Light normalization only: Arabic→Persian character mapping, zero-width space removal Emojis… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/RaahNaameh-1-textual-corpus.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face