ali619/corpus-dataset-normalized-for-persian-farsi
Dataset Summary Persian data of this dataset is a collection of 400k blog posts (RohanAiLab/persian_blog). these posts have been gathered from more than 10 websites. This dataset can be used in different NLP tasks like language modeling, creating tokenizer and text generation tasks. The data in this dataset have been normalized and unnecessary tokens have been removed. Note: If you need Persian and Engish corpus together, click here
578
