PersianML/persian-text-corpus
Persian Corpus (Merged) Dataset Summary Persian Corpus (Merged) is a large-scale, Persian corpus meticulously aggregated from multiple high-quality Persian datasets available on the Hugging Face Hub. Designed to advance Persian NLP research and applications, this corpus consolidates diverse textual sources into a single resource, providing researchers and developers with a robust foundation for training and evaluating language models. Why Use This… See the full description on the dataset page: https://huggingface.co/datasets/PersianML/persian-text-corpus.
Persian Corpus (Merged)
Dataset Summary
Persian Corpus (Merged) is a large-scale, Persian corpus meticulously aggregated from multiple high-quality Persian datasets available on the Hugging Face Hub. Designed to advance Persian NLP research and applications, this corpus consolidates diverse textual sources into a single resource, providing researchers and developers with a robust foundation for training and evaluating language models.
Why Use This Corpus?
By merging fragmented Persian datasets, this resource eliminates data collection bottlenecks and provides a unified, ready-to-use corpus. Its scale and diversity make it particularly valuable for improving model generalization in under-resourced languages like Persian.
Scale
Contains 14+ million examples and ~5.1 billion tokens, making it one of the largest publicly available Persian corpora.
Supported Tasks
- Text Generation: Training Persian text generation models.
- Language Modeling: Developing and evaluating Persian language models.
- General Persian NLP: Suitable for a wide range of Persian Natural Language Processing tasks.
Dataset Creation
This dataset was created by merging several existing Persian datasets available on the Hugging Face Hub. This aggregation simplifies access to a large volume of Persian text for research and development.
Explore Related Datasets
For further exploration, consider these other datasets on the Hugging Face Hub that include Persian text:
- CulturaX (fa): 59M rows
- C4 (fa): 54M rows
- FineWeb2-embedded (fas_Arab): 51M rows
- FineWeb2-HQ (fas_Arab): 5M rows
- GlotCC-V1 (fas-Arab): 3M rows
- xP3x: 1.5M rows
- xlsum (persian): 59k rows
- bluesky (fa): 4.5K rows
