opendatalab/WanJuan-Arabic
💡 Introduction WanJuan-Arabic(万卷丝路-阿拉伯语) corpus, with a volume exceeding 220GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Arabic.
256
