CoolFace
Datasetpublic

Galahallt/Panitikan-34

Panitikan-34 Corpus Panitikan-34 contains Filipino literary texts from 34 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary… See the full description on the dataset page: https://huggingface.co/datasets/Galahallt/Panitikan-34.

sourceHugging Facemitupdated 2y agoView on Hugging Face
0likes12downloads
Dataset Card

Panitikan-34 Corpus

Panitikan-34 contains Filipino literary texts from 34 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary texts and the original author's writing style.

A smaller version of the dataset is also provided which only contains the top 10 Filipino authors who had the most literary works in the website. Kindly check it out at this link.

Dataset Specifications

ItemsCount
No. of tokens724,133
Vocabulary Size60,354
No. of literary works47
No. of authors34