datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
HSC-GalaxiesML-VAE-embeddingsda_resample_part2SPY_Prices_Jan20_Mar25da_resample_part1Panitikan-34
Panitikan-34 Corpus
Panitikan-34 contains Filipino literary texts from 34 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary… See the full description on the dataset page: https://huggingface.co/datasets/Galahallt/Panitikan-34.Panitikan-10
Panitikan-10 Corpus
Panitikan-10 contains Filipino literary texts from 10 Filipino authors from the late 19th to early 20th century. The data was gathered from the Tagalog books category of Project Gutenberg using the scrapy library. Various pre-processing techniques were also applied to the dataset which can also be adopted in other languages as discussed in the paper. Dictionaries, thesauruses, and works translated from other languages were excluded to solely focus on literary… See the full description on the dataset page: https://huggingface.co/datasets/Galahallt/Panitikan-10.GalaxyMergerTabularTensorSampleGalaxyMergerTabularSample
