datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GreekWikipedia
GreekWikipedia
A Greek abstractive summarization dataset collected from the Greek part of Wikipedia, which contains 93,432 articles, their titles and summaries.
This dataset has been used to train our best-performing model GreekWiki-umt5-base as part of our research paper:Giarelis, N., Mastrokostas, C., & Karacapilidis, N. (2024) Greek Wikipedia: A Study on Abstractive Summarization.For information about dataset creation, limitations etc. see the original article.… See the full description on the dataset page: https://huggingface.co/datasets/IMISLab/GreekWikipedia.GreekReddit
GreekReddit
A Greek topic classification dataset collected from Greek subreddits, which contains 6,534 posts, their titles and topic labels.
This dataset has been used to train our best-performing model Greek-Reddit-BERT as part of our research article:
Mastrokostas, C., Giarelis, N., & Karacapilidis, N. (2024). Social Media Topic Classification on Greek Reddit
For information about dataset creation, limitations etc. see the original article.
Supported Tasks and… See the full description on the dataset page: https://huggingface.co/datasets/IMISLab/GreekReddit.Recipes_Greek
