datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lajme-shqip
Lajme Shqip - Albanian News Corpus
Dataset Description
A large-scale Albanian-language news corpus containing over 1.6 million article summaries (~500 MB).
Albanian is a low-resource language with limited publicly available datasets for NLP research. This corpus helps bridge that gap by offering diverse, real-world text data from multiple news sources and categories.
Supported Tasks
Language Modeling / Text Generation: Pre-train or fine-tune… See the full description on the dataset page: https://huggingface.co/datasets/akadriu/lajme-shqip.qysh-me-shqip-scrape-datasetA small/complete web scrape of the old qysh.me website now hosted on dua.com.
