datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-exp031-stateturkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.swefaqpathinen_keezhkanakku-kaarnarpadhu
📚 Dataset Card: கார் நாற்பது (Kaarnarpadhu)
Dataset Summary
கார் நாற்பது (Kaarnarpadhu) is a classical Tamil poetic work belonging to the Pathinen Keezhkanakku tradition. The text derives its name from two defining characteristics:
It consists of 40 poems (நாற்பது செய்யுட்கள்)
Each poem describes the arrival and nature of the monsoon season (கார் காலம்)
Thus, the work came to be known as Kaar Narpadhu.
Title: கார் நாற்பது
Text Type: Seasonal & Emotional Poetry… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/pathinen_keezhkanakku-kaarnarpadhu.
