datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
refusal-exp031-stateeesti-kaanamiskorpus
eesti-kaanamiskorpus — Estonian inflection corpus (11,011 entries)
Word and phrase inflections with case, number, source and licence per entry.
Parallel forms kept as separate entries (crucial for fair scoring!). Built
from Riigikogu stenograms + ERR (CC-BY-SA) and Vabamorf rule-based synthesis
with round-trip validation.
On licences, stated plainly. 11,011 of 13,436 entries are published here;
entries with unresolved rights are withheld from this dataset. Withheld from… See the full description on the dataset page: https://huggingface.co/datasets/pertai/eesti-kaanamiskorpus.turkish-wikipedia-dataset-clean
Turkish Wikipedia Dataset
A cleaned and structured Turkish Wikipedia dataset designed for Turkish language model pretraining, continued pretraining, research, and NLP experiments.
The dataset consists of articles collected from the Turkish Wikipedia (tr.wikipedia.org) and processed into a machine-readable format while preserving important source metadata.
Dataset Summary
Language: Turkish (tr)
Source: Turkish Wikipedia
Domain: General knowledge / encyclopedia… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset-clean.sugarcrm_130_documentation
Source: Sugarcrm 13.0 Dev Documentation
The chunks in the files are diffrent splittet based on the tokenizer conained in the name of the file
cl100k_base: 400 Tokens per chunk
p50k_base: 200 Tokens per chunk
aimperum_kaappiyangal-seevaga_chintamani
📕 Sivaga Chintamani Dataset (சீவக சிந்தாமணி தரவுத்தொகுப்பு)
🧾 Dataset Summary
Sivaga Chintamani (சீவக சிந்தாமணி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics) and is considered the earliest epic chronologically among them.
The epic was composed in Tamil by adapting several Sanskrit Sivagan legends. The original source is believed to be a work known as “Kshatriya Chudamani”.
This dataset presents a structured digital version of Sivaga Chintamani… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-seevaga_chintamani.aimperum_kaappiyangal-kundalakesi
📙 Kundalakesi Dataset (குண்டலகேசி தரவுத்தொகுப்பு)
🧾 Dataset Summary
Kundalakesi (குண்டலகேசி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics).
This epic is distinct in its strong focus on religious debate, renunciation, and philosophical transformation. The work survives only in a fragmentary form, with a limited number of verses available today.
This dataset presents a structured digital collection of all the extant verses of Kundalakesi, preserved in their… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-kundalakesi.carzi-tr-knowledge-base
carzi-tr-knowledge-base
Türkiye merkezli bulut oto servis programı Carzi hakkında Türkçe bilgi bankası.
Kayıt: 106
Boyut: 198.3 KB
Lisans: CC-BY-4.0
Dil: tr
İçerik
Marka / ürün özeti
Blog makaleleri (carzi.com.tr/yazilar)
SEO landing sayfaları
Özellik sayfaları
Sektör çözüm sayfaları
SSS
Özellik & içerik kitabı bölümleri
Dosya
data.jsonl — her satır bir JSON nesnesi:
id, title, url, text, summary, language, category, keywords, published_at… See the full description on the dataset page: https://huggingface.co/datasets/kaancan404/carzi-tr-knowledge-base.swefaqTrnsfr_ALPACAEXAMPLErefusal-exp012-specdecturkish-wikipedia-dataset
Türkçe Kamu Kurumları ve Tarih Sohbet Veri Seti
Bu veri seti, Türkiye'deki kamu kurumları, bakanlıklar, devlet organları, resmi semboller ve tarihi figürler hakkında yapılandırılmış Türkçe sohbet verileri içermektedir. Veriler, güvenilir ve tarafsız bir kaynak olan Türkçe Vikipedi'den otomatik olarak çıkarılmış ve büyük dil modellerini (LLM) ince ayar (fine-tuning) için uygun bir formata dönüştürülmüştür.
Her bir örnek, bir "sistem" talimatı, bir "kullanıcı" sorgusu ve bir… See the full description on the dataset page: https://huggingface.co/datasets/kaan39/turkish-wikipedia-dataset.pathinen_keezhkanakku-kaarnarpadhu
📚 Dataset Card: கார் நாற்பது (Kaarnarpadhu)
Dataset Summary
கார் நாற்பது (Kaarnarpadhu) is a classical Tamil poetic work belonging to the Pathinen Keezhkanakku tradition. The text derives its name from two defining characteristics:
It consists of 40 poems (நாற்பது செய்யுட்கள்)
Each poem describes the arrival and nature of the monsoon season (கார் காலம்)
Thus, the work came to be known as Kaar Narpadhu.
Title: கார் நாற்பது
Text Type: Seasonal & Emotional Poetry… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/pathinen_keezhkanakku-kaarnarpadhu.refusal-exp011-deepinceptionaimperum_kaappiyangal-silappadhikaram
📚 Silappathikaram Dataset (சிலப்பதிகாரம் தரவுத்தொகுப்பு)
🧾 Dataset Summary
Silappathikaram (சிலப்பதிகாரம்) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics).
This dataset presents a structured digital representation of the epic, preserving its literary, cultural, and ethical significance for modern computational use.
The epic was composed by Ilango Adigal (இளங்கோ அடிகள்), the brother of the Chera king Senguttuvan. Renouncing royal life, Ilango Adigal embraced asceticism… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-silappadhikaram.aimperum_kaappiyangal-manimekalai
📘 Manimekalai Dataset (மணிமேகலை தரவுத்தொகுப்பு)
🧾 Dataset Summary
Manimekalai (மணிமேகலை) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics).
It is closely connected to Silappathikaram, sharing the same historical setting, characters, and timeline. Because of this close relationship, Silappathikaram and Manimekalai are referred to as twin epics (இரட்டைக் காப்பியங்கள்).
This dataset presents a structured digital version of the Manimekalai epic, preserving the… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-manimekalai.aimperum_kaappiyangal-valaiyapathi
📗 Valaiyapathi Dataset (வளையாபதி தரவுத்தொகுப்பு)
🧾 Dataset Summary
Valaiyapathi (வளையாபதி) is one of the Aimperum Kaappiyangal (Five Great Tamil Epics).
Unlike the other epics, much of the information about Valaiyapathi—such as the author’s name, period of composition, the hero’s name, and the complete storyline—remains unknown due to the loss of the original manuscript.
Only 72 verses of this epic are available today. This dataset brings together all the surviving… See the full description on the dataset page: https://huggingface.co/datasets/TamilThagaval/aimperum_kaappiyangal-valaiyapathi.cge-fine-tunecge-finetunekaa_sentences
