datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Panza-emails
The Panza Emails dataset
This dataset contains collections of emails of three authentic users (david, isabel, and marcus), with personal information (names, places, etc.) replaced by other ones for donor privacy.
Except for these changes, the language of the emails is genuine. The intention of this dataset is to allow researchers to study strategies for text personalization.
The data was donated explicitly for this purpose. This dataset is ethically collected and fully licensed for… See the full description on the dataset page: https://huggingface.co/datasets/ISTA-DASLab/Panza-emails.Turkce-istatistik-reasoning
Türkçe İstatistik Muhakeme (Chain-of-Thought) Veri Seti
Türkçe'de istatistik konularında düşünce zinciri (chain-of-thought) içeren, soru-cevap formatında bir fine-tuning veri seti. Her örnek, bir kullanıcı sorusu ve modelin hem iç muhakeme sürecini (thinking) hem de nihai cevabını içeren bir asistan yanıtından oluşan bir conversations listesidir.
Veri Seti Özeti
Toplam örnek sayısı
400
Dil
Türkçe
Kapsanan modül sayısı
7
Soru tipleri
Kavramsal… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/Turkce-istatistik-reasoning.i-statements
I-Statements
This dataset has axproximently 5,335 I-statements generated by Qwen2.5-7B-Q4_K_M using Ollama.
Stats
Metric
Value
Entries
5,334
Total tokens (GPT2)
36,032
Total words
29,735
Avg. tokens per entry
6.67
Avg. words per entry
5.57
Word range
3–10
Unique vocab (words)
2,237
Unique verbs
252
We used GPT2's tokenizer to find the token count.
Note: The tokens may vary depending on the tokenizer used.
Use Cases
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/i-statements.home-ASS-istant-sharegpt
