datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Turkce-istatistik-benchmark
Türkçe İstatistik Benchmark
Bu veri seti, Toivo0/Turkce-istatistik-reasoning ana veri setiyle fine-tuning'de kesinlikle kullanılmamış 100 soru-cevap çiftinden oluşur. Fine-tune edilmiş modellerin gerçek performansını, eğitimde hiç görmediği sorularla ölçmek için hazırlanmıştır.
Benchmark iki aşamada oluşturulmuştur:
37 soru, ana veri setinden (400 soru) eğitim öncesinde stratified (modül bazında orantılı) olarak ayrılmış bölümdür.
63 soru, benchmark'ı istatistiksel olarak daha… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/Turkce-istatistik-benchmark.Turkce-istatistik-reasoning
Türkçe İstatistik Muhakeme (Chain-of-Thought) Veri Seti
Türkçe'de istatistik konularında düşünce zinciri (chain-of-thought) içeren, soru-cevap formatında bir fine-tuning veri seti. Her örnek, bir kullanıcı sorusu ve modelin hem iç muhakeme sürecini (thinking) hem de nihai cevabını içeren bir asistan yanıtından oluşan bir conversations listesidir.
Veri Seti Özeti
Toplam örnek sayısı
400
Dil
Türkçe
Kapsanan modül sayısı
7
Soru tipleri
Kavramsal… See the full description on the dataset page: https://huggingface.co/datasets/Toivo0/Turkce-istatistik-reasoning.i-statements
I-Statements
This dataset has axproximently 5,335 I-statements generated by Qwen2.5-7B-Q4_K_M using Ollama.
Stats
Metric
Value
Entries
5,334
Total tokens (GPT2)
36,032
Total words
29,735
Avg. tokens per entry
6.67
Avg. words per entry
5.57
Word range
3–10
Unique vocab (words)
2,237
Unique verbs
252
We used GPT2's tokenizer to find the token count.
Note: The tokens may vary depending on the tokenizer used.
Use Cases
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Harley-ml/i-statements.
