CoolFace
4 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01afkfatih /turkish-gemma-51k Turkish Chat Dataset - Gemma Format Bu dataset, Türkçe sohbet ve talimat takip etme görevleri için hazırlanmış 51.914 konuşma örneği içerir. 📊 Dataset Özeti Dil: Türkçe Format: Chat/Conversation Örnek Sayısı: 51,914 Kaynak: afkfatih/turkishdataset 🎯 Kullanım Alanları Türkçe sohbet botları eğitimi Instruction-tuning Fine-tuning LLM modelleri (Gemma, Llama, vb.) Türkçe doğal dil anlama 📝 Format Her örnek şu yapıya sahiptir: [ { "role":… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-gemma-51k.texttext-generation10K<n<100K1 likes37 downloads1y agoHugging Face02afkfatih /guppylm-turmix-data GuppyLM TurMix Tokenized Dataset This dataset contains the highly curated TurMix Turkish corpus, pre-tokenized using the custom GuppyLM 50K ByteLevel BPE Tokenizer. Not: HuggingFace üzerindeki veri önizleme (Dataset Viewer) özelliği kasten kapatılmıştır (viewer: false). Çünkü bu veri seti, PyTorch eğitimlerinde RAM tasarrufu sağlamak amacıyla doğrudan bellek haritalaması (np.memmap) yapılabilen ham uint16 ikili (binary) formatta (.bin) kaydedilmiştir. 📊 Veri Seti… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/guppylm-turmix-data.text-generation10B<n<100B0 likes31 downloads4mo agoHugging Face03afkfatih /turkish-distilled-5K Turkish Distilled SFT Dataset Temizlenmiş, tek turlu ve messages şemasına sahip Türkçe instruction-tuning verisi. Dataset Summary Kaynak: data/distilled_train.jsonl Şema: messages Toplam temiz örnek: 5957 Train örnek sayısı: 5660 Test örnek sayısı: 297 Bu veri kümesi Unsloth, TRL ve Hugging Face datasets akışlarıyla uyumlu olacak şekilde hazırlanmıştır. Features Her satır şu alanları içerir: id messages source task language quality_score flags… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-distilled-5K.texttext-generation1K<n<10K0 likes29 downloads5mo agoHugging Face04afkfatih /turkish-cpt-dataset Turkish CPT Dataset A high-quality Turkish + English dataset for Continued Pre-Training (CPT) of language models. Dataset Summary Property Value Total examples 1,908,378 Total tokens ~2.19B Turkish ratio ~80% English ratio ~20% Languages Turkish, English Sources Source Language Examples Description wikimedia/wikipedia (tr) TR ~534K Turkish Wikipedia wikimedia/wikipedia (en) EN ~134K English Wikipedia (20% replay)… See the full description on the dataset page: https://huggingface.co/datasets/afkfatih/turkish-cpt-dataset.texttext-generation1M<n<10M0 likes26 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.