itu
Datasets
All datasets matching “itu”itu
ITU-T source mirror
Part of the Open-Telco Telecom Standards Corpus.Sibling mirrors: 3GPP · ETSI · ITU-T · O-RAN · GSMA · TM Forum · CAMARA
This dataset mirrors the ITU Telecommunication Standardization Sector (ITU-T) Recommendations. Every source is kept twice: original/ is the Recommendation as ITU-T released it (PDF), and marked/ is that same document converted to Markdown for search and retrieval, one raw.md per document with any figures extracted beside it. The two trees… See the full description on the dataset page: https://huggingface.co/datasets/GSMA/itu.image-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.itu-aris-lab-acoustic-drone-dataset
ITU ARIS Lab. Acoustic Drone Dataset
Two-class acoustic dataset for drone detection, recorded outdoors with multiple
drone models at ranges from a few metres to roughly two kilometres, annotated by
hand from a single continuous 3.12 h session.
Classes
drone (4 491), background (1 000)
Segment length
1.0 s
Sampling rate
22 050 Hz, mono, 16-bit PCM
Licence
CC-BY-4.0
Read this before you split the data
Segments overlap by 50 %. drone_0004.wav… See the full description on the dataset page: https://huggingface.co/datasets/imm61/itu-aris-lab-acoustic-drone-dataset.image-vqa-turkish
Türkçe Image VQA Veri Seti
Bu veri seti Türkçe görsel soru-cevap (VQA) çiftleri içermektedir.
Kullanım
from datasets import load_dataset
ds = load_dataset("ituperceptron/turkish-image-vqa", split="vqa")
Veri Yapısı
image: Görsel (PIL Image)
image_id: Görselin benzersiz ID'si
vqa: VQA soru-cevap çiftleri (JSON formatında)
turkish_medical_reasoning
Türkçe Medikal Reasoning Veri Seti
Bu veri seti FreedomIntelligence/medical-o1-verifiable-problem veri setinin Türkçeye çevirilmiş bir alt kümesidir.
Çevirdiğimiz veri seti 7,208 satır içermektedir. Veri setinde bulunan sütunlar aşağıda açıklanmıştır:
question: Medikal soruların bulunduğu sütun.
answer_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce yanıtların Türkçeye çevrilmiş hali.*
reasoning_content: DeepSeek-R1 modeli tarafından oluşturulmuş İngilizce akıl… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish_medical_reasoning.turkish-math-186k
Türkçe Matematik Veri Seti
Bu veri seti AI-MO/NuminaMath-1.5 veri setinin Türkçe'ye çevirilmiş bir alt parçasıdır ve paylaştığımız veri setinde orijinal veri setinden yaklaşık 186 bin satır bulunmaktadır.
Veri setindeki sütunlar ve diğer bilgiler ile ilgili detaylı bilgiye orijinal veri seti üzerinden ulaşabilirsiniz. Problem ve çözümlerin çevirileri için gemini-2.0-flash modeli kullanılmıştır ve matematik notasyonları başta olmak üzere veri setinin çeviri kalitesinin üst düzeyde… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/turkish-math-186k.
