CoolFace
Datasetpublic

GoktugD/turkish-keyword-extraction-500k

Turkish Keyword Extraction 500K v2 Yirmi alanda konu ve anahtar sözcük çıkarımı için kısa Türkçe belgeler. Doğrulanmış boyut Train: 490,000 Validation: 5,000 Test: 5,000 Toplam: 500,000 Ana görev sütunları: id, text, keywords, domain Provenance Veri insan mesajlarından, belgelerinden veya web kazımasından alınmamıştır. Tamamı depodaki üretici koduyla deterministik olarak oluşturulur. Her satırda source_type, provenance, generator_version… See the full description on the dataset page: https://huggingface.co/datasets/GoktugD/turkish-keyword-extraction-500k.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes179downloads
5 commits on main
c38fd1a2mo ago

Release v2.0 with provenance and quality audits

GoktugD
48c50fd2mo ago

Release v2.0 with provenance and quality audits

GoktugD
1b46ad42mo ago

Scale dataset to 500,000 validated rows in sharded Parquet

GoktugD
32c860d2mo ago

Publish Turkish synthetic dataset with train/test splits

GoktugD
3e052e92mo ago

initial commit

GoktugD