datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SpaRRTa
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) —
such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP — encode the spatial relations
between objects in a scene, rather than only their semantic identity.
📄 Paper: arXiv:2601.11729
💻 Code: github.com/gmum/SpaRRTa
🧱 Real-world (lego) split: turhancan97/SpaRRTa-Lego
🔬 Attention-analysis split (images +… See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa.synthetic-turkish-passports
Turkish passport dataset - 5, 000 images
Dataset comprises 5,000 meticulously organized files capturing Turkish passports under highly controlled variations, making it an invaluable resource for developing robust document recognition and verification systems. It is specifically designed for training and testing models in passport authentication, biometric data extraction, and identity verification.
By leveraging this dataset containing detailed information from Turkish… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-turkish-passports.sdxl-turbo-sae-labels
SDXL-Turbo SAE Feature Labels
20,480 labeled sparse autoencoder features across 4 UNet attention blocks in SDXL-Turbo, plus 50K generated images with full activation logs.
Built for latent-dance — a real-time audio-reactive music visualizer using SAE steering at 50 FPS.
Important attribution: The SDXL-Turbo sparse autoencoders/checkpoints used here were trained and released by Surkov et al. / EPFL through sdxl-unbox. This dataset does not claim authorship of the SAE training. It… See the full description on the dataset page: https://huggingface.co/datasets/hammamiomar/sdxl-turbo-sae-labels.wroclaw-dwarves
Wrocław Dwarves — Fine-Grained Instance Retrieval
Wrocław is scattered with several hundred small bronze dwarf statues — krasnale — installed
across the city since 2001. They are individually sculpted, but they share a visual vocabulary:
the same scale, the same material, the same crouching poses and hand-held props. Telling one
from another is therefore a fine-grained instance recognition problem rather than a
classification one. Every statue here belongs to the same semantic… See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/wroclaw-dwarves.TurkishFoods-25
⚠️ Important: English version is available below.
TürkSofrası-25 (TurkishFoods-25) Veri Seti
TürkSofrası-25, 25 farklı geleneksel Türk yemeğine ait toplam 11.461 görsel içeren ve yemek tanıma/sınıflandırma amaçlı hazırlanmış bir görüntü veri setidir. Görseller .jpg formatında olup her sınıf için ayrı klasörlerde yer almaktadır.
Veri seti, Hugging Face datasets kütüphanesi biçimindedir ve image (görsel) ile label (etiket) olmak üzere iki özelliğe sahiptir. Etiketler class_label… See the full description on the dataset page: https://huggingface.co/datasets/yunusserhat/TurkishFoods-25.Turkish-VLM-Mix-BenchmarkThis is a Turkish multimodal (image-text-text triplets) dataset consisting of Turkish translated samples from the datasets google/docci, tomg-group-umd/pixelprose, detection-datasets/coco, rafaelpadilla/coco2017, liuhaotian/LLaVA-Instruct-150K, liuhaotian/LLaVA-CC3M-Pretrain-595K, and HuggingFaceM4/FairFace.
The labels are in Turkish and the dataset is in an instruction-tuning format with separate columns for prompts and completion labels.
The original labels (except… See the full description on the dataset page: https://huggingface.co/datasets/ucsahin/Turkish-VLM-Mix-Benchmark.Turkish-Visual-Reasoning-Dataset
Turkish Visual Reasoning Dataset
The Turkish Visual Reasoning Dataset is a Turkish multimodal reasoning dataset designed to evaluate and improve the abstract reasoning capabilities of Vision-Language Models (VLMs).
It was created by adapting established visual reasoning benchmarks into Turkish and combining them with original Turkish BİLSEM preparation questions. The dataset targets challenging reasoning tasks such as logical pattern discovery, spatial reasoning, analogical… See the full description on the dataset page: https://huggingface.co/datasets/Berkesule/Turkish-Visual-Reasoning-Dataset.TurkishFoods-15
⚠️ Important: English version is available below.
TürkSofrası-15 (TurkishFoods-15) Veri Seti
TürkSofrası-15, 15 farklı geleneksel Türk yemeğine ait toplam 7.411 görsel içeren ve yemek tanıma/sınıflandırma amaçlı hazırlanmış bir görüntü veri setidir. Görseller .jpg formatında olup her sınıf için ayrı klasörlerde yer almaktadır.
Veri seti, Hugging Face datasets kütüphanesi biçimindedir ve image (görsel) ile label (etiket) olmak üzere iki özelliğe sahiptir. Etiketler class_label… See the full description on the dataset page: https://huggingface.co/datasets/yunusserhat/TurkishFoods-15.Farfetch.Product.prices.Turkey
Farfetch web scraped data
About the website
The Ecommerce industry in the EMEA region, specifically in Turkey, has shown significant growth in recent years. Companies like Farfetch have established their presence in the competitive Turkish market. Turkeys rapid digital transformation, favourable demographics, and high levels of internet penetration have resulted in the boom of online retail. This proves profitable for ecommerce platforms specialising in luxury fashion… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Farfetch.Product.prices.Turkey.turmeric_leaf_disease_classification
Turmeric Leaf Disease Classification
A dataset for disease classification of turmeric leaves. The dataset contains raw and augmented versions.The raw dataset contains 865 images.Images per class:
Aphids_Disease: 221
Blotch: 238
Healthy_Leaf: 213
Leaf_Spot: 193
The augmented dataset contains 3,496 images.Images per class:
Aphids_Disease: 847
Blotch: 909
Healthy_Leaf: 821
Leaf_Spot: 919
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/turmeric_leaf_disease_classification.SpaRRTa-Lego
SpaRRTa-Lego: Real-World Split of the SpaRRTa Spatial-Relation Benchmark
SpaRRTa-Lego is the real-world counterpart of the synthetic
SpaRRTa benchmark. Scenes are
photographed with toy minifigures and everyday objects, and used for sim-to-real evaluation
of the spatial-relation capabilities of Visual Foundation Models.
📄 Paper: arXiv:2601.11729
💻 Code: github.com/gmum/SpaRRTa
🧩 Synthetic split: turhancan97/SpaRRTa
The task
A 4-way classification problem — Front / Back… See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa-Lego.Net.a.Porter.Product.prices.Turkey
Net-a-Porter web scraped data
About the website
The Net-a-Porter company operates within the Ecommerce industry in the EMEA region, particularly in Turkey. This expansive sector primarily focuses on the buying and selling of goods and services through the internet, showcasing an array of products from various providers on a global scale. Over the past few years, the Ecommerce sector in Turkey has experienced rapid growth, leading to a highly competitive market. Companies… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.Turkey.turmeric_disease_classification
Turmeric Disease Classification
A dataset for disease classification of turmeric. The dataset contains raw and augmented versions.The raw dataset contains 1,063 images.Images per class:
Dry Leaf: 203
Healthy Leaf: 197
Leaf Blotch: 199
Rhizome Disease Root: 182
Rhizome Healthy Root: 282
The augmented dataset contains 4,548 images.Images per class:
Dry Leaf: 812
Healthy Leaf: 985
Leaf Blotch: 995
Rhizome Disease Root: 910
Rhizome Healthy Root: 846
This dataset is indexed on… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/turmeric_disease_classification.turcoins
Dataset Description
We present a novel and comprehensive dataset which contains Turkish Republic coins minted since 1924.
The proposed dataset consists of 11080 coin images from 138 different classes. All images are tightly cropped, RGB and 256x256.
Point of Contact: Huseyin Temiz
Citation Information
@inproceedings{temiz2021turcoins,
title={TurCoins: Turkish republic coin dataset},
author={Temiz, H{\"u}seyin and G{\"o}kberk, Berk and Akarun, Lale}… See the full description on the dataset page: https://huggingface.co/datasets/hsyntemiz/turcoins.
