datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.Cauldron-JA
Dataset Card for The Cauldron-JA
Dataset description
The Cauldron-JA is a Vision Language Model dataset that translates 'The Cauldron' into Japanese using the DeepL API. The Cauldron is a massive collection of 50 vision-language datasets (training sets only) that were used for the fine-tuning of the vision-language model Idefics2.
To create a Japanese Vision Language Dataset, datasets related to OCR, coding, and graphs were excluded because translating them into Japanese… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Cauldron-JA.gaussian-surfels-dtuharis-weapon-detection-dataset-curatedJapan-Open-Driving-Dataset-Sample
Japan Open Driving Dataset Sample
Overview
This repository contains a sample subset of the Japan Open Driving Dataset, a large-scale autonomous driving dataset comprising over 100 hours of driving data collected in Tokyo, Japan.
The data is stored in nuScenes format and can be loaded with the nuscenes-devkit.
In addition to sensor data and 3D annotations, this dataset includes virtual captioned data for training Vision-Language-Model (VLM) and Vision-Language-Action (VLA)… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/Japan-Open-Driving-Dataset-Sample.Krea-2-Turbo-Checkpoint-Format-Benchmark
Krea 2 Turbo ComfyUI Format Fidelity Benchmark
This release is a paired, deterministic comparison of eight Krea 2 Turbo checkpoint formats in ComfyUI: BF16, FP8 Scaled, INT8 ConvRot, MXFP8, NVFP4, INT4 ConvRot W4A4, GGUF Q8_0, and GGUF Q4_K_M. It contains 240 scored 1024×1024 images, saved float32 decoded tensors and final latents, every denoising trajectory, raw metric tables, telemetry, statistical comparisons, and reproduction code.
Main result
BF16 is the… See the full description on the dataset page: https://huggingface.co/datasets/Merserk/Krea-2-Turbo-Checkpoint-Format-Benchmark.gaussian-surfels-bmvsPic-CoSTRIDE-QA-Dataset
STRIDE-QA Dataset
📦 Dataset
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
Category
Description
Object-centric Spatial QA
Spatial relations between two… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset.Easy-Turn-Trainset
Easy Turn: Integrating Acoustic and Linguistic Modalities for Robust Turn-Taking in Full-Duplex Spoken Dialogue Systems
Guojian Li1, Chengyou Wang1, Hongfei Xue1,
Shuiyuan Wang1, Dehui Gao1, Zihan Zhang2,
Yuke Lin2, Wenjie Li2, Longshuai Xiao2,
Zhonghua Fu1,╀, Lei Xie1,╀
1 Audio, Speech and Language Processing Group (ASLP@NPU), Northwestern Polytechnical University
2 Huawei Technologies, China
🎤 Demo Page
🤖 Easy Turn Model
📑 Paper
🌐 Huggingface… See the full description on the dataset page: https://huggingface.co/datasets/ASLP-lab/Easy-Turn-Trainset.Vision-SR1-47KPartImageNetPartImageNet: A Large, High-Quality Dataset of Parts (ECCV 2022)
Project page: https://github.com/TACJu/PartImageNet
arXiv: https://arxiv.org/abs/2112.00933
project_filesRIO-Bench
Read or Ignore? A Unified Benchmark for Typographic-Attack Robustness and Text Recognition in Vision-Language Models
Real-world VLMs must decide when to read text and when to ignore it, e.g., reading traffic signs but not being fooled by text-based attacks on objects.
We propose a unified benchmark, RIO-Bench, to evaluate both typographic-attack robustness and text recognition in VLMs through a novel task called RIO-VQA.
Problem Settings: VLMs Must Adaptively Read… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/RIO-Bench.z-image-turbo-gensingle_turn
InterSyn: A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
This dataset card accompanies the paper
A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text GenerationYukang Feng, Jianwen Sun, Chuanhao Li, Zizhen Li, Jiaxin Ai, Fanrui Zhang, Yifan Chang, Sizhuo Zhou, Shenglin Zhang, Yu Dai, Kaipeng Zhang (2025)
🧠 Introduction
TL;DR InterSyn is a high-quality dataset for instruction‑following, interleaved image–text… See the full description on the dataset page: https://huggingface.co/datasets/finyorko/single_turn.Rex-Omni_muti_turn
Detect Anything via Next Point Prediction
Rex-Omni is a 3B-parameter Multimodal Large Language Model (MLLM) that redefines object detection and a wide range of other visual perception tasks as a simple next-token prediction problem.
News 🎉
[2025-10-31] We release the AWQ quantized version of Rex-Omni, which saves 50% of the storage space. Rex-Omni-AWQ
[2025-10-29] Fine-tuning code is now available.… See the full description on the dataset page: https://huggingface.co/datasets/qq-2/Rex-Omni_muti_turn.file_dataimage-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.STRIDE-QA-Dataset-Mini
STRIDE-QA-Dataset-Mini
STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning in autonomous driving. Constructed from 100 hours of multi-sensor driving data in Tokyo, it offers 16 M QA pairs over 270 K frames with dense annotations including 3D bounding boxes, segmentation masks, and multi-object tracks.
⚠️ Note: STRIDE-QA-Dataset-Mini is provided as a preliminary version and does not fully match the format of the… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Dataset-Mini.TurbulentChannel
PIVtools Turbulent Channel Validation Dataset
Reproducibility capsule for the PIVtools software paper (SoftwareX, submitted). It contains
everything needed to regenerate the paper's validation figures, at three levels of effort:
Plot only — run two scripts against the shipped reference_results/. No PIV processing.
Reprocess — run PIV → calibration → statistics from the shipped images and configs,
then plot from your own outputs.
Re-derive calibration — recompute the camera… See the full description on the dataset page: https://huggingface.co/datasets/MTT69/TurbulentChannel.Vision-SR1-Cold-9KThe dataset_info.json contains all available datasets. If you are using a custom dataset, please make sure to add a dataset description in dataset_info.json and specify dataset: dataset_name before training to use it.
The dataset_info.json file should be put in the dataset_dir directory. You can change dataset_dir to use another directory. The default value is ./data.
Currently we support datasets in alpaca and sharegpt format. Allowed file types include json, jsonl, csv, parquet, arrow.… See the full description on the dataset page: https://huggingface.co/datasets/LMMs-Lab-Turtle/Vision-SR1-Cold-9K.turkish_clip_dataset_with_text_embeddingsThis dataset cleaned and dowloaded version of following dataset: https://huggingface.co/datasets/visheratin/laion-coco-nllb
The main purpose was to extract Turkish captions and download images.
You can use this dataset to fine-tune or create a clip model.
Since there English and Turkish captions you can also use those to create language model?
robocasa_turnonmicrowave_lerobotThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 3000,
"total_frames": 702401,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 3,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:3000"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/daixianjie/robocasa_turnonmicrowave_lerobot.SpaRRTa
SpaRRTa: A Synthetic Benchmark for Evaluating Spatial Intelligence in Visual Foundation Models
SpaRRTa is a synthetic benchmark that probes whether Visual Foundation Models (VFMs) —
such as DINO, DINOv2/v3, MAE, CroCo, VGGT, SPA and CLIP — encode the spatial relations
between objects in a scene, rather than only their semantic identity.
📄 Paper: arXiv:2601.11729
💻 Code: github.com/gmum/SpaRRTa
🧱 Real-world (lego) split: turhancan97/SpaRRTa-Lego
🔬 Attention-analysis split (images +… See the full description on the dataset page: https://huggingface.co/datasets/turhancan97/SpaRRTa.STRIDE-QA-Bench
STRIDE-QA-Bench
STRIDE-QA-Bench provides a standardized benchmark for evaluating spatiotemporal reasoning of Vision-Language Models (VLMs) in autonomous driving.This HuggingFace repository provides the images and JSON files of the benchmark.
For detailed benchmark description and execution code, please refer to STRIDE-QA-Dataset (GitHub).
🗂️ Data Fields
The main data fields are as follows.
Field
Type
Description
question_id
str
Unique question ID.… See the full description on the dataset page: https://huggingface.co/datasets/turing-motors/STRIDE-QA-Bench.image-vqa-turkish
Türkçe Image VQA Veri Seti
Bu veri seti Türkçe görsel soru-cevap (VQA) çiftleri içermektedir.
Kullanım
from datasets import load_dataset
ds = load_dataset("ituperceptron/turkish-image-vqa", split="vqa")
Veri Yapısı
image: Görsel (PIL Image)
image_id: Görselin benzersiz ID'si
vqa: VQA soru-cevap çiftleri (JSON formatında)
crag-mm-single-turn-public
CRAG-MM: Comprehensive multi-modal, multi-turn RAG Benchmark
This repository contains the CRAG-MM dataset, a high-quality conversational benchmark for multimodal assistants. The dataset features conversations about images with varied complexity levels, designed to evaluate AI systems' visual understanding and conversational abilities.
CRAG-MM is a visual question-answering benchmark that focuses on factual questions, offering a unique collection of image and question-answering sets… See the full description on the dataset page: https://huggingface.co/datasets/crag-mm-2025/crag-mm-single-turn-public.synthetic-turkish-passports
Turkish passport dataset - 5, 000 images
Dataset comprises 5,000 meticulously organized files capturing Turkish passports under highly controlled variations, making it an invaluable resource for developing robust document recognition and verification systems. It is specifically designed for training and testing models in passport authentication, biometric data extraction, and identity verification.
By leveraging this dataset containing detailed information from Turkish… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/synthetic-turkish-passports.qwen36-mtp-turbo-kv-analysis
Qwen3.6 MTP Turbo KV Runtime Analysis
This repository is a curated analysis artifact for local Qwen3.6-35B-A3B MTP GGUF inference experiments on Windows CUDA. It compares clean MTP llama.cpp, QuinsZouls llama-next TurboQuant, and the completed subset of Atomic TurboQuant runs under a fixed 64k context, MoE CPU offload, and Unsloth-aligned sampling settings.
The raw benchmark runs included incomplete and capability-incompatible rows. This repo keeps only completed, comparable… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-mtp-turbo-kv-analysis.
