datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
yt_full_image_dataset
Dataset Card for "yt_full_image_dataset"
More Information needed
house_kg_full_dataset
house.kg — Kyrgyzstan Real Estate (multimodal)
A complete snapshot of house.kg, the largest real-estate
board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller
identities, agency ratings, reviews — and 227,294 photographs.
Field names are English; values are kept in the original language (Russian/Kyrgyz),
exactly as the site renders them.
💻 Scraper source code on GitHub →
The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.vqasynth_opencv3d_dataset_3600_wotem_v2_fullfull_pose_retrain_datasetspecies-dataset-full-oaktext-2-image-dpo-human-preferences-full
Text-2-Image DPO Human Preferences (Full)
The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference.
This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see:
datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.vqasynth_opencv3d_dataset_3600_v1_fullhouse_kg_full_dataset_frames
house.kg — Kyrgyzstan Real Estate, over time
Sale and rental listings scraped from house.kg, the largest
real-estate board in Kyrgyzstan, re-measured on a schedule. Field names are
English; values are kept in the original language (Russian), exactly as the site
renders them.
Coverage: 2026-09-08. This is the baseline snapshot; later runs append new partitions.
Subsets
subset
rows
description
listings
25,264
one row per advertisement — current state plus… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset_frames.tubitak-olimpiyat-dataset-full-mcq-pp-v4smdg-full-dataset
Dataset Card for Dataset Name
All the images of the dataset come from this kaggle dataset.
Only fundus images have been collected and some minor modifications have been made to the metadata.
All credit goes to the original authors and the contributor on Kaggle.
Dataset Details
Dataset Description
Standardized Multi-Channel Dataset for Glaucoma (SMDG-19) is a collection and standardization of 19 public datasets, comprised of full-fundus glaucoma images and… See the full description on the dataset page: https://huggingface.co/datasets/bumbledeep/smdg-full-dataset.vqa_caption.dataset-fullvqasynth_opencv3d_dataset_3600_v2_fullicis_dataset_full_dataTFQ-Data-Full
TFQ-Data: A Fine-Grained Dataset for Image Implication
TFQ-Data is a large-scale visual instruction tuning dataset specifically designed to train Multi-modal Large Language Models (MLLMs) on Image Implication and Metaphorical Reasoning.
Unlike standard VQA datasets that focus on literal description, TFQ-Data utilizes a True-False Question (TFQ) format. This format provides high knowledge density and verifiable reward signals, making it an ideal substrate for Visual Reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/MING-ZCH/TFQ-Data-Full.vqa_caption.dataset-fullnew_look_dataset_fullGood_OJT_full_aligned_datafull_pose_semantic_localization_dataset_gazeboboxes_full_controlnet_dataset
Dataset Card for "boxes_full_controlnet_dataset"
FWIW, I didn't get good results with this after 20K training steps for some reason, but feel free to give it a shot!
More Information needed
vqasynth_processed_clever_fulldatapublic_smash_full_dataset_07vqasynth_reasoining_opencv3d_dataset_3600_wotem_v2_full_reasoningfull_dataset_augmentedtubitak-olimpiyat-dataset-full-mcq-pp-v2libero-data-64px-fullVisualPRM300Kv0-full-dataset-mc0-o4-judge-incorrect-step-qwen-formatui-dataset-full-part1multimodal-spectroscopic-seed-datasets-fulltubitak-olimpiyat-dataset-full-mcq-pp-v3memes_dataset_full
Dataset Card for "memes_dataset_full"
More Information needed
