datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
T2I-CoReBench-Images
T2I-CoReBench-Images
📖 Overview
T2I-CoReBench-Images is the companion image dataset of T2I-CoReBench. It contains images generated using 1,080 challenging prompts, covering both composition and reasoning scenarios undere real-world complexities.
This dataset is designed to evaluate how well current Text-to-Image (T2I) models can not only paint (produce visually consistent outputs) but also think (perform reasoning over causal chains, object relations, and logical… See the full description on the dataset page: https://huggingface.co/datasets/lioooox/T2I-CoReBench-Images.gs-images-v2sdxl_images_easy_prompts-artists-seed1mint-1t-html-images-gte6-sample
Size: 6769158 images sampled from Mint-1t-html
Criteria: Data entries with greater than or equal to 6 images (gte6)
Swift-OpenX-Embodiment-wrist-imagessample-images-TADNEsdxl_images_sb_prompts-multi_artist-seed1gbif-plants-images-1m-wdsMinecraft-Depth-Images
D4MCDataset
Dataset created to train the Depth4MC model.
Images and depth label are of size 480x854.
Read more information on the GitHub page: https://github.com/JulianBvW/Depth4MC
sd2hd_images_v2WildDet3D-V3Det-imagesTADNE-sample-images
TADNE sample images
Images generated by the TADNE model.
Note
prediction_results/anime-face-detector
https://github.com/hysts/anime-face-detector
YOLOv3 + HRNetV2
prediction_results/deepdanbooru
https://github.com/KichangKim/DeepDanbooru
model-resnet_custom_v3.h5
prediction_results/deepdanbooru/intermediate_features
Output by the following model
4096-dim
def create_model() -> tf.keras.Model:
path = huggingface_hub.hf_hub_download('hysts/DeepDanbooru'… See the full description on the dataset page: https://huggingface.co/datasets/hysts/TADNE-sample-images.chess-imagesThe Handwritten Chess Scoresheet Dataset contains a set of single and double paged chess scoresheet images with ground truth labels for training and testing.
Images are named as follows: [Game #]_pg[page #].png
Ground truth labels are formatted as follows: [Game #][page #][move #]_[black/white] [ground truth]
Note: ground truth labels for testing - found in "testing_tags.txt" - do not include a page number as they are ground truths for the game represented by the corresponding two pages with… See the full description on the dataset page: https://huggingface.co/datasets/Chesscorner/chess-images.mmathcot1m-images
MMathCoT-1M Images (Sharded)
Local mirror of image assets referenced by the
URSA-MATH/MMathCoT-1M dataset.
Shards: tar files in shards/
Manifest: manifest.csv and manifest.parquet
Each tar’s internal arcname equals the original image_url from the dataset.
synthetic_jawi_imagesocr-vqa-200k_imagesImage collections for OCR-VQA-200K.
Image size: 208,467.
eqben-images
Equivariant Similarity for Vision-Language Foundation Models
ICCV 2023
Tan Wang,
Kevin Lin,
Linjie Li,
Chung-Ching Lin,
Zhengyuan Yang,
Hanwang Zhang,
Zicheng Liu,
Lijuan Wang
Nanyang Technological University, Microsoft Corporation
About
This study explores the concept of equivariance in vision-language foundation models (VLMs), focusing specifically on the multimodal similarity function that is… See the full description on the dataset page: https://huggingface.co/datasets/ytaek-oh/eqben-images.lfhre-images-v2test-imagesSCORE_automatic_imagesscryfall-card-imagestun3d_s3dis_posed_imagesobject_images
Overview
This dataset contains 2D rendered images generated from 3D assets originally sourced from Objaverse and Objathor.The purpose of this dataset is to provide a large-scale collection of photo-realistic renderings for research on vision, multimodal learning, and text-to-3D understanding.
Following prior works such as Diffusion4D and Stable-Zero123,we release only the rendered 2D images (not the original 3D assets) to facilitate efficient experimentation while preserving the… See the full description on the dataset page: https://huggingface.co/datasets/LEGO-Eval/object_images.my-imagesILSVRC_images_10_classtar -xzf images_10_class.tar.gz
images_10_class/
├── 000_tench/
│ ├── 00000.jpg
│ ├── 00001.jpg
│ └── ... (1300 images)
├── 001_goldfish/
│ ├── 00000.jpg
│ └── ...
├── 002_great_white_shark/
│ └── ...
├── 003_tiger_shark/
│ └── ...
├── 004_hammerhead/
│ └── ...
├── 005_electric_ray/
│ └── ...
├── 006_stingray/
│ └── ...
├── 007_cock/
│ └── ...
├── 008_hen/
│ └── ...
└── 009_ostrich/
└── ...
CLASS_INFO = {
"000_tench": "A tench, a freshwater fish"… See the full description on the dataset page: https://huggingface.co/datasets/CCRss/ILSVRC_images_10_class.sdxl_images_easy_prompts-multi_artist-seed0Dalle3-1M-ImagesThis dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically share their best results online, this dataset reflects a diverse and high quality compilation of human preferences and high quality creative works. Captions for the images were generated using 4-bit CogVLM with custom… See the full description on the dataset page: https://huggingface.co/datasets/bitmind/Dalle3-1M-Images.MMEdit_imagesInterMT-Bench-Imagesdriver-images-processed-v1
