datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mscoco_2014_5k_test_image_text_retrieval
MSCOCO (5K test set)
Original paper: Microsoft COCO: Common Objects in Context
Homepage: https://cocodataset.org/#home
5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@inproceedings{lin2014microsoft,
title={Microsoft coco: Common objects in context},
author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/mscoco_2014_5k_test_image_text_retrieval.flickr_1k_test_image_text_retrieval
Flickr30k (1K test set)
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage: https://shannon.cs.illinois.edu/DenotationGraph/
1K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@article{young2014image,
title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr_1k_test_image_text_retrieval.japanese-text-image-retrieval-trainshunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。
query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1つを選んだものをデータセットに含めました。質問を生成する際は以下のプロンプトを使用しました。
あなたは、質問から画像をretrieveするためのモデルをトレーニングするための(質問, 画像)ペアのデータセットを作成するプロジェクトのメンバーである。
プロジェクトは以下のように進める。
step1. ドキュメントPDFを1ページ1枚の画像ファイルに変換する
step2.… See the full description on the dataset page: https://huggingface.co/datasets/oshizo/japanese-text-image-retrieval-train.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13AfriMCQA-speech-image-retrieval
Afri-MCQA speech-image retrieval (MTEB)
Afri-MCQA reshaped for retrieval: find the photograph a spoken question is asking
about. Questions are spoken by native speakers in 16 African languages and are
grounded in culturally relevant images.
Source: Atnafu/Afri-MCQA at revision 8b8c53d, cc-by-nc-4.0, official
test split. Images are stored per language because one photograph can carry
questions in several languages.
Built by scripts/data/afri_mcqa_retrieval/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-speech-image-retrieval.vaani-audio-image-retrieval
Vaani audio–image retrieval (MTEB)
Multilingual audio↔image retrieval over 62 Indian languages, derived from
Project Vaani (IISc Bangalore /
ARTPARK).
Vaani records image-prompted speech: a speaker is shown a photograph and describes it
aloud in their own language. Each recording is therefore grounded in a specific image,
which is what makes audio↔image retrieval well defined without any extra annotation.
Prepared for MTEB as
VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.ilias_image_retrieval_v1cusa-image-text-retrieval
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
This is the source code of our AAAI 2024 paper "Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval"
[ Paper | Appendix ]
Quick Links
Cross-Modal and Uni-Modal Soft-Label Alignment for Image-Text Retrieval
Quick Links
Overview
Usage
Getting Started
Environment Installation
Data Preprocessing
Training & Evaluation
Q&A
Citation
Overview
We propose a… See the full description on the dataset page: https://huggingface.co/datasets/EGOISTyrh/cusa-image-text-retrieval.mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01flickr30k_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-14cc12m_openai-clip-vit-patch32_image_retrieval_top15_start1000000_end3500000mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15cc12m_openai-clip-vit-patch32_image_retrieval_top15_start1000000_end3500000_SHORT500Kmscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairscc12m_openai-clip-vit-patch32_image_retrieval_top4_start1000000_end3000000_DEBUGproduct10k_image_retrievalinsect_image_retrieval
Introduction
This dataset provides CLIP embeddings for images of insects, enabling similarity search and content-based retrieval.
Source
All images were sourced from Francesco/insects-mytwu.
Code
If you wish to create a similar dataset then feel free to use the code provided here:
https://github.com/rishik18/CLIP_embedding_based_retrieval
Sample Usage
from datasets import load_dataset
ds_new = load_dataset("hkanade/insect_image_retrieval")… See the full description on the dataset page: https://huggingface.co/datasets/hkanade/insect_image_retrieval.image-captioning-retrievalA dataset of curated images used for captioning using LLMs and retrieval based on caption(text based not embedding based) using a go server.
cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15_SHORTstanford-cars-image-retrieval-v1mscoco_2014_5k_test_image_text_retrieval
MSCOCO (5K test set)
Original paper: Microsoft COCO: Common Objects in Context
Homepage: https://cocodataset.org/#home
5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@inproceedings{lin2014microsoft,
title={Microsoft coco: Common objects in context},
author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/wdios/mscoco_2014_5k_test_image_text_retrieval.japanese-text-image-retrieval下記のデータセットのtest splitをretrievalの評価データセットとして編集したものです。https://huggingface.co/datasets/shunk031/JDocQA
corpusにはtest splitのquestion_page_numberに含まれるページのみが含まれます。corpusの画像は縦横の長い辺が720pxを超えないように縮小されています。
japanese-text-image-retrieval-small2cub200-image-retrieval-v1multilingual-image-recipe-retrieval
Description
This is the first real-world multicultural cuisine dataset designed for the image–recipe retrieval task. It addresses limitations of existing datasets by providing multilingual recipes spanning five Southeast Asian cuisines (Indonesia, Malaysia, Thailand, Vietnam and India).
Motivation
Existing image–recipe retrieval datasets focus primarily on monolingual settings, whereas in real-world applications, recipes originate from diverse regions worldwide and are… See the full description on the dataset page: https://huggingface.co/datasets/Multimedia-SMU/multilingual-image-recipe-retrieval.image-retrievaljapanese-text-image-retrieval-smallwds_mscoco_2014_5k_test_image_text_retrieval_test
