image-retrieval
mscoco_2014_5k_test_image_text_retrieval
MSCOCO (5K test set)
Original paper: Microsoft COCO: Common Objects in Context
Homepage: https://cocodataset.org/#home
5K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@inproceedings{lin2014microsoft,
title={Microsoft coco: Common objects in context},
author={Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Doll{\'a}r, Piotr and Zitnick, C Lawrence}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/mscoco_2014_5k_test_image_text_retrieval.flickr_1k_test_image_text_retrieval
Flickr30k (1K test set)
Original paper: From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Homepage: https://shannon.cs.illinois.edu/DenotationGraph/
1K test set split from: http://cs.stanford.edu/people/karpathy/deepimagesent/caption_datasets.zip
Bibtex:
@article{young2014image,
title={From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions}… See the full description on the dataset page: https://huggingface.co/datasets/nlphuji/flickr_1k_test_image_text_retrieval.japanese-text-image-retrieval-trainshunk031/JDocQAのtrain splitに含まれるPDFデータを画像化し、NDLOCRでOCRしたテキストとペアにしたデータセットです。OCRは長い辺を1200pxにリサイズした画像に対して実施しました。OCR結果には、読み取りに失敗した際の文字列「〓」が含まれます。本データセットに含めている画像は、長い辺を896px、700px、588pxのいずれかにリサイズしています。どのサイズとするかは主にページに含まれる文字数で決めました。
query列は、OCR結果の文字列に対しQwen/Qwen2.5-14B-Instructで生成したものです。3つの質問を生成させ、ランダムに1つを選んだものをデータセットに含めました。質問を生成する際は以下のプロンプトを使用しました。
あなたは、質問から画像をretrieveするためのモデルをトレーニングするための(質問, 画像)ペアのデータセットを作成するプロジェクトのメンバーである。
プロジェクトは以下のように進める。
step1. ドキュメントPDFを1ページ1枚の画像ファイルに変換する
step2.… See the full description on the dataset page: https://huggingface.co/datasets/oshizo/japanese-text-image-retrieval-train.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13AfriMCQA-speech-image-retrieval
Afri-MCQA speech-image retrieval (MTEB)
Afri-MCQA reshaped for retrieval: find the photograph a spoken question is asking
about. Questions are spoken by native speakers in 16 African languages and are
grounded in culturally relevant images.
Source: Atnafu/Afri-MCQA at revision 8b8c53d, cc-by-nc-4.0, official
test split. Images are stored per language because one photograph can carry
questions in several languages.
Built by scripts/data/afri_mcqa_retrieval/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/AfriMCQA-speech-image-retrieval.vaani-audio-image-retrieval
Vaani audio–image retrieval (MTEB)
Multilingual audio↔image retrieval over 62 Indian languages, derived from
Project Vaani (IISc Bangalore /
ARTPARK).
Vaani records image-prompted speech: a speaker is shown a photograph and describes it
aloud in their own language. Each recording is therefore grounded in a specific image,
which is what makes audio↔image retrieval well defined without any extra annotation.
Prepared for MTEB as
VaaniA2IRetrieval and VaaniI2ARetrieval.… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/vaani-audio-image-retrieval.
