datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MixBench
MixBench: A Benchmark for Mixed Modality Retrieval
MixBench is a benchmark for evaluating retrieval across text, images, and multimodal documents. It is designed to test how well retrieval models handle queries and documents that span different modalities, such as pure text, pure images, and combined image+text inputs.
MixBench includes four subsets, each curated from a different data source:
MSCOCO
Google_WIT
VisualNews
OVEN
Each subset contains:
queries.jsonl: each entry… See the full description on the dataset page: https://huggingface.co/datasets/mixed-modality-search/MixBench.google_search_result
Google Search Results With Source Queries and Images
This dataset contains 396 newdomain samples from the local Continual-LLaVA-NeXT
workspace, enriched with entity-based Google/Serper search results.
Each row includes the original multimodal sample context:
dataset: source dataset name.
source_index: index in the original training JSON.
id: sample id used locally.
image: relative path to the copied image file in this dataset repo.
original_image: original image field from the… See the full description on the dataset page: https://huggingface.co/datasets/leo20000306/google_search_result.pinduoduo-search-filter-order-bulk1000searchengine-kie
搜尋引擎截圖VQA資料集
使用各家搜尋引擎自動化關鍵字搜尋後截圖的VQA資料集
格式如下
Q: 這張搜尋結果截圖的查詢內容與第一筆結果是什麼?<image>請回傳JSON格式。
A: {"搜尋內容": "得到聯系完成比較", "搜尋結果": "用孩子的語言,教他們學會分辨「的」和「得」 - 翻轉教育"}
為避免噪聲,標記僅涵蓋搜尋結果與第一筆結果;關鍵字為隨機字串生成,Json檔案無遵照常見標記結構,請根據檔案路徑進行匹配
由於Huggingface的API是只下載格式trace到的檔案,使用上請不要透過datasets或huggingface-cli,直接git clone https://huggingface.co/datasets/CGTec3/searchengine-kie來得到標記與圖片並進行處理
不保證資料必定正確無誤,請謹慎使用
visual-search-catalog
