document-retrieval
Qwen3-0.6B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-Generalization-BaselineBGE-M3-0.56B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-Generalization-BaselineQwen3-0.6B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-With-ConfigBGE-M3-0.56B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-With-ConfigBGE-M3-0.56B-Fine-Tuned-Telecom-Technical-Documents-Retrieval-Embedding-With-Config-Runpodpaperstack_document_data_retrievaldual-document-retrievalStatutory_Article_Retrieval_From_Legal_Document
document-visual-retrieval-test
Model Card: Document Visual Retrieval Test (internal)
Dataset Overview
This dataset is designed to evaluate the performance of visual retrievers by testing their ability to match a query to a relevant image. Each of the three examples in this dataset contains a text query and an associated image, which is a scanned page from the foundational "Attention is All You Need" paper. The purpose of this dataset is to facilitate the evaluation of visual retrievers, where the… See the full description on the dataset page: https://huggingface.co/datasets/hf-internal-testing/document-visual-retrieval-test.persian-web-document-retrieval
Dataset Summary
Persian Web Document Retrieval is a Persian (Farsi) dataset designed for the Retrieval task. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset consists of real-world queries collected from the Zarrebin search engine and web documents labeled by humans for relevance. It is curated to evaluate model performance in web search scenarios.
Language(s): Persian (Farsi)
Task(s): Retrieval (Web Search)
Source: Collected from Zarrebin… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-web-document-retrieval.multimodal-document-retrieval-20260911-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260911-dataset.DRCD-for-Document-Retrieval-Task DRCD for Document Retrieval (Simplified Chinese)
This dataset is a reformatted version of the Delta Reading Comprehension Dataset (DRCD), converted to Simplified Chinese and adapted for document-level retrieval tasks.
Summary
The dataset transforms the original DRCD QA data into a document retrieval setting, where queries are used to retrieve entire Wikipedia articles rather than individual passages. Each document is the full text of a Wikipedia entry.
The format is compatible with… See the full description on the dataset page: https://huggingface.co/datasets/ihainan/DRCD-for-Document-Retrieval-Task.multimodal-document-retrieval-20260901-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260901-dataset.multimodal-document-retrieval-20260822-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260822-dataset.
