datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
persian-web-document-retrieval
Dataset Summary
Persian Web Document Retrieval is a Persian (Farsi) dataset designed for the Retrieval task. It is a component of the FaMTEB (Farsi Massive Text Embedding Benchmark). This dataset consists of real-world queries collected from the Zarrebin search engine and web documents labeled by humans for relevance. It is curated to evaluate model performance in web search scenarios.
Language(s): Persian (Farsi)
Task(s): Retrieval (Web Search)
Source: Collected from Zarrebin… See the full description on the dataset page: https://huggingface.co/datasets/MCINext/persian-web-document-retrieval.multimodal-document-retrieval-20260911-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260911-dataset.multimodal-document-retrieval-20260901-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260901-dataset.multimodal-document-retrieval-20260822-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260822-dataset.multimodal-document-retrieval-20260812-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260812-dataset.multimodal-document-retrieval-20260802-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260802-dataset.multimodal-document-retrieval-20260921-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260921-dataset.multimodal-document-retrieval-20260723-dataset
Multimodal Document Retrieval Baseline Synthetic Dataset
Summary
This dataset contains 14 training examples and 4
held-out examples for Business documents contain meaning in text, tables, layout, and imagery that text-only retrieval can miss.
Every record is synthetic and includes:
input: query, event, or feature description
label: expected class, route, relation, or evidence category
context: synthetic supporting context
source: fictional source identifier… See the full description on the dataset page: https://huggingface.co/datasets/RKB109/multimodal-document-retrieval-20260723-dataset.korean_document_retrieval_priv
