datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR.multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN
PreFLMR M2KR Dataset Card
Dataset details
Dataset type:
M2KR is a benchmark dataset for multimodal knowledge retrieval. It contains a collection of tasks and datasets for training and evaluating multimodal knowledge retrieval models.
We pre-process the datasets into a uniform format and write several task-specific prompting instructions for each dataset. The details of the instruction can be found in the paper. The M2KR benchmark contains three types of tasks:… See the full description on the dataset page: https://huggingface.co/datasets/BByrneLab/multi_task_multi_modal_knowledge_retrieval_benchmark_M2KR_CN.pointer-retrievalgithub-readme-retrieval-multilingual_beirThis is a copy of https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/github-readme-retrieval-multilingual_beir.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13airbnb-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/airbnb-synthetic-retrieval_beir.twice_kr_news_retrieval
FinNews-Retrieval-ko
Constructed a Retrieval dataset based on Korean financial news articles.
tweet-stock-synthetic-retrieval_beirThis is a copy of https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/tweet-stock-synthetic-retrieval_beir.ms2-peptide-replicate-retrieval
MS2 Peptide-Replicate-Retrieval Benchmark
A benchmark for evaluating spectrum-embedding models for tandem mass
spectrometry (MS2). It is a set of real experimental MS2 spectra, each labelled
with the peptide it was identified as (a peptide-spectrum match, PSM), pooled so
that every peptide is represented by many replicate acquisitions. The task:
does an embedding map replicate spectra of the same peptide close together?
15,649 spectra
1,000 unique peptide/charge labels (≈ 15… See the full description on the dataset page: https://huggingface.co/datasets/chrisagrams/ms2-peptide-replicate-retrieval.twice_tat_qa_retrieval
TATQA-Retrieval-ko
Constructed a TAT (Textual and Tabular) QA dataset based on information collected from Korean financial reports.
twice_kr_market_report_retrieval
FinMarketReport-Retrieval-ko
Constructed a Retrieval dataset related to the stock market, based on Korean Financial Reports.
esg_cid_retrieval
Enhancing Retrieval for ESGLLM via ESG-CID -- A Disclosure Content Index Finetuning Dataset for Mapping GRI and ESRS
Usage
from datasets import load_dataset
# document chunks: train/dev/test_gri/test_esrs
documents = load_dataset("esgllm/esg_cid_retrieval", "documents")
# queries (disclosure text): train/dev/test_gri/test_esrs
queries = load_dataset("esgllm/esg_cid_retrieval", "queries")
# training triplets: train/dev
triplets = load_dataset("esgllm/esg_cid_retrieval"… See the full description on the dataset page: https://huggingface.co/datasets/airefinery/esg_cid_retrieval.opengloss-v2.3-retrieval-pairs
OpenGloss v2.3 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an example paired with its own sense's gloss (positive), and optional sampled cross-headword same-domain negatives. Every pair carries both spans, both reading levels, and live_senses, so a consumer can filter or reweight by… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.3-retrieval-pairs.newbalance_shoe_insole_retrieval_and_packing_0611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 298,
"total_frames": 4967405,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:298"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0611.shoe_insole_retrieval_and_packing0515This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 101,
"total_frames": 206299,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:101"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/shoe_insole_retrieval_and_packing0515.newbalance_shoe_insole_retrieval_and_packing_0604This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 108,
"total_frames": 1635459,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:108"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0604.opengloss-v2.2-retrieval-pairs
Superseded by OpenGloss v2.3 (2026-09-09): tier 6 adds ~12,000 named entities (people, places, organizations, works, events) with entity_type, Wikidata ids and alias_of links, and every proper noun in the release is now typed. v2.2 stays published for reproducibility.
OpenGloss v2.2 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.2-retrieval-pairs.opengloss-v2.1-retrieval-pairs
Superseded by OpenGloss v2.2 (2026-09-08): 148,292 live lexemes and 288,304 senses — tier 5 closes the WordNet gap (38,100 entries imported from Princeton WordNet 3.0 and enriched), inflected-form headwords are folded onto their lemmas, and every inherited field carries a migrate provenance record. v2.1 stays published for reproducibility.
OpenGloss v2.1 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.1-retrieval-pairs.enwikivoyage-retrieval-202605
English Wikivoyage Retrieval 2026-05
I like travel datasets because they are about real places, real constraints,
and the small practical questions people ask before they go somewhere. I am
sharing this English Wikivoyage retrieval corpus in that spirit: as honest
work from a researcher-builder who wants to explore the world, make the
pipeline inspectable, and let other people reuse or challenge the choices.
This is not presented as a finished travel product or a private… See the full description on the dataset page: https://huggingface.co/datasets/vincentss/enwikivoyage-retrieval-202605.open-hermes-2.5-sft-mixture-llama3-inference-retrieval-tokenscc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15opengloss-v2.0-retrieval-pairs
Superseded by OpenGloss v2.1 (2026-09-07): 109,633 lexemes and 250,003 live senses — twice this release's coverage — plus a new opengloss-v2.1-inflections form→lemma lookup. v2.0 stays published for reproducibility.
OpenGloss v2.0 — Retrieval Pairs
Binary-labelled text pairs mined straight off the release with no model call: two example sentences of the same sense (positive), one example from each of two senses of the same headword (the hard word-in-context negative), an… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/opengloss-v2.0-retrieval-pairs.mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13newbalance_shoe_insole_retrieval_and_packing_0529This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 65,
"total_frames": 332561,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:65"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/newbalance_shoe_insole_retrieval_and_packing_0529.tda-gnn-scientific-retrievalkarl-with-retrieval_bm25_v2
Dataset Card for "karl-with-retrieval_bm25_v2"
More Information needed
huili_shoe_insole_retrieval_and_packing_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 68,
"total_frames": 949788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:68"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Vertax/huili_shoe_insole_retrieval_and_packing_0602.mscoco_train_2014_openai_clip-vit-base-patch32_image_caption_retrieval_pairs_2022-09-01huili_shoe_insole_retrieval_and_packing_0602This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_flexiv_rizon4_rt",
"total_episodes": 68,
"total_frames": 949788,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:68"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Xense/huili_shoe_insole_retrieval_and_packing_0602.mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15
