datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
azerbaijani_retriever_corpus
A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training
Dataset Description
This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents.
The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.ko-law-retriever-artifacts-20260622
Korean Legal Retriever Artifacts 2026-06-22
This public dataset repository stores large artifacts for ko-law-retriever
that are too large for GitHub's 100 MB file limit.
Contents
ko_legal_source_rag/: retriever SFT JSONL mixes and metadata.
ko_legal_retriever/ko_legal_fts_full.sqlite: local SQLite FTS retrieval index.
evals/: source-rag evaluation outputs.
reports/: analysis JSON/Markdown reports and evidence-pack outputs.
Scope
These files are… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ko-law-retriever-artifacts-20260622.corpus.latest.vod-retriever-medical-v1.1retriever-princeton-nlp-CharXiv-clean
Description
princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.Qwen2.5-7B-Instruct-vllm-retriever-20251202_093826Tool-REX_train_retriever_50kIII-Retrieverllm-retriever-tasksThis dataset tasks for training in-context example retrievers.aihub_retriever_commonsense일반상식
musique_hotpotqa_graph_retrieverazerbaijani_books_retriever_corpus-reranked
Azerbaijani Books Retrieval Dataset (Reranked)
A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani.
Dataset Configs
The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.hint_retriever_datag-retriever-scene-graphsssf-synthetic-data-for-retriever
Dataset Card for ssf-synthetic-data-for-retriever
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever.retriever_training_dataaihub_retriever_doc_smr문서요약 텍스트
ssf-synthetic-data-for-retriever-openai
Dataset Card for ssf-synthetic-data-for-retriever-openai
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever-openai/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever-openai.aihub_retriever_news뉴스 기사 기계독해 데이터
code-retriever-query-passage-pairsaihub_retriever_admin행정 문서 대상 기계독해 데이터
lavague-retriever-bigretriever-vidore-vdsid_french-clean
Description
vidore/vdsid_french dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@misc{faysse2024colpaliefficientdocumentretrieval,
title={ColPali: Efficient Document Retrieval with Vision Language Models},
author={Manuel Faysse and Hugues… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-vidore-vdsid_french-clean.retriever-manu-tabfquad_retrieving-clean
Description
manu/tabfquad_retrieving dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
https://huggingface.co/datasets/manu/tabfquad_retrieving
aihub_retriever_tech기술과학 문서 기계독해 데이터
combisearch-retrieverscode_retriever_benchmarkmusique_hotpotqa_graph_text_retrieverAgent-Skills-Retrieverdemo-retrieveraihub_retriever_books_smr
