CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LocalDoc /azerbaijani_retriever_corpus A Large-Scale Azerbaijani Corpus for Contrastive Retriever Training Dataset Description This dataset is a large-scale, high-quality resource designed for training Azerbaijani text embedding models for information retrieval tasks. It contains 671,528 training instances, each consisting of a query, a relevant positive document, and 10 hard-negative documents. The primary goal of this dataset is to facilitate the training of dense retriever models using contrastive learning.… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_retriever_corpus.tabularsentence-similarity100K<n<1M0 likes381 downloads1y agoHugging Face02gyung /ko-law-retriever-artifacts-20260622 Korean Legal Retriever Artifacts 2026-06-22 This public dataset repository stores large artifacts for ko-law-retriever that are too large for GitHub's 100 MB file limit. Contents ko_legal_source_rag/: retriever SFT JSONL mixes and metadata. ko_legal_retriever/ko_legal_fts_full.sqlite: local SQLite FTS retrieval index. evals/: source-rag evaluation outputs. reports/: analysis JSON/Markdown reports and evidence-pack outputs. Scope These files are… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ko-law-retriever-artifacts-20260622.textn<1K0 likes377 downloads3mo agoHugging Face03findzebra /corpus.latest.vod-retriever-medical-v1.1text100K<n<1M1 likes210 downloads3y agoHugging Face04CATIE-AQ /retriever-princeton-nlp-CharXiv-clean Description princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives. Citation @article{wang2024charxiv, title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs}, author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.imageimage-feature-extraction1K<n<10K0 likes143 downloads1y agoHugging Face05lihaoxin2020 /Qwen2.5-7B-Instruct-vllm-retriever-20251202_093826text0 likes129 downloads8mo agoHugging Face06Lux1997 /Tool-REX_train_retriever_50ktext10K<n<100K0 likes115 downloads8mo agoHugging Face07qian /III-Retrievertext100K<n<1M1 likes108 downloads3y agoHugging Face08intfloat /llm-retriever-tasksThis dataset tasks for training in-context example retrievers.text100K<n<1M8 likes88 downloads3y agoHugging Face09wisenut-nlp-team /aihub_retriever_commonsense일반상식 text10K<n<100K0 likes49 downloads2y agoHugging Face10gzone0111 /musique_hotpotqa_graph_retrievertext10K<n<100K0 likes44 downloads1y agoHugging Face11LocalDoc /azerbaijani_books_retriever_corpus-reranked Azerbaijani Books Retrieval Dataset (Reranked) A large-scale retrieval dataset built from LocalDoc/books_dataset — a collection of 2,804 Azerbaijani-language books with 7.8M sentences spanning politics, history, literature, science, and more. Designed for training and evaluating information retrieval, semantic search, and RAG pipelines in Azerbaijani. Dataset Configs The dataset consists of three configs that can be joined via passage_id and query_id:… See the full description on the dataset page: https://huggingface.co/datasets/LocalDoc/azerbaijani_books_retriever_corpus-reranked.tabularsentence-similarity1M<n<10M0 likes43 downloads6mo agoHugging Face12yoonholee /hint_retriever_datatabular10K<n<100K0 likes40 downloads1y agoHugging Face13xxhe /g-retriever-scene-graphstabular100K<n<1M0 likes39 downloads2y agoHugging Face14dnth /ssf-synthetic-data-for-retriever Dataset Card for ssf-synthetic-data-for-retriever This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever.text1K<n<10K0 likes36 downloads1y agoHugging Face15liuwenhan /retriever_training_datatext100K<n<1M3 likes36 downloads8mo agoHugging Face16wisenut-nlp-team /aihub_retriever_doc_smr문서요약 텍스트 text100K<n<1M0 likes34 downloads2y agoHugging Face17dnth /ssf-synthetic-data-for-retriever-openai Dataset Card for ssf-synthetic-data-for-retriever-openai This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever-openai/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/dnth/ssf-synthetic-data-for-retriever-openai.text10K<n<100K0 likes29 downloads1y agoHugging Face18wisenut-nlp-team /aihub_retriever_news뉴스 기사 기계독해 데이터 text100K<n<1M0 likes28 downloads2y agoHugging Face19arcee-ai /code-retriever-query-passage-pairstext100K<n<1M2 likes27 downloads3y agoHugging Face20wisenut-nlp-team /aihub_retriever_admin행정 문서 대상 기계독해 데이터 text100K<n<1M0 likes25 downloads2y agoHugging Face21BigAction /lavague-retriever-bigtextn<1K0 likes22 downloads2y agoHugging Face22CATIE-AQ /retriever-vidore-vdsid_french-clean Description vidore/vdsid_french dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives. Citation @misc{faysse2024colpaliefficientdocumentretrieval, title={ColPali: Efficient Document Retrieval with Vision Language Models}, author={Manuel Faysse and Hugues… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-vidore-vdsid_french-clean.imageimage-feature-extraction1K<n<10K0 likes22 downloads1y agoHugging Face23CATIE-AQ /retriever-manu-tabfquad_retrieving-clean Description manu/tabfquad_retrieving dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives. Citation https://huggingface.co/datasets/manu/tabfquad_retrieving imageimage-feature-extraction1K<n<10K0 likes20 downloads1y agoHugging Face24wisenut-nlp-team /aihub_retriever_tech기술과학 문서 기계독해 데이터 text10K<n<100K0 likes17 downloads2y agoHugging Face25holi-lab /combisearch-retrieverstext100K<n<1M0 likes17 downloads4mo agoHugging Face26nankisu0301 /code_retriever_benchmarktextn<1K0 likes16 downloads6mo agoHugging Face27gzone0111 /musique_hotpotqa_graph_text_retrievertext10K<n<100K0 likes15 downloads1y agoHugging Face28Athekunal /Agent-Skills-Retrievertext100K<n<1M0 likes15 downloads5mo agoHugging Face29CarlosMalaga /demo-retrievertextn<1K0 likes13 downloads2y agoHugging Face30wisenut-nlp-team /aihub_retriever_books_smrtext100K<n<1M0 likes13 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.