CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AdithyaSK /RAG_Evalimage1K<n<10K0 likes19k downloads2y agoHugging Face02G4KMU /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.documenttable-question-answering10K<n<100K17 likes6k downloads6mo agoHugging Face03siyrus /BToks-vidore_rag_Infographic-VQA BToks ViDoRe Infographic-VQA This dataset repository contains Lance-format converted data used by the open-source reproduction code for Bottleneck Tokens for Unified Multimodal Retrieval (arXiv:2604.11095). Source Converted from vidore/colpali_train_set. Subset/view: infographic_vqa. This repository does not change upstream ownership, licensing, citation requirements, or usage restrictions. Format The data is stored as Lance tables for the… See the full description on the dataset page: https://huggingface.co/datasets/siyrus/BToks-vidore_rag_Infographic-VQA.imageimage-to-text10K<n<100K0 likes967 downloads3mo agoHugging Face04rl-rag /hle_text_onlyimage1K<n<10K0 likes885 downloads7mo agoHugging Face05racineai /VDR_ibm-research_REAL-MM-RAG VDR_ibm-research_REAL-MM-RAG - Overview Dataset Summary VDR_ibm-research_REAL-MM-RAG is a multimodal dataset that combines text and image data, and support tasks such as DSE retrieval (RAG). Dataset Creation This dataset is a merge and shuffle of the following datasets in the VDR format: ibm-research/REAL-MM-RAG_TechSlides ibm-research/REAL-MM-RAG_TechReport ibm-research/REAL-MM-RAG_FinTabTrainSet ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_ibm-research_REAL-MM-RAG.imagevisual-document-retrieval100K<n<1M8 likes877 downloads10mo agoHugging Face06brics-edtech /nornikel-metallurgy-rag-index Nornikel Metallurgy RAG Knowledge Base (для векторного индекса) Дедуплицированный корпус текстовых фрагментов (и связанных изображений) из технической базы знаний по металлургии/горному делу/обогащению — источник для построения векторного индекса (Annoy) в пайплайне RAG. Индекс НЕ включён в этот репозиторий — эмбеддинги и Annoy-индекс строятся во время выполнения ноутбука Google Colab (на GPU, это быстрее, чем на CPU), используя corpus.jsonl как исходные данные. Модель… See the full description on the dataset page: https://huggingface.co/datasets/brics-edtech/nornikel-metallurgy-rag-index.image10K<n<100K0 likes860 downloads3mo agoHugging Face07BangumiBase /ragnacrimson Bangumi Image Base of Ragna Crimson This is the image base of bangumi Ragna Crimson, we detected 98 characters, 6899 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/ragnacrimson.image1K<n<10K0 likes749 downloads2y agoHugging Face08botay /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/botay/t2-ragbench.documenttable-question-answering10K<n<100K0 likes495 downloads5mo agoHugging Face09ValerianFourel /ragdag-results RAGDAG results Artefacts from RAGDAG - treating a multi-stage retrieval pipeline as a structural causal model and computing path-specific effects exactly by freezing stages, rather than estimating them. Code: https://github.com/ValerianFourel/RAGDAG Layout One directory per collection, named after its ir_datasets id: <dataset-tag>/ REPORT.md human-readable report incl. the PASS/FAIL verdict MANIFEST.json provenance: git SHA, code… See the full description on the dataset page: https://huggingface.co/datasets/ValerianFourel/ragdag-results.image1M<n<10M0 likes486 downloads2mo agoHugging Face10raghavendrad60 /vqa_plant-disease-classification-merged-datasetimage10K<n<100K1 likes460 downloads2y agoHugging Face11ragingbullG9 /fashion-recommender-dataimage1K<n<10K0 likes435 downloads2mo agoHugging Face12ibm-research /REAL-MM-RAG_FinReport REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport.image1K<n<10K8 likes272 downloads2y agoHugging Face13ibm-research /REAL-MM-RAG_TechReport REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport.image1K<n<10K3 likes231 downloads2y agoHugging Face14ibm-research /REAL-MM-RAG_TechSlides REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides.image1K<n<10K2 likes231 downloads2y agoHugging Face15bharat-raghunathan /indian-foods-dataset Dataset Card for Indian Foods Dataset Dataset Summary This is a multi-category(multi-class classification) related Indian food dataset showcasing The-massive-Indian-Food-Dataset. This card has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages English Dataset Structure { "image": "Image(decode=True, id=None)", "target": "ClassLabel(names=['biryani', 'cholebhature'… See the full description on the dataset page: https://huggingface.co/datasets/bharat-raghunathan/indian-foods-dataset.imageimage-classification1K<n<10K7 likes221 downloads3y agoHugging Face16grasson /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/grasson/t2-ragbench.documenttable-question-answering10K<n<100K0 likes210 downloads5mo agoHugging Face17ibm-research /REAL-MM-RAG_TechSlides_BEIR BEIR Version of REAL-MM-RAG_TechSlides Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_TechSlides It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_TechSlides Content: 62 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechSlides_BEIR.image1K<n<10K1 likes201 downloads1y agoHugging Face18ibm-research /REAL-MM-RAG_FinReport_BEIR BEIR Version of REAL-MM-RAG_FinReport Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_FinReport It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_FinReport Content: 19 financial… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinReport_BEIR.image1K<n<10K2 likes200 downloads1y agoHugging Face19ibm-research /REAL-MM-RAG_FinSlides_BEIR BEIR Version of REAL-MM-RAG_FinSlides Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_FinSlides It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_FinSlides Content: 65 quarterly… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides_BEIR.image1K<n<10K1 likes199 downloads1y agoHugging Face20raghad-murad /clinical-skin-diseaseimagen<1K0 likes181 downloads6mo agoHugging Face21ibm-research /REAL-MM-RAG_TechReport_BEIR BEIR Version of REAL-MM-RAG_TechReport Summary This dataset is the BEIR-compatible version of the following Hugging Face dataset: ibm-research/REAL-MM-RAG_TechReport It has been reformatted into the BEIR structure for evaluation in retrieval settings.The original dataset is QA-style (each row is a query tied to a document image).Here, queries, qrels, docs, and corpus are separated into BEIR-standard splits. REAL-MM-RAG_TechReport Content: 17 technical… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_TechReport_BEIR.image1K<n<10K1 likes175 downloads1y agoHugging Face22Fujitsu /agentic-rag-redteam-benchgated WARNING: HARMFUL CONTENT - RESEARCH USE ONLY This dataset contains adversarial prompts, jailbreak attacks, toxic outputs, and other explicitly harmful content generated for AI safety research. Samples include prompt injections, social engineering payloads, misinformation, hate speech, instructions for illegal activities, phishing templates, and other dangerous material. All content is synthetic and produced by automated red-teaming pipelines for the sole purpose of evaluating and improving… See the full description on the dataset page: https://huggingface.co/datasets/Fujitsu/agentic-rag-redteam-bench.imagetext-retrieval10K<n<100K1 likes164 downloads7mo agoHugging Face23ibm-research /REAL-MM-RAG_FinTabTrainSet REAL-MM-RAG_FinTabTrainSet We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM) and filtered by a large language model (LLM). We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. For more information, see the project page:… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet.image10K<n<100K2 likes152 downloads1y agoHugging Face24ibm-research /REAL-MM-RAG_FinSlides REAL-MM-RAG-Bench: A Real-World Multi-Modal Retrieval Benchmark We introduced REAL-MM-RAG-Bench, a real-world multi-modal retrieval benchmark designed to evaluate retrieval models in reliable, challenging, and realistic settings. The benchmark was constructed using an automated pipeline, where queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM to ensure high-quality retrieval evaluation. To simulate real-world… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinSlides.image1K<n<10K2 likes144 downloads2y agoHugging Face25PatelDivyam23 /multimodal-rag-indeximagen<1K1 likes137 downloads10d agoHugging Face26ibm-research /REAL-MM-RAG_FinTabTrainSet_rephrased REAL-MM-RAG_FinTabTrainSet_rephrased We curated a table-focused finance dataset from FinTabNet (Zheng et al., 2021), extracting richly formatted tables from S&P 500 filings. We used an automated pipeline in which queries were generated by a vision-language model (VLM), filtered by a large language model (LLM), and rephrased by an LLM. We generated 48,000 natural-language (query, answer, page) triplets to improve retrieval models on table-intensive financial documents. This is the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-research/REAL-MM-RAG_FinTabTrainSet_rephrased.image10K<n<100K2 likes132 downloads1y agoHugging Face27tomsummerfield /t2-ragbench Dataset Card for T2-RAGBench Project Page | Paper | Code IMPORTANT NOTICE: We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history. Dataset Description Dataset Summary T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/tomsummerfield/t2-ragbench.documenttable-question-answering10K<n<100K0 likes123 downloads6mo agoHugging Face28thang3092004 /ebr-rag-final-data-ingestimagen<1K0 likes98 downloads3mo agoHugging Face29raghad-murad /clinical-skin-disease-imagesimage1K<n<10K0 likes96 downloads6mo agoHugging Face30blamm /retail_visual_rag_pipeline A Visual RAG Pipeline for Few-Shot Fine-Grained Product Classification Paper Accepted at The 12th Workshop on Fine-Grained Visual Categorization (FGVC12) at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2025. Overview The Retail Visual RAG Pipeline Dataset is a subset of the Retail-786k (Retail-786k) image dataset, supplemented with additional textual data per image. Data Image Data: The images are cropped from scanned… See the full description on the dataset page: https://huggingface.co/datasets/blamm/retail_visual_rag_pipeline.imagevisual-question-answering1K<n<10K0 likes76 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.