datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
retriever-princeton-nlp-CharXiv-clean
Description
princeton-nlp/CharXiv dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@article{wang2024charxiv,
title={CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs},
author={Wang, Zirui and Xia, Mengzhou and He, Luxi and… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-princeton-nlp-CharXiv-clean.imnet1k_golden_retrieverretriever-vidore-vdsid_french-clean
Description
vidore/vdsid_french dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@misc{faysse2024colpaliefficientdocumentretrieval,
title={ColPali: Efficient Document Retrieval with Vision Language Models},
author={Manuel Faysse and Hugues… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-vidore-vdsid_french-clean.retriever-manu-tabfquad_retrieving-clean
Description
manu/tabfquad_retrieving dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
https://huggingface.co/datasets/manu/tabfquad_retrieving
imnet1k_flat-coated_retrieverimnet1k_curly-coated_retrieverimnet1k_Labrador_retrieverretriever-vidore-tabfquad_test_subsampled-clean
Description
vidore/tabfquad_test_subsampled dataset that we processed.Although useless, we have created an empty answer column to facilitate the concatenation of this dataset with VQA datasets where only the quesion and image columns would be used to train a Colpali-type model or one of its derivatives.
Citation
@misc{faysse2024colpaliefficientdocumentretrieval,
title={ColPali: Efficient Document Retrieval with Vision Language Models},
author={Manuel Faysse… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/retriever-vidore-tabfquad_test_subsampled-clean.imnet1k_Chesapeake_Bay_retrieverretriever_eval_dmv_datasetsretriever_eval_dmv_datasets_extended
