datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vidore_v3_pharmaceuticalsViDoRe V3 : Pharmaceuticals
This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_pharmaceuticals.vidore_v3_pharmaceuticalsViDoRe V3 : Pharmaceuticals
This dataset, Pharmaceutical, is a corpus of slides from the FDA, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with human-verified relevant pages… See the full description on the dataset page: https://huggingface.co/datasets/Ethicalpirate91/vidore_v3_pharmaceuticals.vidore3_pharmaceuticals_neomme_260m_li
vidore3_pharmaceuticals_neomme_260m_li
Multi-vector (late-interaction) embeddings of ViDoRe pharmaceuticals (vidore/pharmaceuticals), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset vidore/vidore_v3_pharmaceuticals at revision 3abd4aa8a9445fb5538a78a19ba50bd57bd22b5c, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_pharmaceuticals_neomme_260m_li.
