datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
NL-DIR
Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark (CVPR 2025)
Dataset Introduction
NL-DIR consists of 41,795 document images with a diverse range of document types, and each image corresponds to 5 high-quality fine-grained semantic queries, which are generated and evaluated through large language models in conjunction with manual verification, resulting 200K+ queries in total. Following an 8:1:1 ratio partition, the dataset is divided into… See the full description on the dataset page: https://huggingface.co/datasets/nianbing/NL-DIR.NL-DIR-sampleThe NL-DIR dataset comprises 41K authentic document images, in which each image is paired with five high-quality fine-grained semantic queries, generated and evaluated through large language models in conjunction with manual verification.
The complete dataset, along with detailed descriptions, specific formats, usage instructions, and construction methods, will be released soon.
NL_datasetleafjetworx_leafjet
