CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01llamaindex /ParseBench ParseBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics: Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on. Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ParseBench.document100K<n<1M129 likes23k downloads5mo agoHugging Face02llamaindex /ExtractBench ExtractBench Quick links: [🌐 Website] [📜 Paper] [💻 Code] Given a document and a schema, a system returns structured data with evidence. The input is a full document, born-digital or scanned, and a schema written by the user. The output is a schema-valid JSON object, with the source page and a bounding box for each value as evidence. It must return correct, exhaustive values (including repeated records), correctly use null for absent information, and ground each extracted… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/ExtractBench.documentn<1K32 likes20k downloads1mo agoHugging Face03llamaindex /vdr-multilingual-train Multilingual Visual Document Retrieval Dataset This dataset consists of 500k multilingual query image samples, collected and generated from scratch using public internet pdfs. The queries are synthetic and generated using VLMs (gemini-1.5-pro and Qwen2-VL-72B). It was used to train the vdr-2b-multi-v1 retrieval multimodal, multilingual embedding model. How it was created This is the entire data pipeline used to create the Italian subset of this dataset. Each step… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-train.image100K<n<1M31 likes914 downloads2y agoHugging Face04llamaindex /liteparse_bench_smalldocumentn<1K3 likes687 downloads4d agoHugging Face05llamaindex /liteparse_cicd_dataContains files and results used to validate pull-requests in the liteparse repo document1K<n<10K0 likes623 downloads1mo agoHugging Face06tomaarsen /llamaindex-vdr-en-train-preprocessed llamaindex-vdr-en-train-preprocessed This dataset is a preprocessed English subset of llamaindex/vdr-multilingual-train, prepared for training multimodal Sentence Transformer embedding models on document screenshot retrieval. Changes from the original dataset The original llamaindex/vdr-multilingual-train dataset stores hard negatives as a list of ID strings that reference other rows. This dataset makes two key changes: English only: Only the English subset (53,512… See the full description on the dataset page: https://huggingface.co/datasets/tomaarsen/llamaindex-vdr-en-train-preprocessed.image10K<n<100K3 likes503 downloads6mo agoHugging Face07lightonai /llamaindex-vdr-images LlamaIndex VDR Images This dataset contains the document-page images used by lightonai/llamaindex-vdr-fine-tuning. The pair is a reformatted derivative of llamaindex/vdr-multilingual-train for multilingual retrieval fine-tuning. Dataset structure The train split contains: Column Type Description image_filename string Stable key used by the companion fine-tuning dataset. image image Document-page image. Load the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/llamaindex-vdr-images.imagevisual-document-retrieval100K<n<1M0 likes242 downloads2mo agoHugging Face08AlignmentLab-AI /llama-indextext1K<n<10K3 likes159 downloads3y agoHugging Face09lightonai /llamaindex-vdr-fine-tuning LlamaIndex VDR Fine-Tuning This dataset reformats llamaindex/vdr-multilingual-train for retrieval fine-tuning with PyLate. It contains queries in German, English, Spanish, French, and Italian, together with document metadata and the hard negatives provided by the source dataset. Images are stored separately in lightonai/llamaindex-vdr-images. Hard negatives The source dataset mined hard negatives with voyage-3 using a fixed similarity threshold of 0.75.… See the full description on the dataset page: https://huggingface.co/datasets/lightonai/llamaindex-vdr-fine-tuning.textvisual-document-retrieval1M<n<10M0 likes157 downloads2mo agoHugging Face10towardsai-tutors /llama-index-docstextn<1K2 likes126 downloads2y agoHugging Face11llamaindex /vdr-multilingual-test Multilingual Visual Document Retrieval Benchmarks This dataset consists of 15 different benchmarks used to initially evaluate the vdr-2b-multi-v1 multimodal retrieval embedding model. These benchmarks allow the testing of multilingual, multimodal retrieval capabilities on text-only, visual-only and mixed page screenshots. Each language subset contains queries and images in that language and is divided into three different categories by the "pagetype" column. Each category contains… See the full description on the dataset page: https://huggingface.co/datasets/llamaindex/vdr-multilingual-test.image10K<n<100K3 likes122 downloads2y agoHugging Face12luna-code /llamaindextext1K<n<10K0 likes57 downloads3y agoHugging Face13shanghong /llama_index_integration_datatext10M<n<100M0 likes47 downloads1y agoHugging Face14antonioloison /filtered-llamaindex-with-translationsimage100K<n<1M0 likes32 downloads1y agoHugging Face15adalib /llama_index-cond-gentext1K<n<10K0 likes25 downloads3y agoHugging Face16sreddy109 /llama_index_firewall_trivia_qa_200ktext100K<n<1M0 likes25 downloads2y agoHugging Face17mole-code /llama_indextext1K<n<10K0 likes24 downloads2y agoHugging Face18sreddy109 /llama_index_firewall_trivia_qa_200k_shuffledtext100K<n<1M0 likes24 downloads2y agoHugging Face19mhammadkhan /llamaindex_stack Dataset Card for "llamaindex_stack" More Information needed text1K<n<10K0 likes22 downloads3y agoHugging Face20mole-code /llama_index-datatext1K<n<10K0 likes19 downloads2y agoHugging Face21genaiarchitect /llamaindextextn<1K0 likes19 downloads1y agoHugging Face22adalib /llama_index-cond-gen-10text1K<n<10K0 likes16 downloads3y agoHugging Face23adalib /llama_indextext1K<n<10K0 likes16 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.