CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vidore /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K1 likes2k downloads1y agoHugging Face02vidore /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.9k downloads1y agoHugging Face03vidore /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1.8k downloads1y agoHugging Face04vidore /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.8k downloads1y agoHugging Face05mteb /docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.2k downloads8mo agoHugging Face06mteb /infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1.1k downloads8mo agoHugging Face07mteb /arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled. imagedocument-question-answering1K<n<10K0 likes1k downloads8mo agoHugging Face08mteb /tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled. imagedocument-question-answeringn<1K0 likes1k downloads8mo agoHugging Face09vidore /docvqa_test_subsampled Dataset Description This is the test set taken from the DocVQA dataset. It includes collected images from the UCSF Industry Documents Library. Questions and answers were manually annotated. Example of data (see viewer) Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/docvqa_test_subsampled.imagedocument-question-answeringn<1K6 likes897 downloads1y agoHugging Face10lucweber /mlqa__subsampledtext100K<n<1M0 likes742 downloads1y agoHugging Face11vidore /tabfquad_test_subsampled Dataset Description TabFQuAD (Table French Question Answering Dataset) is designed to evaluate TableQA models in realistic industry settings. Using a vision language model (GPT4V), we create additional queries to augment the existing human-annotated ones. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 280 pairs, leaving the rest for training and renaming the different columns. Load the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/tabfquad_test_subsampled.imagedocument-question-answeringn<1K0 likes725 downloads1y agoHugging Face12vidore /infovqa_test_subsampled Dataset Description This is the test set taken from the InfoVQA dataset. includes infographics collected from the Internet using the search query “infographics”. Questions and answers were manually annotated. Questions and answers were manually annotated. Example of data : (see viewer) Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns. Load the dataset from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/infovqa_test_subsampled.imagedocument-question-answeringn<1K3 likes686 downloads1y agoHugging Face13vidore /arxivqa_test_subsampled Dataset Description This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision. Data Curation To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose. Load the dataset from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/arxivqa_test_subsampled.imagedocument-question-answeringn<1K4 likes653 downloads1y agoHugging Face14lucweber /logiqa2__subsampledtabular1K<n<10K0 likes350 downloads1y agoHugging Face15LucasWeber /logiqa2__subsampledtabular1K<n<10K0 likes344 downloads1y agoHugging Face16benjamin /flanv2_subsampletext10M<n<100M0 likes314 downloads2y agoHugging Face17justram /atomic_subsampled_500kimage100K<n<1M1 likes267 downloads2y agoHugging Face18sanderland /CulturaX-subsample-100-bal2val-2text1M<n<10M0 likes242 downloads1y agoHugging Face19yali30 /findingdory-subsampled-96 FindingDory: A Benchmark to Evaluate Memory in Embodied Agents Karmesh Yadav*, Yusuf Ali*, Gunshi Gupta, Yarin Gal, Zsolt Kira Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks. In this repo, we release the FindingDory Subsampled Video Dataset. Each video contains 96 images… See the full description on the dataset page: https://huggingface.co/datasets/yali30/findingdory-subsampled-96.textquestion-answering10K<n<100K1 likes189 downloads1y agoHugging Face20skymizer /mmlu-redux-2.0-ok-subsample-seed-1234text1K<n<10K0 likes171 downloads26d agoHugging Face21sanderland /CulturaX-subsample-100-bal2Subsample of https://huggingface.co/datasets/uonlp/CulturaX for tokenizer training. Subsampled 1/100 samples per file, then sampled all files at 1/5 except English (1/20) and Russian (1/10). Total size: 19.2 GB commpressed, 35.4 GB uncompressed. File size statistics by language (NB: compressed sizes!): Language Size (MB) Percentage --------------------------------------- en 3478.62459 MB 17.678302% ru 2017.44529 MB 10.252617% es… See the full description on the dataset page: https://huggingface.co/datasets/sanderland/CulturaX-subsample-100-bal2.text1M<n<10M0 likes166 downloads1y agoHugging Face22sanderland /CulturaX-subsample-100-bal2val-1text1M<n<10M0 likes159 downloads1y agoHugging Face23zekeZZ /medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choices-subsampled_traktextn<1K0 likes139 downloads2y agoHugging Face24sanderland /CulturaX-subsample-100-bal2val-4text1M<n<10M0 likes127 downloads1y agoHugging Face25lacg030175 /CIC-IoT-2023-neto-subsample CIC-IoT-2023 — Neto-Subsample (1.3M, 46-feature canonical) Stratified subsample (~1,429,753 rows) of the canonical Neto 46.7M dataset (lacg030175/CIC-IoT-2023-neto-full). Same 46-feature schema as the full version. Drop-in replacement for lacg030175/CIC-IoT-2023 (1.3M bencorn-derived, 39 features) for new experiments needing the canonical feature set. Subsample composition: Benign: 200,000 rows Each attack subclass: up to 50,000 rows NaN/Inf preserved (no dropna). Pair with… See the full description on the dataset page: https://huggingface.co/datasets/lacg030175/CIC-IoT-2023-neto-subsample.tabular1M<n<10M0 likes113 downloads5mo agoHugging Face26sanderland /CulturaX-subsample-100-bal2val-3text1M<n<10M0 likes101 downloads1y agoHugging Face27prdev /subsampled-jxm-nomic-unsupervisedtext10M<n<100M0 likes77 downloads2y agoHugging Face28matyaydin /mcq_subsampledtext100K<n<1M0 likes73 downloads1y agoHugging Face29vidore /docvqa_test_subsampled_tesseractimagedocument-question-answeringn<1K0 likes69 downloads1y agoHugging Face30jonathanli /uground-v1-subsampletabular10K<n<100K0 likes64 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.