datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DrivAerML_subsampled_10x
DrivAerML Subsampled 10x
A subsampled and compressed version of the DrivAerML dataset by Ashton et al. (2024), prepared for convenient use with ML frameworks for automotive aerodynamics tasks such as drag and lift coefficient prediction.
Disclaimer
This dataset is provided by Emmi AI for convenience only, on an "as-is" basis, and without any warranty, express or implied. Emmi AI does not own, and does not claim any ownership or rights over, the underlying data. All… See the full description on the dataset page: https://huggingface.co/datasets/EmmiAI/DrivAerML_subsampled_10x.arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled.
docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled.
tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled.
infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled.
docvqa_test_subsampled_beirBEIR version of vidore/docvqa_test_subsampled.
infovqa_test_subsampled_beirBEIR version of vidore/infovqa_test_subsampled.
arxivqa_test_subsampled_beirBEIR version of vidore/arxivqa_test_subsampled.
tabfquad_test_subsampled_beirBEIR version of vidore/tabfquad_test_subsampled.
docvqa_test_subsampled
Dataset Description
This is the test set taken from the DocVQA dataset. It includes collected images from the UCSF Industry Documents Library. Questions and answers were manually annotated.
Example of data (see viewer)
Data Curation
To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns.
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/docvqa_test_subsampled.mlqa__subsampledtabfquad_test_subsampled
Dataset Description
TabFQuAD (Table French Question Answering Dataset) is designed to evaluate TableQA models in realistic industry settings. Using a vision language model (GPT4V), we create additional queries to augment the existing human-annotated ones.
Data Curation
To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 280 pairs, leaving the rest for training and renaming the different columns.
Load the dataset
from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/tabfquad_test_subsampled.arxivqa_test_subsampled
Dataset Description
This is a VQA dataset based on figures extracted from arXiv publications taken from ArXiVQA dataset from Multimodal ArXiV. The questions were generated synthetically using GPT-4 Vision.
Data Curation
To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs. Furthermore we renamed the different columns for our purpose.
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/arxivqa_test_subsampled.infovqa_test_subsampled
Dataset Description
This is the test set taken from the InfoVQA dataset. includes infographics collected from the Internet using the search query “infographics”. Questions and answers were manually annotated.
Questions and answers were manually annotated.
Example of data : (see viewer)
Data Curation
To ensure homogeneity across our benchmarked datasets, we subsampled the original test set to 500 pairs and renamed the different columns.
Load the dataset
from… See the full description on the dataset page: https://huggingface.co/datasets/vidore/infovqa_test_subsampled.logiqa2__subsampledflanv2_subsamplelogiqa2__subsampledatomic_subsampled_500kfindingdory-subsampled-96
FindingDory: A Benchmark to Evaluate Memory in Embodied Agents
Karmesh Yadav*,
Yusuf Ali*,
Gunshi Gupta,
Yarin Gal,
Zsolt Kira
Current vision-language models (VLMs) struggle with long-term memory in embodied tasks. To address this, we introduce FindingDory, a benchmark in Habitat that evaluates memory-based reasoning across 60 long-horizon tasks.
In this repo, we release the FindingDory Subsampled Video Dataset. Each video contains 96 images… See the full description on the dataset page: https://huggingface.co/datasets/yali30/findingdory-subsampled-96.mmlu-redux-2.0-ok-subsample-seed-1234CulturaX-subsample-100-bal2val-1CulturaX-subsample-100-bal2Subsample of https://huggingface.co/datasets/uonlp/CulturaX for tokenizer training.
Subsampled 1/100 samples per file, then sampled all files at 1/5 except English (1/20) and Russian (1/10).
Total size: 19.2 GB commpressed, 35.4 GB uncompressed.
File size statistics by language (NB: compressed sizes!):
Language Size (MB) Percentage
---------------------------------------
en 3478.62459 MB 17.678302%
ru 2017.44529 MB 10.252617%
es… See the full description on the dataset page: https://huggingface.co/datasets/sanderland/CulturaX-subsample-100-bal2.medmcqa-gen-by-zephyr-ft-gpqa-all-sorted-contrastive-with-choices-subsampled_trakCulturaX-subsample-100-bal2val-4CIC-IoT-2023-neto-subsample
CIC-IoT-2023 — Neto-Subsample (1.3M, 46-feature canonical)
Stratified subsample (~1,429,753 rows) of the canonical Neto 46.7M
dataset (lacg030175/CIC-IoT-2023-neto-full). Same 46-feature schema as the
full version. Drop-in replacement for lacg030175/CIC-IoT-2023 (1.3M
bencorn-derived, 39 features) for new experiments needing the canonical
feature set.
Subsample composition:
Benign: 200,000 rows
Each attack subclass: up to 50,000 rows
NaN/Inf preserved (no dropna). Pair with… See the full description on the dataset page: https://huggingface.co/datasets/lacg030175/CIC-IoT-2023-neto-subsample.CulturaX-subsample-100-bal2val-3subsampled-jxm-nomic-unsuperviseddocvqa_test_subsampled_tesseractCulturaX-subsample-100-bal2val-2OpenThoughts3-full-filtered-code-subsampled
