ppp
Datasets
All datasets matching “ppp”GQA
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{hudson2019gqa,
title={Gqa: A new dataset for real-world visual reasoning and compositional question… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/GQA.UKKIIIUKKIILanamnesis-bench
AnamnesisBench
AnamnesisBench is an evaluation benchmark for numerical reliability in LLM research agents.
It focuses on a practical failure mode: an agent writes or accepts a financial research artifact that
looks plausible, but contains a wrong, unsupported, or misattributed number.
The benchmark is not intended as training data. It is a set of test cases, source packets, expected
truth values, and deterministic scoring scripts. You run your own model or verifier, then score… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/anamnesis-bench.planeperturbed
Dataset Card for "planeperturbed"
More Information needed
OCR-VQA
Dataset Card for "OCR-VQA"
More Information needed
