CoolFace
20 results

ppp

pppop7 /GQA Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets This Dataset This is a formatted version of GQA. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models. @inproceedings{hudson2019gqa, title={Gqa: A new dataset for real-world visual reasoning and compositional question… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/GQA.image10M<n<100M0 likes910 downloads9mo agoHugging FacePPPIMM /UKKIII0 likes759 downloads1y agoHugging FacePPPIMN /UKKIIL0 likes507 downloads1y agoHugging Facepppop7 /anamnesis-bench AnamnesisBench AnamnesisBench is an evaluation benchmark for numerical reliability in LLM research agents. It focuses on a practical failure mode: an agent writes or accepts a financial research artifact that looks plausible, but contains a wrong, unsupported, or misattributed number. The benchmark is not intended as training data. It is a set of test cases, source packets, expected truth values, and deterministic scoring scripts. You run your own model or verifier, then score… See the full description on the dataset page: https://huggingface.co/datasets/pppop7/anamnesis-bench.tabulartext-generation1K<n<10K0 likes456 downloads3mo agoHugging Facepppppppppp2 /planeperturbed Dataset Card for "planeperturbed" More Information needed image1K<n<10K1 likes306 downloads3y agoHugging Facepppop7 /OCR-VQA Dataset Card for "OCR-VQA" More Information needed image100K<n<1M0 likes268 downloads9mo agoHugging Face