CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01McGill-NLP /agent-reward-bench AgentRewardBench 💾Code 📄Paper 🌐Website 🤗Dataset 💻Demo 🏆Leaderboard AgentRewardBench: Evaluating Automatic Evaluations of Web Agent TrajectoriesXing Han Lù, Amirhossein Kazemnejad*, Nicholas Meade, Arkil Patel, Dongchan Shin, Alejandra Zambrano, Karolina Stańczak, Peter Shaw, Christopher J. Pal, Siva Reddy*Core Contributor Loading dataset You can use the huggingface_hub library to load the dataset. The dataset is available on Huggingface Hub at… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/agent-reward-bench.imagerobotics1K<n<10K4 likes19k downloads1y agoHugging Face02McGill-NLP /CHASE-Code CHASE: Challenging AI with Synthetic Evaluations The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-Code.imagen<1K0 likes89 downloads2y agoHugging Face03caiyuhu /MCiteBench MCiteBench Dataset MCiteBench is a benchmark for evaluating the ability of Multimodal Large Language Models (MLLMs) to generate text with citations in multimodal contexts. Websites: https://caiyuhu.github.io/MCiteBench Paper: https://arxiv.org/abs/2503.02589 Code: https://github.com/caiyuhu/MCiteBench Data Download Please download the MCiteBench_full_dataset.zip. It contains the data.jsonl file and the visual_resources folder. Data Statistics… See the full description on the dataset page: https://huggingface.co/datasets/caiyuhu/MCiteBench.imagetext-generation1K<n<10K0 likes65 downloads1y agoHugging Face04McGill-NLP /CHASE-QA CHASE: Challenging AI with Synthetic Evaluations The pace of evolution of Large Language Models (LLMs) necessitates new approaches for rigorous and comprehensive evaluation. Traditional human annotation is increasingly impracticable due to the complexities and costs involved in generating high-quality, challenging problems. In this work, we introduce **CHASE**, a unified framework to synthetically generate challenging problems using LLMs without human involvement. For a given task… See the full description on the dataset page: https://huggingface.co/datasets/McGill-NLP/CHASE-QA.imagequestion-answeringn<1K0 likes57 downloads2y agoHugging Face05mcgillailab /newspapers2024image1M<n<10M0 likes18 downloads1y agoHugging Face06JourneyBench /JourneyBench_MCOTimagequestion-answering1K<n<10K1 likes11 downloads2y agoHugging Face07Krishna5T /mcqimagen<1K0 likes5 downloads8mo agoHugging Face08andreidemianko /t-shopping-mcqaimage1K<n<10K1 likes1 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.