CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zjunlp /Chat2Workflow-Evaluation Chat2Workflow Chat2Workflow is a benchmark designed for evaluating the ability of Large Language Models (LLMs) to generate executable visual workflows from natural language instructions. Paper: Chat2Workflow: A Benchmark for Generating Executable Visual Workflows with Natural Language Repository: zjunlp/Chat2Workflow Overview Executable visual workflows are widely used in industrial deployments for their reliability and controllability. Chat2Workflow addresses the… See the full description on the dataset page: https://huggingface.co/datasets/zjunlp/Chat2Workflow-Evaluation.documenttext-generationn<1K4 likes3.1k downloads4mo agoHugging Face02FreedomIntelligence /Medical_Multimodal_Evaluation_Data Evaluation Guide This dataset is used to evaluate medical multimodal LLMs, as used in HuatuoGPT-Vision. It includes benchmarks such as VQA-RAD, SLAKE, PathVQA, PMC-VQA, OmniMedVQA, and MMMU-Medical-Tracks. To get started: Download the dataset and extract the images.zip file. Find evaluation code on our GitHub: HuatuoGPT-Vision. This open-source release aims to simplify the evaluation of medical multimodal capabilities in large models. Please cite the relevant benchmark… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/Medical_Multimodal_Evaluation_Data.imageimage-to-text10K<n<100K29 likes346 downloads2y agoHugging Face03demisama /UGround-Offline-Evaluationimage1K<n<10K1 likes238 downloads2y agoHugging Face04mjf-su /PhysicalAI-US-Evaluation PhysicalAI-US-Evaluation A held-out US evaluation set for the navigation planner: 19,744 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective. Provenance Every record here was drawn — uniformly at random — from the pool of US scenes that were withheld from every training stage of the planner: the base VLA pretraining mix, the reasoning supervised fine-tuning (SFT)… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-US-Evaluation.imagerobotics10K<n<100K0 likes23 downloads4mo agoHugging Face05mjf-su /PhysicalAI-DE-Evaluation PhysicalAI-DE-Evaluation A held-out German evaluation set for the navigation planner: 19,999 records, each pairing a single front-camera frame with the corresponding past trajectory, future ground-truth waypoints, and a natural-language driving objective. Provenance Every record here was drawn from the pool of German scenes that were withheld from every training stage of the planner: the base VLA pretraining mix, the reasoning supervised fine-tuning (SFT) stage, and the… See the full description on the dataset page: https://huggingface.co/datasets/mjf-su/PhysicalAI-DE-Evaluation.imagerobotics10K<n<100K0 likes16 downloads4mo agoHugging Face06kyne0127 /vla-evaluation-v3imagen<1K0 likes16 downloads4mo agoHugging Face07kyne0127 /vla-evaluation-v4imagen<1K0 likes2 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.