datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MAD-Bench
MAD-Bench
A Benchmark for Evaluating Deceptive Behaviors in Multimodal Computer-Use Agents.
As MLLMs and computer-use agents increasingly take control of our desktops, safety concerns must evolve beyond text-based prompt injection. MAD-Bench is the first comprehensive benchmark designed to evaluate deceptive behaviors of multimodal agents — cases where an agent fabricates evidence, falsely reports success, ignores conflicting visual feedback, or otherwise produces dishonest outputs… See the full description on the dataset page: https://huggingface.co/datasets/goldenash/MAD-Bench.adaption-mmmed-autoscientist-gold
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
adaption-mmmed_autoscientist_gold
A refined, high-entropy multimodal clinical benchmark containing 431 perfectly aligned pairs of visual medical artifacts (X-rays, CT scans, ultrasounds, and histopathology profiles) and pre-concatenated case narratives with multiple-choice pathways. Optimized specifically for the AutoScientist Challenge (Healthcare Track) to train and evaluate… See the full description on the dataset page: https://huggingface.co/datasets/asadullahdogarr/adaption-mmmed-autoscientist-gold.aiconf-butterfly-detection-goldenset-extendedExtended goldenset for butterfly detection built from the original goldenset and a validated expansion pass.
Files:
larger_goldenset.json
larger_goldenset.tsv
Columns:
photo_id
image
entity
bbox
Generated at: 2026-04-19 23:31:19 UTC
Rows: 356
