datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
PP
SteelBench: A Diagnostic Benchmark for Vision-Language Models in Industrial Safety Monitoring
SteelBench is a diagnostic benchmark of densely annotated CCTV clips from an
operating integrated steel plant. It is designed to evaluate vision-language
models (VLMs) on real-world industrial action recognition, PPE assessment,
and safety-violation detection — under naturally occurring degradation
(dust, glare, steam, low light), at distances and crowdedness levels that
curated… See the full description on the dataset page: https://huggingface.co/datasets/ThinkingHub/PP.european-flora-fungi-thinking
🌿🍄 European Flora & Fungi Identification Dataset
With Synthetic Thinking Traces for Vision-Language Model Fine-Tuning
A curated dataset of 377 multi-turn conversations covering 39 European plant and mushroom species organized into 14 confusion groups (species commonly mistaken for each other).
🎯 Purpose
This dataset is designed for fine-tuning vision-language models (Gemma 3, LLaVA, Qwen-VL) with thinking traces (<think>...</think>) to perform:
Species… See the full description on the dataset page: https://huggingface.co/datasets/Mightypeacock/european-flora-fungi-thinking.oxford-pets-grpo-think
Oxford-IIIT Pet — GRPO training data (structured reasoning)
GRPO (verl) training data for Oxford-IIIT Pet breed classification with a structured-reasoning prompt: the model emits a scratchpad tagging visible properties (HasProperty), parts (HasA), and setting (AtLocation) before the label. Reward: 0.30 for a well-formed think block, 0.70 for the label match.
Splits: train 2,944 rows, test 3,669 rows.
Schema
column
type
data_source
string
prompt… See the full description on the dataset page: https://huggingface.co/datasets/jucamohedano/oxford-pets-grpo-think.deep-think
