datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unified-agent-trajectories
Unified Benchmark Agent Trajectories
Dataset release: v2.1.1 (2026-09-18)Record format: unified-agent-sft-v1
A growing collection of benchmark agent execution trajectories converted into one
transparent, multimodal, tool-aware representation. These are complete recorded benchmark
runs—not ordinary chat transcripts—including benchmark tasks, model reasoning and answers,
tool calls, tool observations, runtime status, and benchmark scores when available. The
directory layout is… See the full description on the dataset page: https://huggingface.co/datasets/ChrisDing1105/unified-agent-trajectories.UnifiedReward-2.0-T2X-score-data
Dataset Summary
UnifiedReward-2.0-T2X-score-data is added for our UnifiedReward-2.0-qwen-[3b/7b/32b/72b] training.
This dataset enables UnifiedReward-2.0 introducing several new capabilities:
Pairwise scoring for image and video generation assessment on Alignment, Coherence, Style dimensions.
Pointwise scoring for image and video generation assessment on Alignment, Coherence/Physics, Style dimensions.
Welcome to try the latest version, and the inference code is available at… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-2.0-T2X-score-data.unified-vlm-steering-emu35
Emu3.5 (BAAI): activation steering sweeps
Steered text and image generations from Emu3.5 (BAAI), one of the unified vision-language models in the unified-vlm-steering project. Steering adds alpha * v_hat (the per-layer unit difference-of-means vector) to the residual stream at every layer of a layer config.
34B dense autoregressive model, 64 decoder layers (0-indexed). Images are 32x32 IBQ tokens (512 px), generated on BAAI's patched vLLM engine with request-id-keyed CFG. The… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-emu35.unified-vlm-steering-uniar
UniAR (ShareLab-SII/UniAR-RL): activation steering sweeps
Steered text and image generations from UniAR (ShareLab-SII/UniAR-RL), one of the unified vision-language models in the unified-vlm-steering project. Steering adds alpha * v_hat (the per-layer unit difference-of-means vector) to the residual stream at every layer of a layer config.
Qwen3-VL backbone, 36 decoder layers (0-indexed). Images are BSQ tokens rendered by an SD3 decoder (16 decoding steps, 544 px).… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-uniar.unified-vlm-steering-liquid
Liquid (FoundationVision Liquid_V1_7B): activation steering sweeps
Steered text and image generations from Liquid (FoundationVision Liquid_V1_7B), one of the unified vision-language models in the unified-vlm-steering project. Steering adds alpha * v_hat (the per-layer unit difference-of-means vector) to the residual stream at every layer of a layer config.
Gemma-7B backbone, 28 decoder layers (0-indexed). Images are VQGAN codes (512 px), CFG 7.0, top-k 4096, top-p 0.96… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-liquid.UnifiedReward-Flex-SFT-90K
UnifiedReward-Flex-SFT-90K
This repository releases 90K SFT data of UnifiedReward-Flex.
For further details, please refer to the following resources:
📰 Paper: https://arxiv.org/abs/2602.02380
🪐 Project Page: https://codegoat24.github.io/UnifiedReward/flex
🤗 Model Collections: https://huggingface.co/collections/CodeGoat24/unifiedreward-flex
🤗 Dataset: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K
👋 Point of Contact: Yibin Wang
Citation… See the full description on the dataset page: https://huggingface.co/datasets/CodeGoat24/UnifiedReward-Flex-SFT-90K.OpenDetection-80K-Unified-Cleaned
OpenDetection-80K-Unified-Cleaned
OpenDetection-80K-Unified-Cleaned is a large-scale object detection dataset built primarily from general, publicly available images, which make up the majority of the input imagery, together with additional publicly available datasets. This dataset is a unified collection created by combining OpenDetection-15K-Dense-v1.0, OpenDetection-15K-Dense-v2.0, and OpenDetection-50K-Remastered-Cleaned into a single standardized dataset. Every sample… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenDetection-80K-Unified-Cleaned.unified_autonomy_stack_datasets
Unified Autonomy Stack Datasets
This repository contains data released in relation with the Unified Autonomy Stack. Details can be found here
Paper: https://arxiv.org/abs/2605.12735
Code: https://ntnu-arl.github.io/unified_autonomy_stack/
The data was collected using the following platforms in both manually piloted and autonomously operated modes:
AR-1 (Hornbill): A variant of the RMF-Owl collision-tolerant aerial robot.
AR-2 (Magpie): A collision-tolerant aerial robot… See the full description on the dataset page: https://huggingface.co/datasets/ntnu-arl/unified_autonomy_stack_datasets.Face-3D-Unified-Preferences
Face-3D-Unified-Preferences
Face-3D-Unified-Preferences is a high-quality dataset designed for human face depth estimation, 3D face reconstruction, and unified preference learning. The dataset is a mixture of male and female human face portraits, providing a diverse collection of facial appearances, identities, poses, and expressions for training modern computer vision and multimodal AI models. Each sample contains an RGB face image, a dense facial depth map, and a… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Face-3D-Unified-Preferences.unified-vlm-steering-minimal-pairs
Unified VLM Steering: Minimal Pairs
Minimal pairs used to extract steering vectors (difference of means between the two poles) for 7 concepts in the
unified-vlm-steering project: 100 text pairs and 100 image pairs per concept.
Layout
txt/<concept>/pairs.json 100 text pairs: concept, pos_label, neg_label, template, n_pairs,
pairs [{subject, pos, neg}], and the pos / neg sentence lists
img/<concept>/<000-099>/
baseline.png… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-minimal-pairs.Animated-World-150M-v1-UnifiedUnified_Road_Defect_Dataset
Unified Road Defect Dataset
A merged, YOLO-format road-defect detection dataset that combines RDD-2022
(primary, ground-level, 6 countries) with two supplementary aerial/drone
datasets — UAV-PDD2023 (China) and RoadDamageVision (China + Spain) —
into a single 4-class CRDDC schema.
This is a derived dataset. It re-packages and re-labels images from three
independently published sources. All credit for the underlying images and
original annotations belongs to their respective… See the full description on the dataset page: https://huggingface.co/datasets/TamAko783/Unified_Road_Defect_Dataset.unified-vlm-steering-eval
Unified VLM steering eval
Activation steering of unified vision-language models: steered text and image generations, steering vectors and
judge scores. One folder per model; prompts/ is shared. Steering is h <- h + alpha * v_hat everywhere.
model
run
text rows
image rows
concepts with images
published
uniar/
av_v2
21,560
21,560
age, chaos, cleanness, emotion, near_far, size, spatial_lr
2026-09-11
Layout per model: manifest.json (exact config)… See the full description on the dataset page: https://huggingface.co/datasets/saintsauce/unified-vlm-steering-eval.power-ovd-unifiedblv-assistive-vision-unifiedSEA_CulturalGround_MCQs_formatted_with_unifiedrewardpitvqa-unified-vlm
PitVQA Unified VLM Classification Dataset
Surgical workflow classification dataset for training vision-language models on pituitary surgery phase detection, step recognition, and instrument identification.
🔗 GitHub: https://github.com/matheus-rech/pit_project
🤖 Trained Model: mmrech/pitvqa-qwen2vl-unified
📄 Original Dataset: UCL Research Data Repository
Dataset Description
This dataset contains 5,184 surgical frames with classification annotations for surgical… See the full description on the dataset page: https://huggingface.co/datasets/mmrech/pitvqa-unified-vlm.SEA_CulturalGround_OE_formatted_with_unifiedrewardOpenCaption-Unified-10K
OpenCaption-Unified-10K
OpenCaption-Unified-10K is a dense image captioning dataset built from 10,000 images paired with long-form synthetic captions generated using the Qwen3.5 multimodal model. Each caption is produced through a dedicated Qwen3.5 captioning pipeline designed to generate detailed, high-fidelity descriptions of scene composition, subject attributes, spatial relationships, activities, and overall visual context rather than short, generic captions. The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/OpenCaption-Unified-10K.shader-unified6kUnified-Animals-Dataset
Unified Animal Dataset
This dataset contains a large-scale collection of animal images designed for multi-class image classification tasks. It includes 57,329 images across 216 animal categories and is suitable for training, evaluation, and benchmarking of computer vision models.
Dataset Details
Dataset Description
The Unified Animal Dataset is a curated multi-source dataset combining several publicly available animal image datasets into a unified… See the full description on the dataset page: https://huggingface.co/datasets/Prentz/Unified-Animals-Dataset.laser-order-unified-lora-v1-datasetfigmirror-unified
Unified FigMirror Dataset
Canonical release with 550 samples.
dataset_augmentation: 500
paper_derivative: 50
paper_derivative verified_pass: 50
Data unit:
One row in data/train.jsonl is one task / one data point.
Asset files under assets/ are supporting files, not separate data points.
Semantic task families:
chart_style_augmentation: input is a reference/source chart; output is an augmented chart.
paper_figure_reproduction: input is a paper figure reference; output is a… See the full description on the dataset page: https://huggingface.co/datasets/zcahjl3/figmirror-unified.megalith-unifiedfitcheck-unified-allachar-raw-female-unified-512Unified-VideoDA-Generated-FlowsOptical flows associated with our work "We're Not Using Videos Effectively: An Updated Domain Adaptive Video Segmentation Baseline"
See the github for full instructions, but to install run
git lfs install
git clone https://huggingface.co/datasets/hoffman-lab/Unified-VideoDA-Generated-Flows
achar-raw-female-unified-portrait-512unified-vqa-imagesHijja-Dhad-Unified-V4
Dataset Card for "Hijja-Dhad-Unified-V4"
More Information needed
