datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Cambrian-Alignment
Cambrian-Alignment Dataset
Please see paper & website for more information:
https://cambrian-mllm.github.io/
https://arxiv.org/abs/2406.16860
Overview
Cambrian-Alignment is an question-answering alignment dataset comprised of alignment data from LLaVA, Mini-Gemini, Allava, and ShareGPT4V.
Getting Started with Cambrian Alignment Data
Before you start, ensure you have sufficient storage space to download and process the data.
Download the Data Repository… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/Cambrian-Alignment.FIRM-Video
FIRM-Video-SFT-90K
This repository releases the 90K SFT data for FIRM-Video.
The dataset covers three key evaluation dimensions:
Instruction Following (IF): whether the generated video accurately follows the text prompt.
Visual Quality (VQ): perceptual and technical quality, including clarity, sharpness, artifacts, flicker, and overall visual fidelity.
World Coherence (WC): whether the video is coherent with commonsense, temporal consistency, physical plausibility, and… See the full description on the dataset page: https://huggingface.co/datasets/VisionXLab/FIRM-Video.SEA-Visionoro_depth_rewardOKReddit-Visionary
Dataset Summary
OKReddit Visionary is a collection of 50 GiB (~74K pairs) of image Question & Answers. This dataset has been prepared for research or archival purposes.
Curated by: KaraKaraWitch
Funded by: Recursal.ai
Shared by: KaraKaraWitch
Special Thanks: harrison (Suggestion)
Language(s) (NLP): Mainly English.
License: Refer to Licensing Information for data license.
Dataset Sources
Source Data: Academic Torrents by (stuck_in_the_matrix, Watchful1… See the full description on the dataset page: https://huggingface.co/datasets/recursal/OKReddit-Visionary.Agriculture_VisionAgriculture-VisionAgriculture_Visiongemma-vision-pretrain
