datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vstat
VSTAT: Visual State Tracking Benchmark
VSTAT is a video-based benchmark for evaluating the visual state tracking
capability of Multimodal Large Language Models (MLLMs). It contains 834 video
clips paired with 1,500 questions whose answers cannot be inferred from any
single keyframe or short segment.
Dataset Composition
Split
Videos
Questions
synthetic
450
550
self_recorded
80
100
youtube
304
850
Total
834
1,500
Files… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/vstat.muse-k2-vision-pilot-20260910
Muse → K2 bridge: first training experiment
Prepared September 10, 2026. This experiment tests whether training a connector
lets the frozen IFM/K2-Horizon-7B decoder use the existing Muse-Glimmer visual
encoder. It does not retrain the vision encoder or K2, and it does not establish
general screenshot, document, natural-image, or visual reasoning capability.
Authorized budget and selected first hardware
The user authorized an initial inexpensive Hugging Face… See the full description on the dataset page: https://huggingface.co/datasets/txgsync/muse-k2-vision-pilot-20260910.lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private
Dataset Card for Evaluation run of Nitral-AI/Eris_PrimeV3.05-Vision-7B
Dataset automatically created during the evaluation run of model Nitral-AI/Eris_PrimeV3.05-Vision-7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 4 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Nitral-AI-Eris_PrimeV3.05-Vision-7B-private.castbench-leaderboard-databasehan-humanoid-vision-object-detection-metrics-v1
Humanoid Vision Detection Metrics Dataset
Overview
Dataset performa sistem visi humanoid saat mendeteksi objek.
Features
image_brightness_level
object_distance_cm
detection_confidence_score
camera_noise_index
frame_processing_time_ms
Target
detection_accuracy_percent
vision-language-action-papers
Vision-Language-Action (VLA) & Robot Learning Papers — FineSet
A research-paper dataset on Vision-Language-Action (VLA) & Robot Learning Papers, assembled, deduplicated, and quality-scored by
FineSet from arXiv and Semantic Scholar.
📸 This is a dated snapshot — generated 2026-06-19.
It is not auto-updated. Research on Vision-Language-Action (VLA) & Robot Learning Papers moves fast — new papers land on arXiv every
week. Want this same dataset refreshed daily, on a topic you… See the full description on the dataset page: https://huggingface.co/datasets/fineset-io/vision-language-action-papers.gpt-4o-mini-vision-traces
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cfahlgren1/gpt-4o-mini-vision-traces.logvision-index
