datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WN18RRWN18llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.vlwnc-if-vf-universal-class-nanofabricator-v1
Vaelorium Luminex / The Weave NooCathedral InfiLattice / Veyrglass Fabricator "VLWNC-IF-VF" - Universal Class
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific release: v1.0.0 · Hub packaging: hf.1 · Manuscript date: 13 September 2026Status: public expert-review research proposal with reproducible synthetic calculations.
Light-addressed physical compilation for heterogeneous fabrication: a proposed multi-cartridge “light printer in a box” combining… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/vlwnc-if-vf-universal-class-nanofabricator-v1.Mathverse_VLMEvalKitgpt-image-edit-benchmark-results
GPT-Image-Edit — Benchmark Results
This repository contains evaluation results of GPT-Image-Edit across four standard image-editing benchmarks. All scores were computed using the official evaluation scripts provided by each benchmark.
📊 Benchmarks
Benchmark
Metrics
Folder
GEdit-EN
12 editing categories + Avg
gedit/
Complex-Edit
IF, IP, PQ, Overall
complex_edit/
ImgEdit-Full
10 editing operations + Overall
imgedit/
OmniContext
Contextual edit scores… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/gpt-image-edit-benchmark-results.GSM8K-V-VLMEvalKitVLSP2018-ABSA-Hotel
VLSP2018-ABSA-Hotel
Dataset Summary
The VLSP 2018 Hotel corpus is designed for Vietnamese Aspect-Based Sentiment Analysis (ABSA), covering two sub-tasks of Aspect Category Sentiment Analysis (ACSA):
Aspect Category Detection (ACD): identify which Aspect#Category pairs are present in each review.
Sentiment Polarity Classification (SPC): assign one of three sentiment labels (Positive, Negative, Neutral) to each detected Aspect#Category.
This unified CSV contains 5,600… See the full description on the dataset page: https://huggingface.co/datasets/visolex/VLSP2018-ABSA-Hotel.Hateful_Memes_in_VLMThis dataset contains the response of VLMs (InstructBlip, ShareGPT4V, LLaVA and CogVLM) to hateful memes and the annotation to these responses. For more information, please refer to paper "From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language Models."
MME-CoT_VLMEvalKitTox
A cleaned up version of train dataset from kaggle, the Toxic Comment Classification Challenge
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge
https://www.kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge/data?select=train.csv.zip
the alt_format directory contains an alternate format intended for a tutorial.
What was done:
Removed extra spaces and new lines
Removed non-printing characters
Removed punctuation except apostrophe… See the full description on the dataset page: https://huggingface.co/datasets/vluz/Tox.youtube-vlog-personality-recognition-wcpr14The Workshop on Personality Recognition 2014 was a competition based on this dataset. The goal is to predict personality scores from visual features and text transcripts.
Reference Paper: https://infoscience.epfl.ch/server/api/core/bitstreams/e61b4c1b-0c56-4afc-9786-23e9841cb81f/content
B2W-Reviews01OmniMat1K-VLMEvalKit
OmniMatBench 1K Subset for VLMEvalKit
Scope notice: This repository contains OmniMat1K, a 1,000-item
subset with 498 QA and 502 CAL records. It is not the complete
3,171-item OmniMatBench release described in the paper. Evaluation results
produced from this repository must be labeled OmniMat1K and must not be
presented as scores on the complete benchmark.
Release status
This repository contains the 1,000-item OmniMatBench subset prepared for
VLMEvalKit. The… See the full description on the dataset page: https://huggingface.co/datasets/Summer12138/OmniMat1K-VLMEvalKit.VLSP2018-ABSA-Restaurant
VLSP2018-ABSA-Restaurant
Dataset Summary
The VLSP 2018 Restaurant corpus targets the same ACSA sub-tasks (ACD & SPC) on 4,751 Vietnamese restaurant reviews. This unified CSV includes:
12 aspect–category indicator columns, each with values {0, 1, 2, 3}.
A type column for train/dev/test.
A dataset column fixed to VLSP2018-ABSA-Restaurant.
Supported Tasks and Metrics
Aspect Category Detection
Sentiment Polarity Classification
Metrics: Precision, Recall, F1… See the full description on the dataset page: https://huggingface.co/datasets/visolex/VLSP2018-ABSA-Restaurant.ViewSpatial-Bench-vlmevalVLMEvalKit_sourced_DocVQA_valDREAM-1k-VLMEvalKitgdt_vlmevalremote-worker-productivity
Key Features:
Primary Research Focus:
Age vs Productivity correlation
Years of Experience impact on remote work effectiveness
WFH Days per Week optimal balance analysis
Multiple productivity metrics (not just one score)
Dataset Highlights:
1,500 rows - Perfect size for analysis
30+ columns - Rich feature set
Realistic correlations - Built-in meaningful relationships
Clean data - No missing values, proper data types
Multiple target variables - 5 different… See the full description on the dataset page: https://huggingface.co/datasets/Vlad20041903/remote-worker-productivity.CensusVLSP2018-ABSA-Restaurant
VLSP2018-ABSA-Restaurant
Dataset Summary
The VLSP 2018 Restaurant corpus targets the same ACSA sub-tasks (ACD & SPC) on 4,751 Vietnamese restaurant reviews. This unified CSV includes:
12 aspect–category indicator columns, each with values {0, 1, 2, 3}.
A type column for train/dev/test.
A dataset column fixed to VLSP2018-ABSA-Restaurant.
Supported Tasks and Metrics
Aspect Category Detection
Sentiment Polarity Classification
Metrics: Precision, Recall, F1… See the full description on the dataset page: https://huggingface.co/datasets/phuongnt101/VLSP2018-ABSA-Restaurant.VLMEvalKitvlsp_testbook-recommender-dataset
