datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CiQi-VQA
CiQi-Agent
Github | Model | Dataset | Paper
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
Accepted to ECCV 2026
🎯 Overview
CiQi-Agent has been accepted to ECCV 2026.
We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.Monet-SFT-125K
Introduction
This is the SFT dataset for paper "Monet: Reasoning in Latent Visual Space Beyond Images and Language"
Paper: http://arxiv.org/abs/2511.21395
Code: https://github.com/NOVAglow646/Monet
Citation
If you find this work useful, please use the following BibTeX. Thank you for your support!
@misc{wang2025monetreasoninglatentvisual,
title={Monet: Reasoning in Latent Visual Space Beyond Images and Language},
author={Qixun Wang and Yang Shi and Yifei Wang… See the full description on the dataset page: https://huggingface.co/datasets/NOVAglow646/Monet-SFT-125K.HADES
Overview
The HADES benchmark is derived from the paper "Images are Achilles' Heel of Alignment: Exploiting Visual Vulnerabilities for Jailbreaking Multimodal Large Language Models" (ECCV 2024 Oral). You can use the benchmark to evaluate the harmlessness of MLLMs.
Benchmark Details
HADES includes 750 harmful instructions across 5 scenarios, each paired with 6 harmful images generated via diffusion models. These images have undergone multiple optimization rounds, covering… See the full description on the dataset page: https://huggingface.co/datasets/Monosail/HADES.Visual-Math-Eval
Visual Equation Solving Benchmark
This repository contains the dataset introduced in the paper:
Can Vision-Language Models Solve Visual Math Equations? which is currently accepted in EMNLP 2025 (Main)
Despite strong performance in vision and language understanding, Vision-Language Models (VLMs) struggle on tasks requiring integrated perception and symbolic reasoning. This benchmark evaluates VLMs on visual equation solving, where systems of linear equations are represented using… See the full description on the dataset page: https://huggingface.co/datasets/monjoychoudhury29/Visual-Math-Eval.vbpl-vn-legal-corpus
VBPL.vn — Vietnamese Legal Corpus
Văn bản pháp luật Trung ương Việt Nam crawl từ vbpl.vn — Cơ sở dữ liệu quốc gia về pháp luật, Bộ Tư pháp.
Nguồn
vbpl.vn (Bộ Tư pháp Việt Nam)
API: https://vbpl-bientap-gateway.moj.gov.vn/api
Phạm vi: văn bản cấp Trung ương (Quốc hội, Chính phủ, Bộ ngành...)
Cấu trúc
data/
├── vanban.parquet # Bảng chính: 1 dòng = 1 văn bản (~50 cột)
├── vanban.jsonl # Tương đương vanban.parquet dạng JSONL
├──… See the full description on the dataset page: https://huggingface.co/datasets/Monmoonluna/vbpl-vn-legal-corpus.
