datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XHRBench
XHRBench
Ultra-High-Resolution Remote Sensing Understanding and Reasoning
🤗 Hugging Face ·
🤖 ModelScope ·
📄 Paper ·
💻 Code
English | 中文
📚 Introduction
XHRBench evaluates fine-grained perception and complex reasoning in multimodal large language models using ultra-high-resolution remote-sensing imagery. This repository retains the name XHRBench and belongs to the same RSHR benchmark project as RSHR-Bench, with a… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/XHRBench.HARD-VQA
HARD-VQA
Ultra-High-Resolution Aerial VQA with Original Images Embedded per Question
🤗 Hugging Face ·
🟣 ModelScope ·
📊 Statistics
English | 中文:Hugging Face · ModelScope
📚 Introduction
HARD-VQA packages ultra-high-resolution aerial visual question answering data as self-contained Parquet shards. Each row is one multiple-choice question, with all required original JPEG bytes embedded in its ordered images… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/HARD-VQA.RSHR-Bench
RSHR-Bench
Ultra-High-Resolution Remote Sensing Understanding and Reasoning
🤗 Hugging Face ·
🤖 ModelScope ·
📄 Paper ·
💻 Code
English | 中文
📚 Introduction
RSHR-Bench evaluates ultra-high-resolution remote-sensing visual understanding and reasoning in multimodal large language models across single-image, multi-image, and multi-turn settings. This release embeds original-resolution images directly in Parquet shards for use… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/RSHR-Bench.NJU-HARD-Tracking
NJU-HARD-Tracking
Multi-Object Tracking across 122 MP UAV Image Sequences
🤗 Hugging Face · 🟣 ModelScope · 📊 Statistics: HF / MS
English | 中文: Hugging Face · ModelScope
🌍 Overview
NJU-HARD-Tracking provides the multi-object-tracking release of HARD, with full-resolution frames, temporal ordering, and the original instance annotations. It supports studying how detection and association behave across wide-area aerial… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/NJU-HARD-Tracking.NJU-HARD-Detection
NJU-HARD-Detection
Object Detection in 122 MP Wide-Area UAV Imagery
🤗 Hugging Face · 🟣 ModelScope · 📊 Statistics: HF / MS
English | 中文: Hugging Face · ModelScope
🌍 Overview
NJU-HARD-Detection provides the object-detection release of HARD, with full-resolution frames and per-frame pedestrian and vehicle boxes. Detection training and evaluation use the categories and boxes; the retained source identity fields are… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/NJU-HARD-Detection.NJU-HARD
NJU-HARD
🤗 Hugging Face ·
🟣 ModelScope ·
📊 Statistics
English | 中文:Hugging Face · ModelScope
📚 Introduction
NJU-HARD is the deduplicated, full-resolution release of the HARD visual question answering data. It contains 1,563 valid VQA records across 8 task types, using 854 original aerial images. Only images referenced by these valid questions are included.
Every original JPEG is embedded once in a native… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/NJU-HARD.llava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.RL_ManiObj_AugRSVLM_SFT
RSVLM_SFT
Remote-Sensing Data for Vision-Language Instruction Tuning
Project · Paper · Code
English | 中文
📚 Introduction
RSVLM_SFT is the remote-sensing data repository associated with supervised instruction tuning for MF-RSVLM, the model presented in FUSE-RSVLM: Feature Fusion Vision-Language Model for Remote Sensing. It brings together image resources and source-specific annotations used in remote-sensing vision-language… See the full description on the dataset page: https://huggingface.co/datasets/RL-MIND/RSVLM_SFT.RL_Mixed_MCQ_Filtered_Our_PMCRL_Mixed_MCQ_Filtered_Our_PMC_New_Image_TagRL_MC_Stage2gepa-rlm-exp-20260218-031832
