datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MWS-Antifraud-Bench
MWS Antifraud Bench (Validation)
Experimental document-authenticity task for general-purpose multimodal language
models. This is the public validation part of MWS Vision Bench anti-fraud v0.1.
The dataset is released for research and model comparison. It is not a
certification tool, a production fraud-detection system, or a universal
leaderboard that is expected to be resistant to deliberate optimization.
Data
The validation split contains 209 items:
44 ai_gen;… See the full description on the dataset page: https://huggingface.co/datasets/MTSAIR/MWS-Antifraud-Bench.eurosat-rgb
EuroSat (RGB)
Description
A dataset based on Sentinel-2 satellite images covering 13 spectral bands and consisting of 10 classes with 27000 labeled and geo-referenced samples. This is the RGB version of the dataset with visible bands encoded as JPEG images.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/mteb/eurosat-rgb.resisc45
Description
RESISC45 dataset is a publicly available benchmark for Remote Sensing Image Scene Classification (RESISC), created by Northwestern Polytechnical University (NWPU). This dataset contains 31,500 images, covering 45 scene classes with 700 images in each class.
The dataset does not have any default splits. Train, validation, and test splits were based on these definitions here… See the full description on the dataset page: https://huggingface.co/datasets/mteb/resisc45.crisismmd-mteb
CrisisMMD for MTEB
This repository contains two image+text classification configurations derived
from the official QCRI/CrisisMMD
release at revision 10a5626ba112ab9c50369c473a2062bcf205d708.
informative: useful vs. not useful for humanitarian response.
humanitarian: five humanitarian information categories.
For informative and humanitarian, only rows in the authors' official
text_img_agreed_lab splits are retained. Their independently annotated image
and tweet labels are… See the full description on the dataset page: https://huggingface.co/datasets/pranitchawla/crisismmd-mteb.mlcd-mteb-cifar-eval
MLCD vs CLIP on MTEB CIFAR-10/100: integration and evaluation
Evaluation results accompanying the MTEB integration of two MLCD image encoders
(PR #5406, resolving
issue #2571).
Two DeepGlint-AI MLCD encoders were integrated into MTEB, verified against the
reference implementation, and evaluated on the official MTEB CIFAR-10/CIFAR-100
image-classification tasks alongside size-matched OpenAI CLIP baselines.
What was measured
Official MTEB image classification: 5… See the full description on the dataset page: https://huggingface.co/datasets/b4ph/mlcd-mteb-cifar-eval.glami-1m-mteb
GLAMI-1M MTEB multimodal classification
This is an MTEB-ready derivative of the official
glami/glami-1m
release for multilingual image+text fashion classification. The source is
pinned at revision befda45d8d4e8b8082bb8a1912d1f9eb9483991c and remains
licensed under Apache-2.0.
Each example contains the official product image, name and description
joined as text, and the official category ID as label. The complete
116,004-row human-labeled test split is unchanged.
To keep… See the full description on the dataset page: https://huggingface.co/datasets/artist/glami-1m-mteb.war-gov-uap-release-1
Department of War UAP Release 1 — structured corpus
The first tranche of declassified U.S. government records on Unidentified
Anomalous Phenomena (UAP / UFOs), released by the Department of War on
8 May 2026 under the Presidential Unsealing and Reporting System for
UAP Encounters (PURSUE) directive.
This dataset is a structured, machine-readable companion to the source
material at https://www.war.gov/UFO/. It pairs every original document
with VLM-extracted page text, cropped… See the full description on the dataset page: https://huggingface.co/datasets/MTSlive/war-gov-uap-release-1.City_mapThis dataset contains over 600 maps images of 45 various city around the world.
For trouble-shooting with the dataset, you may use this script to identify potentially corrupted files.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
