datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BigEarthNetV2-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-providedpre-conversion of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take precedence. Please refer to the authoritative… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-LMDB.BigEarthNetV2-Lithuania-Summer-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN — Lithuania Summer Subset (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-provided pre-conversion of a subset of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-Lithuania-Summer-LMDB.entitynet
EntityNet
33M web images paired with 46M texts, harvested by querying image search with 135k visual entities from Wikidata.
EntityNet is the training set from Using Knowledge Graphs to Harvest Datasets for Efficient CLIP Model Training (GCPR 2025). Instead of scraping the web and filtering afterwards, we start from a knowledge graph: we extract visual entities and their attributes from Wikidata, turn them into search queries, and collect the returned images together with their… See the full description on the dataset page: https://huggingface.co/datasets/lmb-freiburg/entitynet.LMOD-Plus-Sample
LMOD+ — Sample Subset
A 1,076-instance sample of LMOD+, a large-scale multimodal
ophthalmology benchmark for developing and evaluating multimodal large language models (MLLMs).
This repository is a preview subset intended for quickly inspecting the data format, prototyping
evaluation harnesses, and running smoke tests. The full benchmark contains 32,633 instances across
12 ophthalmic conditions and 5 imaging modalities.
📄 Paper: ACM Transactions on Computing for Healthcare… See the full description on the dataset page: https://huggingface.co/datasets/Euanyu/LMOD-Plus-Sample.OmniBenchmark-1K
OmniBenchmark-1K
OmniBenchmark-1K is a challenging benchmark for Class-Incremental Continual Learning designed to evaluate performance on very long task sequences, ranging from 100 to over 300 non-overlapping tasks.
The dataset was introduced in the paper Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts.
GitHub: https://github.com/LMMMEng/CaRE
Paper: Hugging Face | arXiv
Description
OmniBenchmark-1K provides a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/LMMM2025/OmniBenchmark-1K.
