datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
output_pdf_lmdbBigEarthNetV2-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-providedpre-conversion of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take precedence. Please refer to the authoritative… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-LMDB.crossfps-lmdb-2607sidd-medium-lmdbBigEarthNetV2-Lithuania-Summer-LMDB
TU Berlin
RSiM
DIMA
BigEarth
BIFOLD
reBEN — Lithuania Summer Subset (pre-converted to LMDB)
⚠️ Unofficial mirror. This is an unofficial, community-provided pre-conversion of a subset of the BigEarthNet v2.0 (reBEN) dataset into LMDB format. It is provided as a convenience for researchers who wish to get started quickly without running the full conversion pipeline. In case of any discrepancy, the original publication and the original files always take… See the full description on the dataset page: https://huggingface.co/datasets/hackelle/BigEarthNetV2-Lithuania-Summer-LMDB.sf-lmdb-a14-16k-trunc_lowto merge: python shard_lmdb.py merge --shards-dir . --output merged/
sf-lmdb-a14-16k-trunc_highto merge: python shard_lmdb.py merge --shards-dir . --output merged/
calvin_lmdbcalvin_lmdb_with_depthmoleculenet-unimol-lmdb
MoleculeNet Uni-Mol LMDB Data
Pre-split MoleculeNet datasets in LMDB format from Uni-Mol (ICLR 2023).
Split Protocol
Scaffold splitting (Bemis-Murcko with includeChirality=True)
Ratio: 8:1:1 (train:valid:test)
Following GEM (Fang et al., 2022) and Uni-Mol (Zhou et al., ICLR 2023)
Data Format
Each dataset directory contains train.lmdb, valid.lmdb, and test.lmdb. Each LMDB entry is a pickled dict with keys:
smi: SMILES string
atoms: atom type list
coordinates:… See the full description on the dataset page: https://huggingface.co/datasets/leon14159/moleculenet-unimol-lmdb.kfold-multimer-lmdb
kfold-multimer-lmdb
Training data for the multimer / protein-ligand fine-tune in
hsjang0/k-fold-multimer-tuning.
Pre-tokenized complexes as LMDB, plus the eval splits and label tables the configs read.
Do not clone the whole thing to run one experiment. It is 201 GB and a single arm reads a
handful of its sources. The repo's scripts/fetch_data.py takes a config name, works out
which sources that config actually needs, and pulls only those.
python scripts/fetch_data.py --config… See the full description on the dataset page: https://huggingface.co/datasets/hyosoon0/kfold-multimer-lmdb.vipe_long_lmdbseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.3D-scancomp-objaverse-test-lmdb
tags:
- 3D
- scan
- completion
- SDF
- voxel
- lmdb
size_categories:
- 10B<n<100B
Latent Uncertainty-Aware Multi-View SDF Scan Completion
own_lmdb_datavln_trajectory_lmdbcalvin_lmdbmalayalam-ocr-lmdbLMDB_Overhead_Geopose
📦 Geocentric Height Regression Dataset
Large-scale, preprocessed dataset for dense height regression of anthropogenic objects from single satellite RGB imagery. Optimized for streaming PyTorch training with PyArrow and IterableDataset.
📖 Overview
This dataset contains 256×256 patches extracted from satellite RGB imagery and corresponding Above-Ground Level (AGL) height maps (DSM). Each patch includes a binary validity mask and capture metadata, enabling robust… See the full description on the dataset page: https://huggingface.co/datasets/Lexxurius/LMDB_Overhead_Geopose.poc-slt-shapenet-test-lmdbPOC-SLT: Partial Object Completion with SDF Latent Transformers
FFHQ512.lmdblsun-bedroom-256-lmdb
LSUN Bedroom 256x256 (LMDB) for DiffAE
This dataset repo hosts the LSUN Bedroom train split packaged as a single LMDB
(bedroom256.lmdb) compatible with DiffAE’s legacy LMDB reader and the new
HF-backed bedroom_hf wrapper.
Related project: https://github.com/konpatp/diffae
What’s inside
bedroom256.lmdb/data.mdb
bedroom256.lmdb/lock.mdb
Format
Keys: 256-0000000, 256-0000001, ... (7-digit zero-padded index)
Values: RGB images encoded as WebP (quality ~90), resized to 256x256
Total… See the full description on the dataset page: https://huggingface.co/datasets/konpat/lsun-bedroom-256-lmdb.wan-2-1-14b-ode-lmdbdm-cad.lmdb
Multimodal CAD Dataset
A comprehensive multimodal dataset for CAD (Computer-Aided Design) understanding and generation tasks.
Description
This dataset contains multiple modalities of CAD data, designed for research in CAD understanding, generation, and multi-modal learning. Each CAD model is represented in several complementary formats.
Data Modalities
Modality
Directory
Format
Description
Description
cad_desc/
JSON
Natural language descriptions of… See the full description on the dataset page: https://huggingface.co/datasets/ZackaryJing/dm-cad.lmdb.fasttext-lmdbwan_animation_lmdbpoc-slt-abc-test-lmdb
tags:
- SDF
- Completion,
- Transformers,
- 3D
datasets:
Meshes from ABC dataset
POC-SLT: Partial Object Completion with SDF Latent Transformers
lmd_bass_1000_autotokenizedCelebA178-all-lmdbode_pairs_lmdb
