datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sat-image-boundingbox-sft-full
NU-TONIC raw SFT Full
Satellite imagery and aligned land-cover outputs packaged as image–text rows for fine-tuning in SFT format. JSONL user prompts name the modality (satellite imagery vs. overhead context) where it matters.
Provenance
Locations: GeoGuessr-style POIs (source: stochastic/random_streetview_images_pano_v0.0.2)
Optical: Sentinel-2 multispectral optical COGs from a public STAC catalog, blue/green/red or visual preview, percentile-stretched to uint8.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-image-boundingbox-sft-full.stargate_s04e01_100topkdiverse_text2vid
librivox-catalog-full
LibriVox Catalog Full
A full backup of the LibriVox catalog from the LibriVox API, including cover art. (Does not include audio files themselves)
Last updated: 1/26/25
ist-vqa-zip-full
IST-VQA — Multilingual Scene Text VQA Dataset
Curated dataset for Indic Scene Text Visual Question Answering covering three languages:
Bengali (bn) · Hindi (hi) · Tamil (ta)
Summary
Bengali
Hindi
Tamil
Total
Total Images
2,308
2,097
2,322
6,727
Total VQA Pairs
2,306
2,097
2,321
6,724
— Crawl images
1,615
1,570
1,394
4,579
— IndicSTR12 images
357
173
336
866
— Synthetic images
336
354
592
1,282
Sources: Real-world web-crawled scene photos… See the full description on the dataset page: https://huggingface.co/datasets/somusan/ist-vqa-zip-full.Arabic_dataset_13M_translated_cleaned_v2_jsonl_format_ViT-B-16-plus-240-fulldata-v2DatasetDict({
train: Dataset({
features: ['index', 'embeddings', 'en_caption', 'ar_caption', 'nr_words', 'url'],
num_rows: 12166802
})
})
SopriBench-fullPaper: https://arxiv.org/abs/2606.06784
SopriBench is the first large-scale multimodal benchmark dedicated to evaluating privacy inference risks on social media platforms. It contains 50 user profiles and 1,569 images with fine-grained annotations of 28 personal attributes across 4 dimensions: basic information, socioeconomic status, lifestyle, and sensitive information.
The dataset is designed to facilitate research on privacy-preserving AI, adversarial privacy attacks, and multimodal… See the full description on the dataset page: https://huggingface.co/datasets/Nina-Huang/SopriBench-full.vlfeedback_fullbackup-thai_dolma-gamble_web_full
