datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SenseNova-Vision-Corpus-50M
Vision as Unified Multimodal Generation
English | 简体中文
This repository contains the dataset for the paper Vision as Unified Multimodal Generation.
SenseNova Vision Corpus 50M
Overview
SenseNova Vision Corpus 50M (SN-VC-50M) is a large-scale multimodal vision corpus designed for unified training across diverse visual understanding and geometry-oriented tasks. The dataset is curated to address a common limitation of existing… See the full description on the dataset page: https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M.blip3-grounding-50m
BLIP3-GROUNDING-50M Dataset
Overview
The BLIP3-GROUNDING-50M dataset is designed to enhance the ability of Vision-Language Models (VLMs) to ground semantic concepts in visual features, which is crucial for tasks like object detection, semantic segmentation, and understanding referring expressions (e.g., "the object to the left of the dog"). Traditional datasets often lack the necessary granularity for such tasks, making it challenging for models to accurately localize and… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-grounding-50m.blip3-grounding-50mDatacomp-50mocr-mlt-50m
OCR-MLT-50M: Multilingual OCR Corpus
A large-scale multilingual OCR dataset spanning 50 languages and 50.2 million image-text pairs.Designed for training and evaluating robust multilingual text recognition systems across diverse scripts and domains.
📄 Paper |
🤗 Model |
🔥 Demo |
💻 GitHub |
🏆 Leaderboard |
📊 Weights & Biases
🔥 News
[2025-11-15] OCR-MLT-50M is now available on Hugging Face! Download here
[2025-10-28] Our paper is accepted at CVPR 2025!… See the full description on the dataset page: https://huggingface.co/datasets/interfaze-ai/ocr-mlt-50m.
