datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
open-pmc-18m
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc-18m.cwicr-vector-db-bgem3-v3
CWICR Vector Database — BGE-M3 V3 Snapshots
Production Qdrant snapshots for CWICR (Construction Works Items, Costs & Resources) — a multilingual catalogue of construction rate databases covering 30 countries / language locales. Each snapshot encodes one country's rate book using the BAAI/bge-m3 embedder and is ready to restore directly into a Qdrant server for hybrid semantic search.
These snapshots are the V3 production artifacts produced by the OpenConstructionEstimate / CWICR… See the full description on the dataset page: https://huggingface.co/datasets/DataDrivenConstruction/cwicr-vector-db-bgem3-v3.open-pmc
OPEN-PMC
Arxiv: Arxiv
|
Code: Open-PMC Github
|
Model Checkpoint: Hugging Face
Dataset Summary
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
Extracted images from research articles.… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/open-pmc.GeoFidelity-Bench
GeoFidelity-Bench
GeoFidelity-Bench evaluates whether generated street-view images match a
requested location at the level of named street blocks. The release contains
109 named street blocks from 25 cities, 7,117 curated Mapillary reference
images, generated images from six open-weight text-to-image models, prompt
control metadata, and 109-target benchmark result summaries. The generated-image index covers
15,696 released JPEG files across six models, six prompt or control… See the full description on the dataset page: https://huggingface.co/datasets/moss-vector-714/GeoFidelity-Bench.onevision1.5MobileGym-ConAct-Trajectories
MobileGym-ConAct-Trajectories
Dataset Viewer · MobileGym · MemGUI-Agent · Paper
Abstract
MobileGym-ConAct-Trajectories is a release of successful mobile GUI-agent rollouts collected in the MobileGym simulator. Each trajectory is selected from judge-verified rollouts using a deterministic per-task rule: retain the shortest structurally valid success, then break ties by source run and episode ID. The release preserves screenshots, the rendered prompt supplied to the… See the full description on the dataset page: https://huggingface.co/datasets/Ma-Vector/MobileGym-ConAct-Trajectories.HumaniBench
HumaniBench: A Human-Centric Benchmark for Large Multimodal Models Evaluation
**HumaniBench** is a benchmark for evaluating large multimodal models (LMMs) using real-world, human-centric criteria. It consists of 32,000+ image–question pairs across 7 tasks:
✅ Open/closed VQA
🌍 Multilingual QA
📌 Visual grounding
💬 Empathetic captioning
🧠 Robustness, reasoning, and ethics
Each example is annotated with GPT-4o drafts, then verified by experts to ensure quality and… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/HumaniBench.music-ocr-vectors
🎵 Music-OCR-Vectors
A comprehensive dataset of hand-drawn music scores paired with their digital vector and text representations.
📖 About the Dataset
Music-OCR-Vectors is a freely available, open-source dataset designed for Optical Music Recognition (OMR), machine learning, and computer vision research. It provides a bridge between handwritten musical notation and machine-readable formats.
For every hand-drawn music sheet in the dataset, the following ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/Dddixyy/music-ocr-vectors.VLDBench
VLDBench: Evaluating Multimodal Disinformation with Regulatory Alignment
📜 Paper (Preprint)
📄 VLDBench Evaluating Multimodal Disinformation with Regulatory Alignment (arXiv)
Website
Link
Dataset Summary
VLDBench is a multimodal dataset for news disinformation detection, containing text, images, and metadata extracted from various news sources. The dataset includes headline, article text, image descriptions, and images stored as byte arrays… See the full description on the dataset page: https://huggingface.co/datasets/vector-institute/VLDBench.vector-100k
VectorOS Vector 100k SimSat VLM Dataset
VectorOS Vector 100k is a high-fidelity multimodal instruction dataset for fine-tuning vision-language models on geospatial epidemiology tasks. It was built for the VectorOS hackathon project and targets LiquidAI/LFM2.5-VL-450M.
The dataset contains 100,000 chat-style examples derived from 10,000 geospatial chips across 30 AOIs. Every accepted chip has a real SimSat Sentinel-2 true-color view, a real SimSat Sentinel-2 NIR-red-green false-color… See the full description on the dataset page: https://huggingface.co/datasets/Alfaxad/vector-100k.VectorGym
VectorGym: A Multi-Task Benchmark for SVG Code Generation and Manipulation
Dataset Description
VectorGym is a unified corpus for training and evaluating multimodal models on complex vector graphics understanding and manipulation. This dataset provides high-quality, human-annotated data supporting four key SVG tasks:
Sketch-to-SVG: Converting hand-drawn sketches into clean vector graphics
Text-to-SVG: Generating SVG content from natural language descriptions
SVG… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/VectorGym.image-gen-vector-consistencyDataset for upcomming paper "Evaluating Consistency of Image Generation Models with Vector Similarity"
vectors2vibes-discogs-metadata
Vectors2Vibes Discogs Metadata
Metadata for 24.6k tracks derived from Discogs Data and MTG Discogs-VI-YT. No audio files.
Note that earliest release year data is derived from MusicBrainz, as Discogs release year data is sparse and often unreliable.
Quick Facts:This dataset contains 24,689 tracks (release dates spanning from 1890-2026).
The top 5 represented decades are: 1960s (19.39%), 1970s (15.09%), 1980s (14.81%), 1950s (14.96%), and 1990s (14.27%).
The top 5 represented genres… See the full description on the dataset page: https://huggingface.co/datasets/vectors2vibes/vectors2vibes-discogs-metadata.arXiv-AI-papers-multi-vector
Overview
This is a dataset containing individual pages from the top-40 most cited AI papers on arXiv](https://arxiv.org/abs/2412.12121) from the period 2023-01-01 to 2024-09-30.
Only the first 10 pages from each paper is included.
The dataset includes an image of each page as well as a multi-vector embedding using vidore/colqwen2-v1.0.
noto-emoji-vector-512-svg
Dataset Card for "noto-emoji-vector-512-svg"
More Information needed
africa-synth-malaria-vector-surveillance-control-all
Vector Surveillance & Control | Africa (World Health Organization)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help researchers examine… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-malaria-vector-surveillance-control-all.africa-synth-malaria-vector-range-expansion-all
Vector Range Expansion & Disease Emergence | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-malaria-vector-range-expansion-all.vector_lin3art_style_wn-datasettest_demo
Dataset Summary
Placeholder
You can load the dataset via:
import datasets
data = datasets.load_dataset('GEM/wiki_lingua')
The data loader can be found here.
website
None (See Repository)
paper
https://www.aclweb.org/anthology/2020.findings-emnlp.360/
authors
Faisal Ladhak (Columbia University), Esin Durmus (Stanford University), Claire Cardie (Cornell University), Kathleen McKeown (Columbia University)
Dataset Overview
Where to… See the full description on the dataset page: https://huggingface.co/datasets/vector/test_demo.vector_lin3art_style_wnfinal-vector-motifsvector_datavectorized_objectsObject image, Text Description (best fit to SD1.5~)
ibm-hls-burn-vectorizedwordlistvector_wyzvectors2vibes-yt-thumbnails
Vectors2Vibes YT Thumbnails
22.5k YouTube thumbnail JPGs (480x360px, RGB, 90% quality) for Vectors2Vibes Discogs-derived tracks.
Structure
thumbnails/
├── dQw4w9WgXcQ.jpg # YT video_id.jpg
├── abc123def456.jpg
└── ... (22.5k total, ~2GB)
Sister Repos
Repo
Content
Access
Metadata
Discogs/YT metadata
Public
19-4-embeddingsdirectv-zocalos-agosto-5fps_vectors
Dataset Card for "directv-zocalos-agosto-5fps_vectors"
More Information needed
vector-lora-dataset
