datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
L-Mind
L-Mind: A Multimodal Dataset for Neural-Driven Image Editing
This dataset is part of the NeurIPS 2025 paper: "Neural-Driven Image Editing", which introduces LoongX, a hands-free image editing approach driven by multimodal neurophysiological signals.
📄 Overview
L-Mind is a large-scale multimodal dataset designed to bridge Brain-Computer Interfaces (BCIs) with generative AI. It enables research into accessible, intuitive image editing for individuals with limited motor… See the full description on the dataset page: https://huggingface.co/datasets/Lance1573/L-Mind.encyclopaedia-britannica-lance
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/NationalLibraryOfScotland/encyclopaedia-britannica-lance.encyclopaedia-britannica-lance-test
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test.droidBalitaNLPA Filipino multi-modal language dataset for text+visual tasks. Consists of 351,755 Filipino news articles (w/ associated images) gathered from Filipino news outlets.
Description
Total # of articles: 351,755
80-10-10 split for training, validation, and testing.
Dataset field descriptions:
title - Article title
body - Article body. Separated into paragraphs
image - Article image
website… See the full description on the dataset page: https://huggingface.co/datasets/LanceBunag/BalitaNLP.BDD100K-enricheddocvqa-lance
DocVQA (Lance Format)
A Lance-formatted version of DocVQA, a benchmark for visual question answering over document images such as industry and government scans, multi-page reports, forms, and receipts, redistributed via lmms-lab/DocVQA (DocVQA config). Each row carries the page image as inline JPEG bytes, the question and reference answer span(s), the original DocVQA question-type tags, UCSF Industry Documents Library provenance, and paired CLIP embeddings for the image and the… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/docvqa-lance.textvqa-lance
TextVQA (Lance Format)
A Lance-formatted version of TextVQA — visual question answering where the question requires reading text in the image (street signs, product labels, screen captures) — sourced from lmms-lab/textvqa. Each row carries the image bytes, the question, the 10 reference annotator answers, the OCR tokens detected by the source pre-processing, OpenImages-style scene tags, and paired CLIP image and question embeddings — all available directly from the Hub at… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/textvqa-lance.handwriting-ocr
Handwriting OCR (Lance Format)
This Lance-formatted version of the Doctor's Handwritten Prescription BD dataset contains 4,680 cropped PNG images of handwritten medicine names from Bangladesh. Each row keeps the original image bytes with the medicine and generic-name labels, plus deterministic search metadata derived from those labels. The dataset contains three source-preserved splits: train, validation, and test.
[!NOTE]
Training note: The same samples appear repeatedly… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/handwriting-ocr.spot-the-diffcoco-captions-2017-lance
COCO Captions 2017 (Lance Format)
A Lance-formatted version of the COCO Captions 2017 corpus, redistributed via lmms-lab/COCO-Caption2017. Each row is one image with 5–7 human-written captions, a cosine-normalized CLIP image embedding, and a cosine-normalized CLIP text embedding of the canonical caption — all stored inline and available directly from the Hub at hf://datasets/lance-format/coco-captions-2017-lance/data.
Key features
Inline JPEG bytes in the image… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/coco-captions-2017-lance.coco-detection-2017-lance
COCO 2017 Object Detection (Lance Format)
A Lance-formatted version of the COCO 2017 object detection benchmark, sourced from detection-datasets/coco. Each row is one image with its inline JPEG bytes, the full per-image list of bounding boxes, COCO 80-class category ids and names, per-object areas, an OpenCLIP image embedding, and pre-built indices — all available directly from the Hub at hf://datasets/lance-format/coco-detection-2017-lance/data.
Key features
Inline… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/coco-detection-2017-lance.encyclopaedia-britannica-lance-test2
Encyclopaedia Britannica (1771-1860) - Lance Format
This dataset contains 155,388 digitized pages from the Encyclopaedia Britannica, spanning editions from 1771 to 1860. The data is stored in Lance format for efficient streaming and lazy image loading.
Dataset Details
Total Pages: 155,388
Total Volumes: 195
Format: Lance (columnar format with blob storage for images)
Source: National Library of Scotland (NLS)
License: Public Domain (CC0)
Loading the Dataset… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/encyclopaedia-britannica-lance-test2.vqav2-lance
VQAv2 (Lance Format)
A Lance-formatted version of VQAv2 — open-ended visual question answering on COCO images — sourced from lmms-lab/VQAv2. Each row is one (image, question, 10 annotator answers) triple with paired CLIP image and question embeddings drawn from the same shared space, plus the VQAv2 question_type / answer_type taxonomy and the consensus multiple_choice_answer — all available directly from the Hub at hf://datasets/lance-format/vqav2-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/vqav2-lance.bpl-card-catalog-lance-fulllaion-1m
LAION-Subset (Lance Format)
A Lance-formatted slice of the LAION image-text corpus (~1M rows) with inline JPEG bytes, CLIP image embeddings (img_emb), full metadata, and a pre-built ANN index — all available directly from the Hub at hf://datasets/lance-format/laion-1m/data/train.lance.
Key features
Inline JPEG bytes in the image column — no sidecar files, no image folders.
Pre-computed CLIP image embeddings (img_emb, 768-dim) with a bundled IVF_PQ index for… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/laion-1m.textvqa-lance-colab
TextVQA VLM Fine-Tuning Demo
A small Lance-formatted subset of TextVQA, source via pre-baked operations from this repo and fine-tuning a VLM on the subset. Each row is one visual question-answering example over an image that contains scene text: inline image bytes, a natural-language question, 10 reference answers, OCR tokens, image-class labels, and paired 512-dimensional image/question embeddings are stored together in a Lance table.
The train split also includes precomputed… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/textvqa-lance-colab.mnist-lance
MNIST (Lance Format)
A Lance-formatted version of the classic MNIST handwritten-digit dataset covering 70,000 28×28 grayscale digits across ten balanced classes. Each row carries inline PNG bytes, the digit label, the human-readable class name, and a cosine-normalized CLIP image embedding, all backed by a bundled IVF_PQ vector index plus scalar indices on the label columns and available directly from the Hub at hf://datasets/lance-format/mnist-lance/data.
Key features… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/mnist-lance.gelsight-mini-gelsight-20260617
Dual GelSight Tactile Dataset 20260617
This dataset contains synchronized marker-mask tactile captures from two GelSight-style sensors:
gelsight_mini: GelSight Mini camera stream
gelsight: custom UVC GelSight-style camera stream
The data was collected on 2026-06-17 for real-world finetuning/adaptation of UniForce-style tactile models.
Directory Layout
marker/<sensor>/<indenter>/<frame_id>.jpg
collection_log.txt
The marker folders contain marker mask images… See the full description on the dataset page: https://huggingface.co/datasets/LancetRobotics/gelsight-mini-gelsight-20260617.chartqa-lance
ChartQA (Lance Format)
A Lance-formatted version of ChartQA, a benchmark for question answering over scientific and business charts that demands a mix of logical and visual reasoning, redistributed via lmms-lab/ChartQA. Each row carries the chart image as inline JPEG bytes, the natural-language question and reference answer(s), a question-type tag (human vs augmented), and paired CLIP embeddings for the image and the question — all available directly from the Hub at… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/chartqa-lance.fashion-mnist-lance
Fashion-MNIST (Lance Format)
A Lance-formatted version of Fashion-MNIST covering 70,000 28×28 grayscale clothing images across ten balanced apparel classes. Each row carries inline PNG bytes, the integer label, the human-readable class name, and a cosine-normalized CLIP image embedding, all backed by a bundled IVF_PQ vector index plus scalar indices on the label columns and available directly from the Hub at hf://datasets/lance-format/fashion-mnist-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fashion-mnist-lance.kitti-2d-detection-lance
KITTI 2D Object Detection (Lance Format)
A Lance-formatted version of the KITTI 2D Object Detection benchmark, sourced from nateraw/kitti so no manual signup or download from cvlibs.net is required. Each row is a single driving frame with inline JPEG bytes, the full set of 2D and 3D object annotations stored as parallel per-object lists, plus a cosine-normalized OpenCLIP ViT-B-32 image embedding — all available directly from the Hub at… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/kitti-2d-detection-lance.pascal-voc-2012-segmentation-lance
Pascal VOC 2012 Segmentation (Lance Format)
A Lance-formatted version of the Pascal VOC 2012 semantic segmentation split, sourced from nateraw/pascal-voc-2012. Each row pairs an inline JPEG image with the per-pixel PNG segmentation mask and a cosine-normalized OpenCLIP ViT-B-32 image embedding, so a single columnar table carries both annotation modalities and the features needed to retrieve, curate, and train against them — all available directly from the Hub at… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/pascal-voc-2012-segmentation-lance.oxford-pets-lance
Oxford-IIIT Pet (Lance Format)
A Lance-formatted version of the Oxford-IIIT Pet dataset — 7,390 cat and dog photos across 37 breeds — sourced from pcuenq/oxford-pets. Each row carries the inline JPEG bytes, the breed name, a species flag distinguishing cats from dogs, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/oxford-pets-lance/data.
Key features
Inline JPEG bytes in the image column — no sidecar… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/oxford-pets-lance.eurosat-lance
EuroSAT (Lance Format)
A Lance-formatted version of EuroSAT, the canonical Sentinel-2 RGB land-cover benchmark, sourced from blanchon/EuroSAT_RGB. Each row is a single 64×64 RGB tile with its integer class id, the human-readable class name, and a cosine-normalized OpenCLIP image embedding — all stored inline and available directly from the Hub at hf://datasets/lance-format/eurosat-lance/data.
Key features
Inline JPEG bytes in the image column — no sidecar TIF folders… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/eurosat-lance.dual-gelsight-marker-20260617
Dual GelSight Marker-Mask Dataset 20260617
This dataset contains synchronized marker-mask tactile captures from two GelSight-style sensors:
gelsight_mini: GelSight Mini camera stream
gelsight: custom UVC GelSight-style camera stream
Only marker-mask images are included. Raw RGB camera frames were intentionally omitted from this Hub upload.
Directory Layout
marker/<sensor>/<indenter>/<frame_id>.jpg
Sensors
marker/gelsight_mini: 3280x2464… See the full description on the dataset page: https://huggingface.co/datasets/LancetRobotics/dual-gelsight-marker-20260617.flickr30k-lance
Flickr30k (Lance Format)
A Lance-formatted version of Flickr30k, redistributed via lmms-lab/flickr30k. Each row is one image with 5 human-written captions, a cosine-normalized CLIP image embedding, and a cosine-normalized CLIP text embedding of the canonical caption — all stored inline and available directly from the Hub at hf://datasets/lance-format/flickr30k-lance/data.
Key features
Inline JPEG bytes in the image column — no sidecar files, no image folders.
Paired… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/flickr30k-lance.food101-lance
Food-101 (Lance Format)
A Lance-formatted version of Food-101, the fine-grained dish-classification benchmark of 101,000 photos spread evenly across 101 dish classes, sourced from ethz/food101. Each row carries the inline JPEG bytes, the integer label, the human-readable label_name, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/food101-lance/data.
Key features
Inline JPEG bytes in the image column — no… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/food101-lance.bpl-card-catalog-lanceclevr-change
