datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
handwriting-ocr
Handwriting OCR (Lance Format)
This Lance-formatted version of the Doctor's Handwritten Prescription BD dataset contains 4,680 cropped PNG images of handwritten medicine names from Bangladesh. Each row keeps the original image bytes with the medicine and generic-name labels, plus deterministic search metadata derived from those labels. The dataset contains three source-preserved splits: train, validation, and test.
[!NOTE]
Training note: The same samples appear repeatedly… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/handwriting-ocr.mnist-lance
MNIST (Lance Format)
A Lance-formatted version of the classic MNIST handwritten-digit dataset covering 70,000 28×28 grayscale digits across ten balanced classes. Each row carries inline PNG bytes, the digit label, the human-readable class name, and a cosine-normalized CLIP image embedding, all backed by a bundled IVF_PQ vector index plus scalar indices on the label columns and available directly from the Hub at hf://datasets/lance-format/mnist-lance/data.
Key features… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/mnist-lance.fashion-mnist-lance
Fashion-MNIST (Lance Format)
A Lance-formatted version of Fashion-MNIST covering 70,000 28×28 grayscale clothing images across ten balanced apparel classes. Each row carries inline PNG bytes, the integer label, the human-readable class name, and a cosine-normalized CLIP image embedding, all backed by a bundled IVF_PQ vector index plus scalar indices on the label columns and available directly from the Hub at hf://datasets/lance-format/fashion-mnist-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/fashion-mnist-lance.oxford-pets-lance
Oxford-IIIT Pet (Lance Format)
A Lance-formatted version of the Oxford-IIIT Pet dataset — 7,390 cat and dog photos across 37 breeds — sourced from pcuenq/oxford-pets. Each row carries the inline JPEG bytes, the breed name, a species flag distinguishing cats from dogs, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/oxford-pets-lance/data.
Key features
Inline JPEG bytes in the image column — no sidecar… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/oxford-pets-lance.eurosat-lance
EuroSAT (Lance Format)
A Lance-formatted version of EuroSAT, the canonical Sentinel-2 RGB land-cover benchmark, sourced from blanchon/EuroSAT_RGB. Each row is a single 64×64 RGB tile with its integer class id, the human-readable class name, and a cosine-normalized OpenCLIP image embedding — all stored inline and available directly from the Hub at hf://datasets/lance-format/eurosat-lance/data.
Key features
Inline JPEG bytes in the image column — no sidecar TIF folders… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/eurosat-lance.food101-lance
Food-101 (Lance Format)
A Lance-formatted version of Food-101, the fine-grained dish-classification benchmark of 101,000 photos spread evenly across 101 dish classes, sourced from ethz/food101. Each row carries the inline JPEG bytes, the integer label, the human-readable label_name, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/food101-lance/data.
Key features
Inline JPEG bytes in the image column — no… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/food101-lance.cifar10-lance
CIFAR-10 (Lance Format)
A Lance-formatted version of CIFAR-10 covering 60,000 32×32 RGB images across ten balanced object classes. Each row carries inline PNG bytes, the integer label, the human-readable class name, and a cosine-normalized CLIP image embedding, all backed by a bundled IVF_PQ vector index plus scalar indices on the label columns and available directly from the Hub at hf://datasets/lance-format/cifar10-lance/data.
Key features
Inline PNG bytes in the… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/cifar10-lance.stanford-cars-lance
Stanford Cars (Lance Format)
A Lance-formatted version of the Stanford Cars fine-grained benchmark — 8,144 photographs across 196 make/model/year classes — sourced from Multimodal-Fatima/StanfordCars_train. Each row carries the inline JPEG bytes, the integer class id, a BLIP-generated caption inherited from the source mirror, and a cosine-normalized CLIP image embedding, all available directly from the Hub at hf://datasets/lance-format/stanford-cars-lance/data.
Key… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/stanford-cars-lance.imagenet-1k-val-lance
ImageNet-1k Validation (Lance Format)
A Lance-formatted version of the canonical 50,000-image ImageNet-1k (ILSVRC2012) validation split, sourced from benjamin-paine/imagenet-1k. Each row is one image with its integer class id, a string class name, and a cosine-normalized OpenCLIP image embedding — all stored inline and available directly from the Hub at hf://datasets/lance-format/imagenet-1k-val-lance/data. The 1.28 M ImageNet-1k train split (155 GB) is intentionally out of scope… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/imagenet-1k-val-lance.
