datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
military_vehicles
Citation
If you use this dataset, please cite the following paper:
@article{kricheli2024error,
title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge},
author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha and Ozgur, Spencer and Shakarian, Paulo},
journal={arXiv preprint arXiv:2407.15192},
year={2024}
}
Leiniao_Datasetmdsarfdetr-segmentation-leibniz-dataset
Dataset Card for Leibniz's Manuscripts (Instance Segmentation Dataset)
This dataset comprises instance segmentation annotations in raw COCO format, used to train an RF-DETR-Seg-nano model for the automatic recognition of textual, graphical, and mathematical expression zones within the manuscripts of the philosopher and mathematician Gottfried Wilhelm Leibniz (17th-early 18th c.).
Dataset Details
Uses
Direct Use
This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/rfdetr-segmentation-leibniz-dataset.TIGeR-Bench
Paper: https://arxiv.org/abs/2406.05814
ParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/matthew-l-leidos/ParseBench.htr_leibniz_dataset_v1
Dataset Card for Leibniz's Manuscripts (HTR Ground Truth)
This dataset is composed of trascribed and manually corrected folios of Gottfried Wilhelm Leibniz's manuscripts, together with automatically aligned ground truth in order to train or fine-tune Handwritten Text Recognition (HTR) models.
Dataset Details
This ground truth was produced to fine-tune existing HTR models for the recognition of Leibniz's handwriting, with the aim of assisting scholars in the… See the full description on the dataset page: https://huggingface.co/datasets/DenisaBumba/htr_leibniz_dataset_v1.MSE-Bench
MSE-Bench: A Benchmark for Multi-turn Session Image Editing
Introduction
MSE-Bench (Multi-turn Session image Editing Benchmark)
is a benchmark designed to evaluate multi-turn image editing systems under realistic editing workflows. Given a source image and a series of editing instructions, the goal is for a model to apply these edits cumulatively to produce a final image that reflects all the requested changes.
MSE-Bench consists of 100 test instances, each representing… See the full description on the dataset page: https://huggingface.co/datasets/leigangqu/MSE-Bench.Copiale_Lines
Copiale Lines
Copiale Lines is a line-level image-to-text dataset for historical cipher decipherment. It contains cropped line images from the Copiale manuscript paired with plaintext ground truth.
This dataset was presented in the paper Learning to Decipher from Pixels -- A Case Study of Copiale (HistoCrypt 2026).
Code: https://github.com/leitro/Decipher-from-Pixels-Copiale
Dataset Structure
The dataset is split into:
train: 1,269 samples
valid: 175 samples… See the full description on the dataset page: https://huggingface.co/datasets/leitro/Copiale_Lines.MSE-Bench-resultsXieNet
Dataset Card for XieNet
This is the repaired version of GAPartNet dataset, which we use as the simulation dataset for Vi-TacMan.
Description
We identified numerous object meshes in the original dataset that lack proper cap geometry, so we manually repaired these meshes to ensure completeness. The following images (object id: 47296) exemplify the type of geometric defects found and our corrections:
GAPartNet (Original)… See the full description on the dataset page: https://huggingface.co/datasets/Leiyao-Cui/XieNet.PPE_v10validppe24679-hw1-image-register
24-679 HW1 (Fall 2026): Business Message Register Images
leixiang25/24679-hw1-image-register
Digitally rendered screenshot-style images of short, fictional business messages, labeled by register.
1 = formal (high-context business register); 0 = casual (low-context register). Created by Lei Xiang
for 24-679 Homework 1 at Carnegie Mellon University. The dataset is related to Context, a cross-cultural
deal interpreter for Western operators working with Japanese and Chinese… See the full description on the dataset page: https://huggingface.co/datasets/leixiang25/24679-hw1-image-register.leicester_loaded_annotations_binary
Dataset Card for "leicester_loaded_annotations_binary"
More Information needed
flickr30k-qwen3vl-baseline
Flickr30k Qwen3-VL Baseline Captions (Test Split)
This dataset is based on the Mozilla/flickr30k-transformed-captions-gpt4o test split and contains 1,000 images from the original Flickr30k dataset.It includes both the original metadata and newly generated baseline captions produced using the Qwen3-VL-2B-Instruct vision-language model.
Contents
Each entry includes:
image — the original Flickr30k image
alt_text — GPT-4o transformed caption from Mozilla's de-biasing… See the full description on the dataset page: https://huggingface.co/datasets/leinms/flickr30k-qwen3vl-baseline.leicester_loaded_annotations
Dataset Card for "leicester_loaded_annotations"
More Information needed
ImageNet50
ImageNet50 Dataset
This repository contains the dataset for the paper "Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge".
Citation
If you use this dataset in your research, please cite the following paper:
@article{kricheli2024error,
title={Error Detection and Constraint Recovery in Hierarchical Multi-Label Classification without Prior Knowledge},
author={Kricheli, Joshua Shay and Vo, Khoa and Datta, Aniruddha… See the full description on the dataset page: https://huggingface.co/datasets/leibnitz-lab/ImageNet50.flickr30k-qwen3vl-baseline-improved_prompt
Flickr30k Qwen3-VL Few-Shot Styled Captions (Test Split)
This dataset is derived from the Mozilla/flickr30k-transformed-captions-gpt4o test subset and contains 1,000 images.It preserves the original metadata and adds new few-shot–guided baseline captions generated with the Qwen3-VL-2B-Instruct model.
The goal of this dataset is to provide a consistent, controlled captioning style enforced through three-shot visual prompting.
📌 Dataset Contents
Each sample contains:… See the full description on the dataset page: https://huggingface.co/datasets/leinms/flickr30k-qwen3vl-baseline-improved_prompt.rlflickr30k-qwen3vl-baseline-evaluation-by-gpt-4o-2024-08-06flickr30k-qwen3-vl-2b-sft-trl-with-judgeknot_dvrkleishfinevision-mcq-v4imgbedllava-next-mcq-v3zapovednik_combined_v2llava-next-mcq-v2-25kfashion_image_caption-100-v2
Dataset Card for "fashion_image_caption-100-v2"
More Information needed
