datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
dataposterlung-tumour-study
Combining graph neural networks and computer vision methods for cell nuclei classification in lung tissue
This is the dataset of the article in the title. It contains 85 patches of 1024x1024 pixels from H&E stained WSIs of 9 different patients. It contains two main classes: tumoural (2) and non tumoural (1). Due to the difficulty of the problem, 153 cells were labelled as uncertain. For technical reasons, we decided to eliminate them in the train and validation set and we… See the full description on the dataset page: https://huggingface.co/datasets/Jerry-Master/lung-tumour-study.project2-agentic-langdata-es
Agentic Language Learning — ES Dataset
Auto-prepared via Data Ingestion & Augmentation pipeline (Functions 1 & 2).
Contents
Clean images: data/train/chunk_*
Augmented images: data/aug/chunk_*
Metadata: metadata/es_clean.csv (+ aug if available)
Each CSV has columns path, text, lang, split.
master_dataset_v2guide_and_masterSloMoBlurTo cite this dataset in a publication, please use:
@misc{mahmud2025deblurringwildrealworldimage,
title={Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos},
author={Syed Mumtahin Mahmud and Mahdi Mohd Hossain Noki and Prothito Shovon Majumder and Abdul Mohaimen Al Radi and Sudipto Das Sukanto and Afia Lubaina and Md. Mosaddek Khan},
year={2025},
eprint={2506.19445},
archivePrefix={arXiv},
primaryClass={cs.CV}… See the full description on the dataset page: https://huggingface.co/datasets/masterda/SloMoBlur.shanghai_master_plan_beirThis is a copy of https://huggingface.co/datasets/jinaai/shanghai_master_plan reformatted into the BEIR format. For any further information like license, please refer to the original dataset.
Disclaimer
This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at)… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/shanghai_master_plan_beir.haotian_data-GPS-AR-Lopti-master
Do Not Let Low-Probability Tokens Over-Dominate in RL for LLMs
Authors: Zhihe Yang$^{1}$, Xufang Luo$^{2*}$, Zilong Wang$^2$, Dongqi Han$^2$, Zhiyuan He$^2$, Dongsheng Li$^2$, Yunjian Xu$^{1*}$,
($^*$ for corresponding authors)
The Chinese University of Hong Kong, Hong Kong SAR, China
Microsoft Research Asia, Shanghai, China
Introduction
In this study, we identify a critical yet underexplored issue in RL training: low-probability tokens disproportionately… See the full description on the dataset page: https://huggingface.co/datasets/happynew111/haotian_data-GPS-AR-Lopti-master.lane_master
Dataset Card for "lane_master"
More Information needed
third-eye_master_datasettext-2-video-human-preferences-kling-v2.1-master
Rapidata Video Generation Kling v2.1 Master Human Preference
In this dataset, ~60k human responses from ~20k human annotators were collected to evaluate Kling v2.1 Master video generation model on our benchmark. This dataset was collected in roughtly 30 min using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-kling-v2.1-master.marvel-masterpieces-with-3dmesh
Dataset Card for reconstructions
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces-with-3dmesh.flooringmanga-colorization-masterTypeCoffee_16x16OmnigraphCanvas-Dataset
📚 OmnigraphCanvas Dataset
The OmnigraphCanvas Dataset is a massive, high-fidelity collection of raw manga pages compiled specifically for multimodal machine learning, panel segmentation, temporal graph reasoning, and character re-identification.
OmnigraphCanvas provides tens of thousands of high-resolution pages across diverse art styles, making it the perfect foundation for training robust Vision-Language Models (VLMs) and OmniGraph architectures.
📖 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/mastersubhajit/OmnigraphCanvas-Dataset.Master-Reservoir
Master Reservoir
A merged, standardized aerial object-detection dataset combining DetFly and AOD4 into a single YOLO-format collection with a unified class mapping.
Motivation
Master Reservoir is a unified aerial object detection dataset created by merging multiple public datasets into a common YOLO annotation format.
The objective is to provide a single, standardized dataset for training and evaluating drone detection models while reducing inconsistencies in… See the full description on the dataset page: https://huggingface.co/datasets/AkshayUmesh/Master-Reservoir.lane_master2
Dataset Card for "lane_master2"
More Information needed
marvel-masterpieces
Dataset Card for marvel_masterpieces
Wait! Before you go, ❤️ the dataset! Let's get this trending!
This is a FiftyOne dataset with 255 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/marvel-masterpieces")
#… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/marvel-masterpieces.TypeCoffee_32x32STRAY-DOGS-MASTER-VAULTrutube-masterit is the master repository for holding all the scraped 30M Rutube data.
project2-agentic-langdata-en
Agentic Language Learning — EN Dataset
Auto-prepared via Data Ingestion & Augmentation pipeline (Functions 1 & 2).
Contents
Clean images: data/train/chunk_*
Augmented images: data/aug/chunk_*
Metadata: metadata/en_clean.csv (+ aug if available)
Each CSV has columns path, text, lang, split.
TypeCoffee_128x128TypeCoffee_8x8shanghai_master_plan
Shanghai Master Plan Document Retrieval
The master plan document is taken from here. Each page is annotated with a human written query. The text_description column contains OCR text extracted from the images using EasyOCR.
Citation
@manual{shanghai_masterplan_2018,
title = {Shanghai Master Plan 2017--2035: Striving for the Excellent Global City},
author = {{Shanghai Municipal People’s Government Urban Planning and Land Resource Administration Bureau}}… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/shanghai_master_plan.osworld_tasks_filesMaSTr1325_512x384
MaSTr1325
A Maritime Semantic Segmentation Training Dataset for Small‐Sized Coastal USVs
Overview
MaSTr1325 (Maritime Semantic Segmentation Training Dataset) is a large-scale collection of real-world images captured by an unmanned surface vehicle (USV) over a two-year span in the Gulf of Koper (Slovenia). It was specifically designed to advance obstacle-detection and segmentation methods in small-sized coastal USVs. All frames are per-pixel annotated into three main… See the full description on the dataset page: https://huggingface.co/datasets/Wilbur1240/MaSTr1325_512x384.2025-24679-image-dataset-Stefanov
