datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tcga-wsi-uni2h-features
TCGA WSI UNI2H Features
Dataset Summary
This dataset provides tile-level UNI2-h embeddings extracted from TCGA whole-slide images (WSIs) using a reproducible, auditable pipeline designed for computational pathology research.
Data is organized by project (for example TCGA-HNSC) and currently exposes:
features/ containing H5 feature files with tile-level embeddings
vis/ containing overlay images for quality inspection and pipeline verification
[!IMPORTANT]
Unlike the… See the full description on the dataset page: https://huggingface.co/datasets/W8Yi/tcga-wsi-uni2h-features.turkey-all-universitiesCertainly! Here’s the dataset description in Markdown format:
All Universities in Turkey Dataset
Description
This dataset contains detailed information about various universities. Each record represents a single university and includes attributes such as the university's name, type, city, website, address, logo URL, and a button for accessing additional details. This data is typically extracted from a web page listing universities.
Fields
1. id… See the full description on the dataset page: https://huggingface.co/datasets/h8st6ptv/turkey-all-universities.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.MARS-Hyperspectral-EnMAP-PRISMA-v2025
MARS-Hyperspectral dataset (EnMAP and PRISMA) - version v2025
Updated version of the dataset. More information will be added soon.
university-1652
University-1652: Drone-based Geo-localization Benchmark 🚁
University-1652 is a multi-view dataset for drone-based geo-localization, annotating 1652 buildings across 72 universities (ACM Multimedia 2020, paper). Cited in 500+ papers, it supports Drone → Satellite localization and Satellite → Drone navigation.
🔗 Official code & baseline: layumi/University1652-Baseline
· Leaderboard: State-of-the-art results
🔐 Access
This dataset is gated: click "Request… See the full description on the dataset page: https://huggingface.co/datasets/layumi/university-1652.Unsafe2SafeMARS-Hyperspectral
MARS-Hyperspectral dataset
This repository contains the data for the work publicly presented at:
[ArXiv preprint] Růžička, Mateo-García, Irakulis-Loitxate et al., Operational machine learning for remote spectroscopic detection of CH4 point sources, arXiv preprint arXiv:2511.07719 (2025).
[Poster] Růžička, Mateo-García, Irakulis-Loitxate et al., Machine Learning Models for Multi-sensor Detection of Methane Leaks in Hyperspectral Data, June. 23-27, 2025, ESA Living Planet Symposium… See the full description on the dataset page: https://huggingface.co/datasets/UNEP-IMEO/MARS-Hyperspectral.UnicEdit-10M
CVPR 2026 | UnicEdit-10M: Large-scale Image Editing Dataset
🔗 Quick Links
📄 Paper: UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
💻 Code: GitHub - WeChatCV/UnicBench
🌐 Project Page: UnicEdit-10M
🤗 Benchmark: UnicBench
🌟 Support Us: If you find this dataset or our work useful, please verify it by giving us a star on GitHub! Your support encourages us to keep open-sourcing… See the full description on the dataset page: https://huggingface.co/datasets/xiaotanhua/UnicEdit-10M.universal-lesion-segmentation
Universal Lesion Segmentation Datasets
A collection of public medical imaging datasets for lesion segmentation in CT scans. These are the datasets exactly as downloaded from their original sources.
Datasets
This repository contains the following datasets:
CECT - Liver (primary). Luo J, Wang X, Zhang Y, et al. Comprehensive multi-phase three-dimensional contrast-enhanced CT imaging dataset for primary liver cancer. Scientific Data. 2025;12(1):768.… See the full description on the dataset page: https://huggingface.co/datasets/nielsRocholl/universal-lesion-segmentation.UniML3D
UniML3D
UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons — Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects — brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/UniML3D.community-suspicious-samples
This dataset is part of the ULE-CIBERLAB Project: Transfer of knowledge in cybersecurity for the country's business fabric, funded by the European Union NextGeneration-EU, Recovery, Transformation and Resilience Plan, through INCIBE.
MALWARE-SAMPLES DATASET
Disclaimer: This repository may contain real samples of malware that can be executed (.exe) and artifacts related with their execution in CAPEv2 sandbox (JSON/HTML reports, screenshots, dropped files). DO NOT execute any of… See the full description on the dataset page: https://huggingface.co/datasets/unileon-robotics/community-suspicious-samples.dragon
Dataset Card for DRAGON
🧾 ArXiv Preprint
DRAGON is a large-scale Dataset of Realistic imAges Generated by diffusiON models.
The dataset includes a total of 2.5 million training images and 100,000 test images generated using 25 diffusion models, spanning both recent advancements and older, well-established architectures.
Dataset Details
Dataset Description
The remarkable ease of use of diffusion models for image generation has led to a proliferation of… See the full description on the dataset page: https://huggingface.co/datasets/lesc-unifi/dragon.OmniScience
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
🚀 2026-05-17: This work was accepted by the KDD 2026 Dataset & Benchmark Track. 🚀 2026-05-01: The OmniScience dataset surpassed 20,000 average monthly downloads. 🚀 2026-01-21: The OmniScience dataset ranked Top 8 on Hugging Face Datasets Trending (Top 1 on Image Caption Filed). 🚀 2026-01-17: The OmniScience dataset surpassed 5,000 downloads within 5 days of its release. 🚀 2026-01-12:… See the full description on the dataset page: https://huggingface.co/datasets/UniParser/OmniScience.Sea-Undistort
Dataset Card for Sea-Undistort
Sea-Undistort is a synthetic dataset for through-water image restoration in high-resolution airborne bathymetry. It contains 1,200 scenes with four 512×512 RGB images per scene: (1) ground/no water, (2) undistorted/no waves, (3) no sunglint, (4) distorted (all effects). Each scene comes with structured per-image metadata describing camera, water, sky/illumination, and seafloor parameters. Images were procedurally rendered in Blender to emulate… See the full description on the dataset page: https://huggingface.co/datasets/maxkromer/Sea-Undistort.Kairos_AFEIThe AFEI Corpus: Holonic Scarcity Dynamics & The Shadow Lineage
By Sebastiaan van der Heide (Unityinsight / Kairos_AFEI)
===
The AFEI Corpus and the AFEI Research Methodology are meant to be used to analyze societal and institutional dynamics, with a particular focus on institutional iatrogenesis, epistemic entrapment and epistemic entrainment.
===
AFEI Origins:
The following commits contain descriptions and files which show the entire creation and operational blueprint of the AFEI Methodology… See the full description on the dataset page: https://huggingface.co/datasets/Unityinsight/Kairos_AFEI.Uni-GUI-Desktop-1
Uni-GUI-Desktop-1
A large-scale desktop GUI agent trajectory dataset, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,685
Total Steps
~36K
Platform
Desktop (1920x1080)
Applications
10 categories
Coordinate System
Normalized to [0, 999]
Applications
App
Description
chrome
Web browsing tasks
gimp… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-Desktop-1.FUSU-Fine_grained_Urban_Semantic_Understanding
About:
FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:
Bi-temporal high-resolution satellite RGB images with fine-grained annotations.
Monthly revisited Sentinel-2 and Sentinel-1 images.
Details:
1.… See the full description on the dataset page: https://huggingface.co/datasets/sp-juni/FUSU-Fine_grained_Urban_Semantic_Understanding.UniDoc-Bench
UNIDOC-BENCH Dataset
A unified benchmark for document-centric multimodal retrieval-augmented generation (MM-RAG).
Dataset Description
UNIDOC-BENCH is the first large-scale, realistic benchmark for multimodal retrieval-augmented generation (MM-RAG) and Visual Question Answering (VQA) built from 70,000 real-world PDF pages across eight domains. The dataset extracts and links evidence from text, tables, and figures, then generates 1,700+ multimodal QA pairs spanning… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/UniDoc-Bench.UnlearnCanvas
Dataset Card for UnlearnCanvas
This dataset card introduces "UnlearnCanvas", a high-resolution stylized image dataset for benchmarking generative modeling tasks, in particular for machine unlearning in diffusion models. Developed to address the societal concerns arising from diffusion models, such as harmful content generation, copyright disputes, and the perpetuation of stereotypes and biases, UnlearnCanvas aims at facilitating the evaluation and improvement of machine unlearning… See the full description on the dataset page: https://huggingface.co/datasets/OPTML-Group/UnlearnCanvas.IMAGE_UNDERSTANDINGA key question for understanding multimodal performance is analyzing the ability for a model to have basic
vs. detailed understanding of images. These capabilities are needed for models to be used in
real-world tasks, such as an assistant in the physical world. While there are many dataset for object detection
and recognition, there are few that test spatial reasoning and other more targeted task such as visual prompting.
The datasets that do exist are static and publicly available, thus… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/IMAGE_UNDERSTANDING.MARS-S2L
MARS-S2L dataset
This repository contains the MARS-S2L dataset to be released with the publication:
Mateo-Garcia, Allen, Irakulis-Loitxate et al., Artificial intelligence for methane detection: from continuous monitoring to verified mitigation, arXiv:2511.21777
The MARS-S2L dataset contains approximately 87,000 pairs of Sentinel-2 and Landsat images with more than 5,600 manually verified plumes.
It requires approximately 100GB of hard-disk storage.
Top emitter locations and… See the full description on the dataset page: https://huggingface.co/datasets/UNEP-IMEO/MARS-S2L.reef-guidance-system
Dataset Card for Reef Guidance System
This dataset provides imagery used for training and evaluation of models in the Reef Guidance System. All imagery was collected by the Australian Institute of Marine Science using the ReefScan™ Transom Marine Monitoring System.
If you use this dataset in your work, please cite the associated paper: AI-driven dispensing of coral reseeding devices for broad-scale restoration of the Great Barrier Reef (citations provided at bottom of this… See the full description on the dataset page: https://huggingface.co/datasets/QCR-Underwater-Perception/reef-guidance-system.Uni-GUI-OpenMobile
Uni-GUI-OpenMobile
A mobile GUI agent trajectory dataset collected on open-source Android applications via AndroidWorld, used as part of the training data for UI-MOPD (Multi-platform On-Policy Distillation for Continual GUI Agent Learning).
Dataset Statistics
Metric
Value
Trajectories
2,640
Total Steps
~25.9K
Platform
Android Mobile (1080x2400)
Applications
19 open-source apps
Coordinate System
Normalized to [0, 1000]… See the full description on the dataset page: https://huggingface.co/datasets/UI-MOPD/Uni-GUI-OpenMobile.UniVBench
UniVBench
UniVBench is a unified benchmark for video generation and video editing tasks, covering text-guided video editing, reference-guided video generation, captioning, and multimodal evaluation scenarios.
Dataset Download
To download the whole UniVBench dataset, run the following command in your terminal:
from the code here: https://github.com/JianhuiWei7/UniVBench/blob/main/download.py
python ./download.py
UniVBench Benchmark Directory Structure
Here is… See the full description on the dataset page: https://huggingface.co/datasets/JianhuiWei/UniVBench.kitchen-workspace-understanding-safe-manipulation
Kitchen Workspace Understanding & Safe Manipulation
Generated by datapack-import.ts
This dataset mirrors public data-pack render outputs from Physicl.
Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/kitchen-workspace-understanding-safe-manipulation.LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR
Vero-2.5M-unfiltered
Vero-2.5M-unfiltered
[!Note]
This repository contains the full unfiltered dataset used to construct Vero-600k and Vero-1.6M, before question and answer filtering.
Note that task categories are not balanced in this dataset.
Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with vision-language models. This repository contains the Vero-2.5M-unfiltered dataset, a curation of 2.5M reinforcement learning samples… See the full description on the dataset page: https://huggingface.co/datasets/zlab-princeton/Vero-2.5M-unfiltered.gta-data-files-universalUNS-SSDS-2025Dataset from UNS SSDS 2025 Competition
omni6d-test-unseen
Dataset Card for Omni6D (test_unseen + test + Real subsets)
Installation
pip install -U fiftyone
Usage
import fiftyone as fo
from huggingface_hub import snapshot_download
# Download the dataset snapshot to the current working directory
snapshot_download(
repo_id="<username>/omni6d-test-unseen",
local_dir=".",
repo_type="dataset",
)
# Load dataset from current directory using FiftyOne's native format
dataset = fo.Dataset.from_dir(… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/omni6d-test-unseen.
