datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pd12m
PD12M
This is a curated PD12M dataset for use with the II-Commons project.
Dataset Details
Dataset Description
This dataset comprises a curated Public Domain 12M image collection, refined by filtering for active image links. EXIF data was extracted, and images underwent preprocessing and feature extraction using SigLIP 2. All vector embeddings are normalized 16-bit half-precision vectors optimized for L2 indexing with vectorchord.… See the full description on the dataset page: https://huggingface.co/datasets/Intelligent-Internet/pd12m.genhome3d-1280
GenHome3D-1280
1,280 validated household and spatial-design assets in USDZ format, organized
across 64 categories.
Explore the visual catalog ·
Browse the GitHub repository ·
Download the versioned release ·
Read the generation method
Dataset summary
Assets
1,280
Categories
64
Assets per category
20
Runtime format
USDZ
Units
Meters
Asset license
CC BY 4.0
Technical validation
1,280/1,280 pass
Package validation
1… See the full description on the dataset page: https://huggingface.co/datasets/linxy97/genhome3d-1280.open-hdri-1k
Open HDRI 1K
A consolidated, public-domain (CC0-1.0) collection of 3,491 equirectangular HDR environment maps at 1K resolution, gathered from five free HDRI libraries: Poly Haven, BlenderKit, ambientCG, CGEES and Open HDRI.
Every map is stored as a linear, high-dynamic-range .exr file alongside a tonemapped .jpg preview, with a per-asset metadata row (dimensions, source, author, license, SHA-256 checksum and tags).
Contents
Source
Assets
Author(s)
License… See the full description on the dataset page: https://huggingface.co/datasets/leodriesch/open-hdri-1k.skin-cancer-ham10000-datasetrobocurate-synth100
synth100 — 100 generated clips for validating Pre-Contact Level Filtering
100 episodes drawn (seed 20260824) from the 952-episode multi-object generation set, packaged so
Stage-5 filtering can be run on them without re-deriving anything. Every input the filter needs
travels with the package, in the space it is consumed in.
Read section 1 before using this. The single most important fact about this data is not in the
file layout, and getting it wrong invalidates any score… See the full description on the dataset page: https://huggingface.co/datasets/glory-hyeok/robocurate-synth100.IS110_repo
IS110 ecology — data and analysis catalogue
Full-length transposon systematics + rearrangement + movement analysis
of the IS110 family across 10 bacterial species. All catalogues are
built from a common set of ~92k IS110 element records with confirmed
empty-vs-filled boundaries (Cross_reference_IS pipeline).
Author : Kuang Hu (kh36969@berkeley.edu) — pc_rubinlab, UC Berkeley
Species covered : Escherichia coli, Klebsiella pneumoniae,
Enterobacter hormaechei, Salmonella enterica… See the full description on the dataset page: https://huggingface.co/datasets/hukuang/IS110_repo.conceptual-12m-mbart-50-multilingualllava-15-rlmpq-vlm-eval-results
RL-MPQ VLM Evaluation Artifacts
Complete figures, tables, galleries, and raw benchmark CSVs for the extended VLM evaluation.
Dataset: AvoCahDoe/llava-15-rlmpq-vlm-eval-results
Collections (by base VLM)
RL-MPQ VLM — LLaVA-1.5-13B — HF collection
RL-MPQ VLM — LLaVA-1.5-7B — HF collection
RL-MPQ VLM — LLaVA-Next Mistral-7B — HF collection
RL-MPQ VLM — Qwen2-VL-7B — HF collection
Model repos
RL-MPQ High Fidelity →… See the full description on the dataset page: https://huggingface.co/datasets/AvoCahDoe/llava-15-rlmpq-vlm-eval-results.JL1-CUP-2024-Second-Format
JL1 CUP 2024 — Second-track format for semantic change detection
Bi-temporal 256×256 RGB patches with per-pixel semantic maps at times T1/T2 and a binary change map, aligned with the data split described in the literature for the JL1 cropland change-detection benchmark (Second Track / JL1-Second style layout).
Source
Resource
URL
JL1 Mall contest information
contest page
JL1 data / resources portal
resrepo
Data are provided by the JL1 / Jilin-1 ecosystem… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/JL1-CUP-2024-Second-Format.CAMELYON17
CAMELYON17
1. Tổng quan
CAMELYON17 là dataset mở rộng của CAMELYON16, gồm ảnh WSI hạch bạch huyết canh gác từ 5 trung tâm y tế khác nhau (multi-center), với 1000 WSI (5 slide/bệnh nhân x 200 bệnh nhân). Bài toán chính là phân loại di căn theo 4 mức tại cấp lymph-node (negative/isolated tumor cells/micro-metastases/macro-metastases) và tổng hợp thành pN-stage tại cấp bệnh nhân.
Nguồn dữ liệu: AWS Open Data, s3://camelyon-dataset/CAMELYON17/ (region us-west-2, truy… See the full description on the dataset page: https://huggingface.co/datasets/okbro1234/CAMELYON17.conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
in-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
startup-Investments-analysis
📊 StartUp Investments EDA
1. Background & Objectives
This project explores a comprehensive dataset of startup investments (sourced from Crunchbase) to uncover the primary factors that predict a startup's survival and trajectory in a competitive market.
Through this Exploratory Data Analysis (EDA), we analyze historical funding data, investment rounds, and market categories to determine which variables drive specific company outcomes - namely, whether a business… See the full description on the dataset page: https://huggingface.co/datasets/lia-prop13/startup-Investments-analysis.TACK_Tunnel_Data
TACK Tunnel Data (TTD): A Benchmark Dataset for Deep Learning-Based Defect Detection in Tunnels
Tunnels are essential elements of transportation infrastructure, but are increasingly affected by ageing and deterioration mechanisms such as cracking. Regular inspections are required to ensure their safety, yet traditional manual procedures are time-consuming, subjective, and costly. Recent advances in mobile mapping systems and Deep Learning (DL) enable automated visual… See the full description on the dataset page: https://huggingface.co/datasets/Niumengru-123/TACK_Tunnel_Data.MM-Food-100K
Overview
This project aims to introduce and release a comprehensive food image dataset designed specifically for computer vision tasks, particularly food recognition, classification, and nutritional analysis. We hope this dataset will provide a reliable resource for researchers and developers to advance the field of food AI. By publishing on Hugging Face, we expect to foster community collaboration and accelerate innovation in applications such as smart recipe recommendations… See the full description on the dataset page: https://huggingface.co/datasets/Humanbased-AI/MM-Food-100K.MarineLife-16K
Dataset Card for MarineLife-16K
We introduce MarineLife-16K, a marine-domain video benchmark designed to evaluate the video understanding capabilities of Vision-Language Models (VLMs). MarineLife-16K contains 2,000 video-text pairs and 16,080 video-question-answer pairs across a collection of 2,000 marine videos, including 12,080 multiple-choice questions and 4,000 open-ended questions. The benchmark emphasizes specialized marine knowledge, visual reasoning, temporal… See the full description on the dataset page: https://huggingface.co/datasets/MarineLife-16K/MarineLife-16K.STARCOP_allbands_Train1
STARCOP dataset
STARCOP dataset: Semantic Segmentation of Methane Plumes with Hyperspectral Machine Learning Models 🌈🛰️Authors: Vít Růžička, Gonzalo Mateo-Garcia, Luis Gómez-Chova, Anna Vaughan, Luis Guanter and Andrew Markham
Fast data preview in: dataset_exploration.ipynb Main repository: github/spaceml-org/STARCOP
Task:
Methane is the second most important greenhouse gas contributor to climate change; at the same time its reduction has been denoted as one of the… See the full description on the dataset page: https://huggingface.co/datasets/previtus/STARCOP_allbands_Train1.GLUE3D
GLUE3D: General Language Understanding Evaluation for 3D Point Clouds
Data repository containing all necessary data for the GLUE3D evaluation benchmark.
GLUE3D is a Q&A benchmark for evaluation of 3D-LLMs object understanding capabilities. It is built around 128 richly textured
surfaces spanning creatures, objects, architecture and transport. Each surface is provided as a 50 k-point RGB point cloud,
a 8K-point RGB point cloud, a 512 × 512 RGB rendering, and five RGB-D multiviews.… See the full description on the dataset page: https://huggingface.co/datasets/giorgio-mariani-1/GLUE3D.taiwan-dtm-2025-terrarium-z13
2025 年版全臺灣 20 m DTM — Terrarium z13
這個 Dataset 將內政部公開的 2025 年版全臺灣 20 公尺網格數值地形模型(DTM)轉為 ShadeMap 可直接讀取的 Terrarium RGB XYZ tiles。
官方資料來源:https://data.gov.tw/dataset/176927
原始資料授權:政府資料開放授權條款-第 1 版。本 repository 為衍生格式,請保留官方來源與授權資訊。
內容
terrain/13/{x}/{y}.png:256×256 RGB PNG,XYZ / Web Mercator tile addressing。
tile-index.csv:每張 tile 的區域、有效像素比例、來源高程範圍與 Terrarium 量化誤差。
build-summary.json:建置摘要。
source-manifest.json:原始 ZIP/TIFF SHA-256、解析度、範圍與 CRS 決策。… See the full description on the dataset page: https://huggingface.co/datasets/yhzkiki/taiwan-dtm-2025-terrarium-z13.food17-flaggedgcfairbench-v2
GCFairBench-100 (v3.2) — DOPP-CRAFT-GC experiment images
Interactive gallery: https://nati1221-craft-gc-gallery.static.hf.space/Human evaluation (Wave 2 recruiting): https://huggingface.co/spaces/nati1221/craft-gc-human-eval
2,500 unique text-to-image outputs for the DOPP-CRAFT-GC Springer / PanAfriCon research evaluation.
Field
Value
Prompts
100 (GCFairBench-100: 40 SSA, 15 each other region)
Methods
Base SD, PromptAug, PromptAug-Explicit, FairImagen-GC… See the full description on the dataset page: https://huggingface.co/datasets/nati1221/gcfairbench-v2.Aperture_Lab_Synthetic_Aperture_Sonar_v1
ApertureLab Synthetic SAS Dataset
Version 1.0 (September 2026). Author: Isaac Gerg. Made with
ApertureLab; samples, statistics and the
generation pipeline are described on the
dataset page.
1000 simulated synthetic aperture sonar (SAS) images, each an 80 m along-track
by 200 m range swath from a HISAS 1030-class 100 kHz sonar on a straight
track, beamformed by time-domain back-projection at 2.5 cm pixels and
delivered as dynamic-range-compressed (DRC) TIFF LZW images with COCO… See the full description on the dataset page: https://huggingface.co/datasets/idg101/Aperture_Lab_Synthetic_Aperture_Sonar_v1.test-dataset-v1alistairking_public-company-esg-ratings-dataset
Public Company ESG Ratings Dataset
ESG ratings for over 700 mid / large-cap companies across various industries
Dataset Info
Source: Kaggle
Original Size: 0.04 MB
Kaggle Downloads: 5,844
Files: 1
Files
data.csv
Mirrored from Kaggle
yfcc15mYFCC15m dataset from https://github.com/openai/CLIP/blob/main/data/yfcc100m.md.
The subset is obtained by filtering the original YFCC100m (yfcc100m_dataset.sql) using the photo ids from https://github.com/openai/CLIP/blob/main/data/yfcc100m.md.
The script to rebuild the data from the original YFCC100m is provided at build_yfcc15m.py.
ham10000-skin-lesion-classifier-dataPhone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.conceptual-12m-multilingual-marian-esCLIP-ViT-L-14-336-L20-features
OpenAI/CLIP-ViT-L/14@336 Layer 20 features, CLIP+BLIP labels
Feature activation max visualization of the 4096 Features @ L20
CLIP+BLIP labels (may or may not describe what a neuron truly encodes!)
⚠️ May contain sensitive images, albeit abstract. Use responsibly!
Examples:
E2AM_ConvNeXtV2_CIFAR10
