datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
danbooru2025-metadata
🎨 Danbooru 2025 Metadata
Latest Post ID: 9,158,800
(as of Apr 16, 2025)
📁 About the DatasetThis dataset provides structured metadata for user-submitted images on Danbooru, a large-scale imageboard focused on anime-style artwork.
Scraping began on January 2, 2025, and the data are stored in Parquet format for efficient programmatic access.Compared to earlier versions, this snapshot includes:
More consistent tag history tracking
Better coverage of older or previously… See the full description on the dataset page: https://huggingface.co/datasets/trojblue/danbooru2025-metadata.visual_ai_at_neurips2025
Dataset Card for neurips-2025-vision-papers
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/visual_ai_at_neurips2025")
# Launch the App
session = fo.launch_app(dataset)
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/visual_ai_at_neurips2025.ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.IOAI2025
International Olympiad in Artificial Intelligence (IOAI 2025, Beijing, China)
About IOAI 2025
The 2nd International Olympiad in Artificial Intelligence (IOAI 2025) took place in Beijing, China, from August 2 to 9, 2025, hosted by Beijing National Day School (BNDS) under the patronage of UNESCO.
Contest Rules: Full rules encompassing the Individual, Team, and GAITE contests are available here.
Syllabus: The official syllabus outlining the AI topics contestants should… See the full description on the dataset page: https://huggingface.co/datasets/IOAI-official/IOAI2025.e621-2025
e621-2025
e621-2025 is an update to e621-2024, a large-scale furry image dataset retrieved from e621. It contains only the new images uploaded to e621 after the e621-2024 dataset was prepared. A copy of the metadata for the new images is included in metadata/new_posts.parquet.
This dataset was prepared using the daily database export for 2025-07-26. A complete copy of the database export (including updated metadata for images in the e621-2024 dataset) is included in the metadata… See the full description on the dataset page: https://huggingface.co/datasets/boxingscorpionbagel/e621-2025.ikea-us-products-2025
IKEA US Product Dataset (July 2025)
This dataset is a structured snapshot of ~30,000 IKEA US products, scraped from the official IKEA US website in July 2025.
It contains product metadata (titles, descriptions, categories, materials, care instructions, etc.) and associated product images.
Contents
products-us.jsonl — one JSON object per product with structured fields.
images-us/ — the first "hero" image for each product, downloaded via image_downloader_first.py.… See the full description on the dataset page: https://huggingface.co/datasets/jeffreyszhou/ikea-us-products-2025.CulturalBiases-2025Preprint : [https://arxiv.org/pdf/2505.14729?]
isbi2025-ps3c_224x224visual_ai_at_neurips2025_jina_with_ocr
Dataset Card for harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr")
# Launch the App… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_jina_with_ocr.OmniBenchmark-1K
OmniBenchmark-1K
OmniBenchmark-1K is a challenging benchmark for Class-Incremental Continual Learning designed to evaluate performance on very long task sequences, ranging from 100 to over 300 non-overlapping tasks.
The dataset was introduced in the paper Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts.
GitHub: https://github.com/LMMMEng/CaRE
Paper: Hugging Face | arXiv
Description
OmniBenchmark-1K provides a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/LMMM2025/OmniBenchmark-1K.isbi2025-ps3cLeafScan-CornDefoliation2025
# LeafScan-CornDefoliation2025-V1.0 Dataset
A multi-level dataset for corn leaf defoliation assessments research. Provides corn leaves processed in both video and image form. Paired to the LeafScan research project at https://github.com/KevynAngueira/LeafScan.
Fields: 7 sampled sites across 3 states (IA, IN, OH)
Plants: 18 total plants
Leaves: 149 leaves total, with leaf number 7-21
Media: 1000+ videos & images across healthy, defoliated, and simulated conditions
Zenodo:… See the full description on the dataset page: https://huggingface.co/datasets/KevynAngueira/LeafScan-CornDefoliation2025.2025-24679-image-dataset
Dataset Card for Kaikai vs Georgie Image Dataset
Dataset Description
This dataset was created as part of a course project for 24-679.It supports binary image classification of two student-created characters, Kaikai and Georgie, to explore dataset creation, augmentation, and reproducibility workflows.
Dataset Summary
Binary classification dataset (Kaikai vs Georgie).
Student-created images with augmentations.
Educational purpose only, not intended for… See the full description on the dataset page: https://huggingface.co/datasets/cassieli226/2025-24679-image-dataset.archaeological-sites-caa2025
Archaeological Site Dataset (CAA UK 2025)
Dataset Summary
This dataset provides a comprehensive multi-channel remote sensing dataset for training machine learning models to detect archaeological sites. The dataset combines Sentinel-2 satellite imagery, FABDEM elevation data, and derived spectral indices to create 11-channel representations of 1×1 km grid cells at 10m resolution.
Key Features:
Multi-modal data: 6 spectral bands + 3 spectral indices + 2 terrain features… See the full description on the dataset page: https://huggingface.co/datasets/lldbrett/archaeological-sites-caa2025.ikea-us-products-2025
IKEA US Product Dataset (July 2025)
This dataset is a structured snapshot of ~30,000 IKEA US products, scraped from the official IKEA US website in July 2025.
It contains product metadata (titles, descriptions, categories, materials, care instructions, etc.) and associated product images.
Contents
products-us.jsonl — one JSON object per product with structured fields.
images-us/ — the first "hero" image for each product, downloaded via image_downloader_first.py.… See the full description on the dataset page: https://huggingface.co/datasets/doniariz/ikea-us-products-2025.2025-24679-image-dataset
🏞️ Nature vs. Not Nature Image Dataset
📝 Dataset Summary
This is a student-created image dataset designed for binary image classification. The dataset consists of 31 original photographs, each manually labeled as either depicting a nature scene or not_nature (e.g., man-made objects, indoor scenes).
To facilitate model training, a larger augmented split is provided, expanding the dataset to 310 images through a series of documented, label-preserving transformations. The… See the full description on the dataset page: https://huggingface.co/datasets/zacCMU/2025-24679-image-dataset.2025-24679-image-dataset
Dataset Card for ccm/2025-24679-image-dataset
Dataset Details
Dataset Description
This dataset consists of images labeled as recycling (0) or trash (1). It was created as part of a classroom exercise in supervised learning and data augmentation, with the goal of giving students practice in building and evaluating image classification pipelines.
Curated by: Fall 2025 24-679 course at Carnegie Mellon University
Shared by [optional]: Christopher McComb… See the full description on the dataset page: https://huggingface.co/datasets/ccm/2025-24679-image-dataset.CHEST-XRAY-CPE-OPH2025
CHEST-XRAY-CPE-OPH2025
Description
This dataset is a curated subset of Chest X-Ray Images (Pneumonia) from kaggle.It was specifically prepared for educational purposes in the KMUTT CPE OpenHouse 2025 workshop.
Only selected classes of Thai food are included, and corrupted images were removed to ensure smooth training and evaluation. The dataset provides labeled images of Chest X-Ray, useful for practicing deep learning workflows such as preprocessing, training, and… See the full description on the dataset page: https://huggingface.co/datasets/Thinnaphat/CHEST-XRAY-CPE-OPH2025.2025-24679-HW1-Images-mkarthik
🥄 2025-24679-HW1-Images (Cutlery Classification)
📌 Purpose
This dataset was created for an academic assignment on data collection and augmentation.It is intended to support binary image classification tasks: detecting whether an image contains cutlery or not.
📊 Composition
Total size (original): 32 images
21 containing cutlery
11 without cutlery
Augmented size: ~352 images (32 originals + 320 synthetic variants)
Image size: 224×224 pixels (resized… See the full description on the dataset page: https://huggingface.co/datasets/madhavkarthi/2025-24679-HW1-Images-mkarthik.NNDL_HW5_S2025
Dataset Card for "NNDL_HW5_S2025"
This is a dataset created for neural networks and deep learning course at University of Tehran. The original data can be accessed at https://www.kaggle.com/datasets/emmarex/plantdisease/data
More Information needed
synchro-April2025-cluster-labeled-highMag
IFCB Plankton Labeled (Cluster-Sorted)
This dataset contains labeled images of phytoplankton collected with the Planktivore Imaging System. Images were preprocessed with a zero-padding and resized to the standard size used for ViT_b_16
The dataset was originally constructed by clustering unlabeled ROI images using deep features from a ViT model.Clusters were then saved locally and manually curated into taxonomic labels and higher-order groups.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/patcdaniel/synchro-April2025-cluster-labeled-highMag.visual_ai_at_neurips2025_colmodernvbert
Dataset Card for Voxel51/visual_ai_at_neurips2025
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_colmodernvbert")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_colmodernvbert.2025-24679-image-dataset-Stefanovvisual_ai_at_neurips2025_jina
Dataset Card for Voxel51/visual_ai_at_neurips2025
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_jina")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_jina.flickr-lifeboat-commons-1k-2025
Flickr Commons 1K Collection
Dataset Description
This dataset is a Flickr Data Lifeboat converted to a machine learning ready/ Hugging Face datasets compatible format. Data Lifeboats are digital preservation archives created by the Flickr Foundation to ensure long-term access to meaningful collections of Flickr photos and their rich community metadata.
What is a Data Lifeboat?
Data Lifeboats are self-contained archives designed to preserve not just images… See the full description on the dataset page: https://huggingface.co/datasets/flickr-foundation/flickr-lifeboat-commons-1k-2025.2025-24679-image-dataset
Car Classification Dataset - Original
Dataset Description
This dataset contains 31 original car images collected for binary classification tasks. Images are captured from various angles and in different lighting conditions.
Dataset Summary
Images: 31 high-quality car photographs
Resolution: 224x224 pixels
Format: RGB images (converted from HEIC/PNG)
Labels: Binary classification (sedan vs SUV)
Usage
This dataset is designed for:
Image… See the full description on the dataset page: https://huggingface.co/datasets/Anyuhhh/2025-24679-image-dataset.2025-24679-HW1-images
Dataset Card for Chair Presence Image Dataset
This dataset consists of original and augmented images labeled to indicate whether a chair is present in the scene (1) or not (0). It was built as part of a coursework project on dataset creation and augmentation.
Dataset Details
Dataset Description
The dataset contains 30 original images and 300 augmented images. Each sample is resized to 224x224 pixels.
Labels: Binary (0 = no chair, 1 = chair present).… See the full description on the dataset page: https://huggingface.co/datasets/SebastianAndreu/2025-24679-HW1-images.tw-vqa-2025-reasoning
台灣觀光景點資料集 2025
⚠️ 重要聲明
資料來源與版權
本資料集中的圖片來自 Google 圖片搜尋,僅供學術研究與教育用途。
版權聲明:
圖片版權歸原始版權所有者所有
本資料集不主張對任何圖片擁有版權
圖片僅用於研究目的,不得用於商業用途
使用限制:
✅ 學術研究
✅ 教育用途
✅ 非商業性機器學習模型訓練
❌ 商業用途
❌ 再分發圖片
❌ 侵犯原始版權所有者權益的用途
免責聲明:
使用者有責任確保其使用方式符合相關版權法律。資料集提供者不對因使用本資料集而產生的任何版權糾紛負責。
如果您是圖片的版權所有者並希望移除您的圖片,請聯絡資料集維護者。
資料集描述
這是一個台灣觀光景點圖像資料集,包含約 1000 個台灣景點的圖片及其詳細描述。每個樣本都經過 LLM 生成詳細的視覺描述和問答對。
資料收集方式:
圖片來源:Google 圖片搜尋
景點資訊:公開觀光資訊
LLM 生成內容:使用大型語言模型生成圖片描述和問答
資料集統計
訓練集… See the full description on the dataset page: https://huggingface.co/datasets/treeleaves30760/tw-vqa-2025-reasoning.visual_ai_at_neurips2025_nomic
Dataset Card for Voxel51/visual_ai_at_neurips2025
This is a FiftyOne dataset with 1134 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("harpreetsahota/visual_ai_at_neurips2025_nomic")
# Launch the App
session =… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/visual_ai_at_neurips2025_nomic.
