datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenGameArt-CC0
Dataset Card for OpenGameArt-CC0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons 0 (CC0) license, making them effectively public domain works. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC0.bo_or_not
Dataset Card for bo-dataset
This is a FiftyOne dataset with 169 samples designed for binary classification of Bo (Barack Obama's Portuguese Water Dog) versus other pets.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/bo_or_not")
#… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/bo_or_not.OpenGameArt-OGA-BY-4.0
Dataset Card for OpenGameArt-OGA-BY-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution 4.0 (OGA-BY-4.0) license. The dataset includes various types of game assets such as 2D art, music, sound effects, and associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions and metadata are in English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-4.0.OpenGameArt-CC-BY-SA-3.0
Dataset Card for OpenGameArt-CC-BY-SA-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 3.0 Unported (CC-BY-SA-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-3.0.openclipart
Dataset Card for OpenClipart.org SVG Images
Dataset Summary
This dataset contains 178,604 public domain SVG vector clipart images collected from OpenClipart.org. OpenClipart.org is a community-driven platform where artists share vector clip art explicitly released into the public domain (CC0). The dataset includes the SVG content along with comprehensive metadata such as titles, descriptions, artist names, creation dates, tags, and image URLs. The SVG files in this… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/openclipart.OpenGameArt-OGA-BY-3.0
Dataset Card for OpenGameArt-OGA-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the OpenGameArt Attribution (OGA-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions and metadata are in… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-OGA-BY-3.0.OpenGameArt-CC-BY-3.0
Dataset Card for OpenGameArt-CC-BY-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 3.0 (CC-BY-3.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-3.0.polyvore-outfits
Polyvore Outfits (Refactored Version)
This repository provides a refactored version of the Polyvore Outfits dataset, originally introduced in the paper "Learning Type-Aware Embeddings for Fashion Compatibility" by Mariya I. Vasileva et al.
📌 Overview
The goal of this refactoring is to improve usability and developer experience. While the core data remains identical to the original, the file structure and JSON schemas have been standardized to make it easier to load and… See the full description on the dataset page: https://huggingface.co/datasets/owj0421/polyvore-outfits.OpenGameArt-Mixed-Licenses
Dataset Card for OpenGameArt-Mixed-Licenses
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are available under multiple licenses simultaneously. This dataset includes assets where creators have made their work available under two or more license options. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata, all… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-Mixed-Licenses.cc12m-cleaned
CC12m-cleaned
This dataset builds on two others: The Conceptual Captions 12million dataset, which lead to the LLaVa captioned subset done by
CaptionEmporium
(The latter is the same set, but swaps out the (Conceptual Captions 12million) often-useless alt-text captioning for decent ones_
I have then used the llava captions as a base, and used the detailed descrptions to filter out
images with things like watermarks, artist signatures, etc.
I have also manually thrown out all… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/cc12m-cleaned.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.OpenGameArt-CC-BY-4.0
Dataset Card for OpenGameArt-CC-BY-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-4.0.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.OpenGameArt-CC-BY-SA-4.0
Dataset Card for OpenGameArt-CC-BY-SA-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution-ShareAlike 4.0 International (CC-BY-SA-4.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-CC-BY-SA-4.0.OpenGameArt-CC-BY-4.0
Dataset Card for OpenGameArt-CC-BY-4.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the Creative Commons Attribution 4.0 International (CC-BY-4.0) license. The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily… See the full description on the dataset page: https://huggingface.co/datasets/openlyne/OpenGameArt-CC-BY-4.0.openclipart
Dataset Card for OpenClipart.org SVG Images
Dataset Summary
This dataset contains 178,604 public domain SVG vector clipart images collected from OpenClipart.org. OpenClipart.org is a community-driven platform where artists share vector clip art explicitly released into the public domain (CC0). The dataset includes the SVG content along with comprehensive metadata such as titles, descriptions, artist names, creation dates, tags, and image URLs. The SVG files in this… See the full description on the dataset page: https://huggingface.co/datasets/kawwaaa/openclipart.openm3chest-labels
OpenM3Chest Labels (OM3C)
JSON label files and Series UIDs from the OpenM3Chest dataset, prepared for fine-tuning medical vision-language models such as MedGemma.
Raw imaging data (DICOM) can be downloaded from IDC (Imaging Data Commons) using the Series Instance UIDs provided in unique_keys.txt.
Dataset Summary
OpenM3Chest is a medical multimodal multitask dataset for diagnosing chest abnormalities with a focus on lung cancer screening. The original raw data comes… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels.OpenHotelsSample
OpenHotels Representative Sample
This repository contains a representative sample of OpenHotels for review and inspection. It mirrors the full OpenHotels release structure: image files are stored in tar shards under shards/, and metadata files describe the gallery, non-object query images, object-centric query images, and hotel classes.
The sample is intended for data-quality inspection, not benchmark reporting. Use the full OpenHotels dataset for final evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/imagingforgood/OpenHotelsSample.OmniBenchmark-1K
OmniBenchmark-1K
OmniBenchmark-1K is a challenging benchmark for Class-Incremental Continual Learning designed to evaluate performance on very long task sequences, ranging from 100 to over 300 non-overlapping tasks.
The dataset was introduced in the paper Scaling Continual Learning to 300+ Tasks with Bi-Level Routing Mixture-of-Experts.
GitHub: https://github.com/LMMMEng/CaRE
Paper: Hugging Face | arXiv
Description
OmniBenchmark-1K provides a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/LMMM2025/OmniBenchmark-1K.bo_or_not
Dataset Card for bo-dataset
This is a FiftyOne dataset with 169 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("andandandand/bo_or_not")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/andandandand/bo_or_not.OpenHotels-Updated
OpenHotels-Updated
OpenHotels-Updated is a large-scale hotel image retrieval benchmark built from hotel-room imagery and associated hotel metadata. The dataset is designed for hotel-scale retrieval: given a query image, a system must retrieve the matching hotel from a large gallery containing both true matching classes and many distractor hotel classes.
Dataset Structure
The release contains tar-sharded image files under shards/ and four metadata files:
shards/… See the full description on the dataset page: https://huggingface.co/datasets/imagingforgood/OpenHotels-Updated.solar-flare-hmi-datasplitsThis dataset is intended to be used for training/testing solar flare forecasting models. It contains various data splits (in json format) of SDO/HMI magnetogram images compiled by
Boucheron, L.E., et al., 2023, Sci Data 10, 825, https://doi.org/10.1038/s41597-023-02628-8.
Splits "train", "val", "test" corresponds to the original data splits provided by Boucheron et al., while the other splits are created by downsampling the No-Flare and C flare class to obtain
more balanced splits and… See the full description on the dataset page: https://huggingface.co/datasets/inaf-oact-ai/solar-flare-hmi-datasplits.znanio-others
Dataset Card for Znanio.ru Other Educational Materials
Dataset Summary
This dataset contains 3,092 educational files from the platform [znanio.ru] (https://znanio.ru) that were not categorized in the main groups. Znanio.ru is a resource for teachers, educators, students, and parents that provides a variety of educational content.
Languages
The dataset is primarily in Russian, with potential multilingual content:
Russian (ru): The majority of the content… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/znanio-others.OpenHotels
OpenHotels
OpenHotels is a large-scale hotel image retrieval benchmark built from hotel-room imagery and associated hotel metadata. The dataset is designed for hotel-scale retrieval: given a query image, a system must retrieve the matching hotel from a large gallery containing both true matching classes and many distractor hotel classes.
Dataset Structure
The release contains tar-sharded image files under shards/ and four metadata files:
shards/… See the full description on the dataset page: https://huggingface.co/datasets/imagingforgood/OpenHotels.LiveCompose-outpainted-17k
LiveCompose Outpainted 17k (Tar Shards)
Dataset Summary
This is a tar-sharded release of the LiveCompose outpainted image dataset.
It is prepared for hosting on Hugging Face without hitting the per-directory file limit.
Contents:
training_pairs.jsonl metadata
outpainted_tar/outpainted_part_001.tar ... outpainted_part_006.tar
Why Tar Shards
Hugging Face dataset repos currently enforce a limit on file count per directory.
The original outpainted/ directory… See the full description on the dataset page: https://huggingface.co/datasets/LiveCompose/LiveCompose-outpainted-17k.orion-dataset
Dataset
The ORION dataset is a curated collection of satellite imagery and triage labels used to fine-tune the VLM for orbital image classification. Images are fetched from SimSat's Mapbox API and paired with classification prompts and ground-truth labels.
Dataset Structure
images/
low_ocean_pacific_nemo.png
med_city_chicago.png
high_port_rotterdam.png
...
train_dataset.jsonl
val_dataset.jsonl
test_dataset.jsonl
images/: 512x512 RGB satellite images fetched from… See the full description on the dataset page: https://huggingface.co/datasets/Saransh-cpp/orion-dataset.oceanguard-marine-debris-eval-1000
OceanGuard AI — Marine Debris Evaluation Hold-out (annotations only)
The held-out evaluation split used to report the LoRA adapter delta in the
OceanGuard AI Kaggle Gemma 4 Good Hackathon submission
(Global Resilience track + Unsloth bonus track).
Important — this repository contains only the annotations and metadata.
The 1 000 underwater / coastal images are not redistributed here. They
come from three pre-existing third-party datasets, each with its own
license. Reviewers and… See the full description on the dataset page: https://huggingface.co/datasets/asferrer/oceanguard-marine-debris-eval-1000.
