datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
food-dataset
Food Dataset
An image classification dataset of food photos organized into 201 categories (folders), with 35,046 images total (~924 MB).
Each top-level folder is a category (e.g. adana kebab, sushi, waffles, tiramisu, ...) containing JPEG images of that food/dish. This follows the standard Hugging Face imagefolder layout, so it loads directly with:
from datasets import load_dataset
ds = load_dataset("webbrain-one/food-dataset")
Structure
<category… See the full description on the dataset page: https://huggingface.co/datasets/webbrain-one/food-dataset.guiact_websingle_test
Dataset Card for GUIAct Web-Single Dataset - Test Set
This is a FiftyOne dataset with 1410 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/guiact_websingle_test")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/guiact_websingle_test.1k_Website_Screenshots_and_Metadata
Dataset Card for 1000 Website Screenshots with Metadata
Dataset Summary
Silatus is sharing, for free, a segment of a dataset that we are using to train a generative AI model for text-to-mockup conversions. This dataset was collected in December 2022 and early January 2023, so it contains very recent data from 1,000 of the world's most popular websites. You can get our larger 10,000 website dataset for free at: https://silatus.com/datasets
This dataset includes:
High-res… See the full description on the dataset page: https://huggingface.co/datasets/silatus/1k_Website_Screenshots_and_Metadata.mind2web_multimodal_test_website
Dataset Card for Multimodal Mind2Web "Cross-Website" Test Split
Note: This dataset is the test split of the Cross-Website dataset introduced in the paper.
This is a FiftyOne dataset with 1019 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/mind2web_multimodal_test_website.phishing-website-screenshots
Phishing Website Screenshots
A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models.
Contents
Label
label
Images
legitimate
0
7,924
phishing
1
446
Total
8,370
Screenshots were captured at a desktop viewport (1920×1080) as PNG images.
Structure
legitimate/<brand>/<page>.png
phishing/<source>/<page>.png
metadata.csv
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/shresthsamyak/phishing-website-screenshots.typeface_dataset
Typeface Dataset
Designed by COLIGNUM/webXOS (x.com/colignum)
Under Development.
github.com/webxos for more info.
Dataset Information
Name: colignum_typeface_dataset_512px_2026-01-22T01-21-36-081Z
Total Characters: 155
Resolution: 512×512 pixels
Format: PNG + CSV/Parquet
Generated: 2026-01-22
Character Sets Included
A-Z Uppercase (26 characters)
a-z Lowercase (26 characters)
0-9 Numbers (10 characters)
Symbols (32 characters):… See the full description on the dataset page: https://huggingface.co/datasets/webxos/typeface_dataset.showui-web-processed
ShowUI-Web Processed
Flattened, normalized, and scenario-split version of showlab/ShowUI-web.
Each row is a single (instruction, UI element) pair with normalized bounding-box coordinates.
Schema
Column
Type
Description
sample_id
string
Unique row identifier ({row}_{element})
screenshot_id
string
Groups elements from the same screenshot
image_relpath
string
Relative path to the screenshot image
scenario
string
Website/domain inferred from the image path… See the full description on the dataset page: https://huggingface.co/datasets/e1879/showui-web-processed.datastrike_BCI
________ _____________________ ______________________________.___ ____ __.___________
\______ \ / _ \__ ___/ _ \ / _____/\__ ___/\______ \ | |/ _|\_ _____/
| | \ / /_\ \| | / /_\ \ \_____ \ | | | _/ | < | __)_
| ` \/ | \ |/ | \/ \ | | | | \ | | \ | \
/_______ /\____|__ /____|\____|__ /_______ / |____| |____|_ /___|____|__ \/_______ /
\/… See the full description on the dataset page: https://huggingface.co/datasets/webxos/datastrike_BCI.imagenet-w21-webp-wds
Dataset Summary
This is a copy of the full Winter21 release of ImageNet in webdataset tar format with WEBP encoded images. This release consists of 19167 classes, 2674 fewer classes than the original 21841 class Fall11 release of the full ImageNet.
The classes were removed due to these concerns: https://www.image-net.org/update-sep-17-2019.php
This is the same contents as https://huggingface.co/datasets/timm/imagenet-w21-wds but encoded in webp at ~56% of the size, shard count… See the full description on the dataset page: https://huggingface.co/datasets/timm/imagenet-w21-webp-wds.webpii
WebPII Dataset
Split
Samples
Train
40,384
Test
4,481
Total
44,865
For more information, see webpii.github.io
underworld_dataset_v2
_ _ _ _______ ___________ _ _ ___________ _ ______
| | | | \ | | _ \ ___| ___ \ | | || _ | ___ \ | | _ \
| | | | \| | | | | |__ | |_/ / | | || | | | |_/ / | | | | |
| | | | . ` | | | | __|| /| |/\| || | | | /| | | | | |
| |_| | |\ | |/ /| |___| |\ \\ /\ /\ \_/ / |\ \| |___| |/ /
\___/\_| \_/___/ \____/\_| \_|\/ \/ \___/\_| \_\_____/___/
Underworld Dataset v2
Generated with WEBXOS UNDERWORLD LANDSCAPE GENERATOR… See the full description on the dataset page: https://huggingface.co/datasets/webxos/underworld_dataset_v2.ghosthunter-RL
Ghost Hunter RLHF Dataset Sample
This dataset sample contains screenshots captured during gameplay of "Ghost Hunter" (8-bit FPS). Each image corresponds to a moment when the player successfully destroyed a ghost with the precision auto-fire. The dataset is intended for reinforcement learning from human feedback (RLHF) tasks, such as training a preference model to distinguish between "good" and "bad" shots.
by webXOS 2026 (github.com/webxos for more info and apps)
The gym is… See the full description on the dataset page: https://huggingface.co/datasets/webxos/ghosthunter-RL.underworld_dataset_v3
_ _ _ _______ ___________ _ _ ___________ _ ______
| | | | \ | | _ \ ___| ___ \ | | || _ | ___ \ | | _ \
| | | | \| | | | | |__ | |_/ / | | || | | | |_/ / | | | | |
| | | | . ` | | | | __|| /| |/\| || | | | /| | | | | |
| |_| | |\ | |/ /| |___| |\ \\ /\ /\ \_/ / |\ \| |___| |/ /
\___/\_| \_/___/ \____/\_| \_|\/ \/ \___/\_| \_\_____/___/
UNDERWORLD Dataset v3
Visualizes the Fast Inverse Square Root (FISR / Quake III)… See the full description on the dataset page: https://huggingface.co/datasets/webxos/underworld_dataset_v3.timelink_dataset
___ ___ ___ ___ ___ ___
/\ \ ___ /\__\ /\ \ /\__\ ___ /\__\ /\__\
\:\ \ /\ \ /::| | /::\ \ /:/ / /\ \ /::| | /:/ /
\:\ \ \:\ \ /:|:| | /:/\:\ \ /:/ / \:\ \ /:|:| | /:/__/
/::\ \ /::\__\ /:/|:|__|__ /::\~\:\ \ /:/ /… See the full description on the dataset page: https://huggingface.co/datasets/webxos/timelink_dataset.danbooru2023-webp-4Mpixel
Danbooru 2023 webp: A space-efficient version of Danbooru 2023
This dataset is a resized/re-encoded version of danbooru2023.
Which removed the non-image/truncated files and resize all of them into smaller size.
This dataset already be updated to latest_id = 7,832,883.
Thx to DeepGHS!
Notice: content of updates folder and deepghs/danbooru_newest-webp-4Mpixel have been merged to 2000~2999.tar, You can ignore all the content in updates folder safely!
Details
This… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-webp-4Mpixel.Web_UI
Web UI Dataset
This dataset contains web pages, their screenshots across different devices, and images extracted from the web pages. Scrolling videos are stored separately in the 'video' folder. It is intended for use in machine learning tasks related to web design, computer vision, and data analysis.
Dataset Summary
Total web pages: 34
Total images: 492
Total screenshots: 102
Total videos: 34
Contents
For each web page, the dataset includes:
URL of the web… See the full description on the dataset page: https://huggingface.co/datasets/GoofyGoof/Web_UI.phishing-website-screenshots
Phishing Website Screenshots
A dataset of 8,370 full-page website screenshots labelled as legitimate or phishing, intended for training and evaluating visual phishing-detection models.
Contents
Label
label
Images
legitimate
0
7,924
phishing
1
446
Total
8,370
Screenshots were captured at a desktop viewport (1920×1080) as PNG images.
Structure
legitimate/<brand>/<page>.png
phishing/<source>/<page>.png
metadata.csv
metadata.csv… See the full description on the dataset page: https://huggingface.co/datasets/gpm123/phishing-website-screenshots.audioform_dataset
AAA UUUUUUUU UUUUUUUUDDDDDDDDDDDDD IIIIIIIIII OOOOOOOOO FFFFFFFFFFFFFFFFFFFFFF OOOOOOOOO RRRRRRRRRRRRRRRRR MMMMMMMM MMMMMMMM
A:::A U::::::U U::::::UD::::::::::::DDD I::::::::I OO:::::::::OO F::::::::::::::::::::F OO:::::::::OO R::::::::::::::::R M:::::::M M:::::::M
A:::::A U::::::U U::::::UD:::::::::::::::DD I::::::::I OO:::::::::::::OO… See the full description on the dataset page: https://huggingface.co/datasets/webxos/audioform_dataset.webXOS-blackhole-synthetic
webXOS_galaxy_synthetic v1.0
_______ ___ _______ _______ ___ _ __ __ _______ ___ _______
| _ || | | _ || || | | | | | | || || | | |
| |_| || | | |_| || || |_| | | |_| || _ || | | ___|
| || | | || || _| | || | | || | | |___
| _ | | |___ | || _|| |_ | || |_| || |___ | ___|
| |_| || || _… See the full description on the dataset page: https://huggingface.co/datasets/webxos/webXOS-blackhole-synthetic.guiact_websingle_test
Dataset Card for GUIAct Web-Single Dataset - Test Set
This is a FiftyOne dataset with 1410 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from fiftyone.utils.huggingface import load_from_hub
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = load_from_hub("Voxel51/guiact_websingle_test")
# Launch the App
session = fo.launch_app(dataset)… See the full description on the dataset page: https://huggingface.co/datasets/ZhuOnR/guiact_websingle_test.mnist-webdataset-png
MNIST WebDataset PNG
The MNIST dataset with samples stored as PNG images and compiled into the WebDataset format.
DALI/JAX Example
The following code shows how this dataset can be loaded into JAX arrays by DALI.
from nvidia.dali import pipeline_def
import nvidia.dali.fn as fn
import nvidia.dali.types as types
from nvidia.dali.plugin.jax import DALIGenericIterator
from nvidia.dali.plugin.base_iterator import LastBatchPolicy
def get_data_iterator(batch_size, dataset_path):… See the full description on the dataset page: https://huggingface.co/datasets/hayden-donnelly/mnist-webdataset-png.
