datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Indonesian-running-photos
Dataset Card for Fotoyu Album Archive
This dataset stores photo and video collections archived from Fotoyu albums using the potoyu-tree-downloader application. It is designed to act as a high-speed Cloudflare-backed CDN for serving static media assets, as well as providing a dataset for image/video classification and machine learning model training.
Dataset Details
Dataset Description
The dataset aggregates scraped photo galleries and video albums… See the full description on the dataset page: https://huggingface.co/datasets/TierKun/Indonesian-running-photos.Co-Spy-Bench
CO-SPY: Combining Semantic and Pixel Features to Detect Synthetic Images by AI (CVPR 2025)
With the rapid advancement of generative AI, it is now possible to synthesize high-quality images in a few seconds. Despite the power of these technologies, they raise significant concerns regarding misuse.
To address this, various synthetic image detectors have been proposed. However, many of them struggle to generalize across diverse generation parameters and emerging generative models.
In… See the full description on the dataset page: https://huggingface.co/datasets/ruojiruoli/Co-Spy-Bench.anemia-survey-dataset
Anemia Detection — Multi-Modal Clinical SEWA Rural Dataset
Organisation: SEWA Rural — Society for Education, Welfare and Action (Rural), Jhagadia, Gujarat, India
Dataset: sewa-rural-care/anemia-survey-dataset
Contact: sewarural@ymail.com
Version: 1.0 — July 2026
Dataset Summary
This dataset supports research into non-invasive, smartphone-based anemia
screening applicable to low-resource and rural healthcare settings. It was
collected by SEWA Rural — a non-profit… See the full description on the dataset page: https://huggingface.co/datasets/sewa-rural-care/anemia-survey-dataset.Rule-VLN
Rule-VLN Dataset
Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements.
This dataset accompanies the paper:
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.Runway_Frames_t2i_human_preferences
Rapidata Frames Preference
This T2I dataset contains roughly 400k human responses from over 82k individual annotators, collected in just ~2 Days using the Rapidata Python API, accessible to anyone and ideal for large scale evaluation.
Evaluating Frames across three categories: preference, coherence, and alignment.
Explore our latest model rankings on our website.
If you get value from this dataset and would like to see more in the future, please consider liking it.… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Runway_Frames_t2i_human_preferences.ImageTime_Benchmark
ImagineTime Benchmark
This dataset repository contains the public benchmark assets for ImagineTime, released with the paper “Can Image Models Imagine Time?”
Paper: arXiv:2606.10620
ImagineTime evaluates whether image generation models can produce ordered 2x2 motion sheets with coherent entities, spatial relations, state transitions, interactions, and task constraints.
Contents
cases/
750 benchmark cases. Each case includes process specs, prompts… See the full description on the dataset page: https://huggingface.co/datasets/Xin-Rui/ImageTime_Benchmark.rule34_full
Rule34 Full Dataset
This is the full dataset of rule34.xxx. And all the original images are maintained here.
Information
Images
There are 11336807 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-10 21:23:24 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_size
file_url
13078768
13078768.jpeg
1024
1024
image/jpeg
1boy 1girls ai_generated ass bubble_butt… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34_full.rule34xyz
Dataset Card for rule34.xyz
Dataset Summary
This dataset contains information about image files from rule34.xyz, a booru-style imageboard. The dataset includes metadata for 590,983 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset metadata is… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34xyz.flippd-depop-verify
Flippd Depop benchmark — public verification subset
Anonymized subset for independently reproducing the Flippd Depop leave-X-out
recommendation benchmark without model weights or the full dataset.
metadata.jsonl — one row per listing: img_id, seller_id (anonymized),
gender, category, color0, brand, price.
embeddings.npz — precomputed vectors (model outputs, not weights) for
every encoder scored in the study: resnet, clip, fc_nohead, fc_head
(the study's private encoder)… See the full description on the dataset page: https://huggingface.co/datasets/rayna-rules/flippd-depop-verify.bilingual-ocr-ru-en-synthetic
Bilingual OCR RU-EN Synthetic Dataset
This synthetic dataset is designed for bilingual text recognition (OCR) and script classification tasks (cyrillic / latin) at the word and short-line level.
Why are numbers, mathematical symbols, and the Greek alphabet included in the generation?
When creating synthetic OCR datasets, including an expanded set of characters (digits, mathematical signs, and Greek letters) is a deliberate step aimed at two main goals:… See the full description on the dataset page: https://huggingface.co/datasets/tehnik-tehnolog/bilingual-ocr-ru-en-synthetic.rule34lol-images-part2
Dataset Card for rule34lol-images-part2
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.yolo-rubber-ducks
Rubber Duck Detection Dataset
Overview
This dataset contains 192 annotated images of rubber ducks, specifically curated for object detection tasks. It was used for experimentation related to the YOLOv8n Rubber Duck Detector model.
NOTE: I DO NOT RECOMMEND USING THIS DATASET AT THIS TIME. There is an open and ongoing discussion around the use of the datasets that were combined for this.See related licensing discussion on the forum
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/brainwavecollective/yolo-rubber-ducks.image-aesthetic-scores
Rule34.nexus · Licence: Rule34.nexus Derived Dataset Licence 1.0
Rule34.nexus Image Aesthetic Scores
1. Overview
This dataset contains per-image aesthetic predictions for images in the Rule34.nexus corpus.
Predictions were generated using
discus0434/aesthetic-predictor-v2-5. Source images are not
included in this dataset — only opaque post identifiers, the source image's SHA-256 hash,
the post's content type, and the predicted score.… See the full description on the dataset page: https://huggingface.co/datasets/rule34nexus/image-aesthetic-scores.flippd-verify
Flippd benchmark — public verification subset
Anonymized subset for independently reproducing the Flippd leave-X-out
recommendation benchmark without model weights or the full dataset.
metadata.jsonl — one row per listing: img_id, seller_id (anonymized),
gender, category, color0, brand, price.
embeddings.npz — precomputed vectors (model outputs, not weights) for
every encoder scored in the study: resnet, clip, fc_nohead, fc_head
(the study's private encoder), aligned to ids.… See the full description on the dataset page: https://huggingface.co/datasets/rayna-rules/flippd-verify.coffee_rust_multispec_classification
Coffee Rust Multispec Classification
A dataset for image classification of Coffee Rust Multispec Classification. The dataset contains 1,120 images across 2 classes: NoRust, Rust.Images per class:
NoRust: 273
Rust: 847
This dataset is indexed on https://project-agml.github.io/ as part of the AgML python library.
Citation
@article{arocatrujillo2025colombian,
title={Colombian coffee tree leaves multispectral images dataset},
author={Aroca-Trujillo, Jorge Luis… See the full description on the dataset page: https://huggingface.co/datasets/Project-AgML/coffee_rust_multispec_classification.yolo-rubber-ducks
Rubber Duck Detection Dataset
Overview
This dataset contains 192 annotated images of rubber ducks, specifically curated for object detection tasks. It was used for experimentation related to the YOLOv8n Rubber Duck Detector model.
NOTE: I DO NOT RECOMMEND USING THIS DATASET AT THIS TIME. There is an open and ongoing discussion around the use of the datasets that were combined for this.See related licensing discussion on the forum
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/phrogzx/yolo-rubber-ducks.rule34world
Dataset Card for rule34.world
Dataset Summary
This dataset contains information about image files from rule34.world, a booru-style imageboard. The dataset includes metadata for 580,977 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34world.hmdb51-pick-run-stand
HMDB51 — Pick / Run / Stand (frames procesados)
Subconjunto procesado del dataset HMDB51 para un caso de uso de
clasificación de productividad de empleados en almacén mediante visión por
computadora: distinguir entre trabajador activo (recogiendo / corriendo)
e inactivo/pausado (parado).
Clases
Clase HMDB51
Etiqueta de negocio
pick
Activo (Pick Up)
run
Activo (Running)
stand
Inactivo/Pausado (Standing)
Estadísticas
492 videos… See the full description on the dataset page: https://huggingface.co/datasets/treborDev/hmdb51-pick-run-stand.rule34lol-images-part1
Dataset Card for rule34lol-images-part1
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.Louis.Vuitton.Product.prices.Russia
Louis Vuitton web scraped data
About the website
The luxury fashion industry in the EMEA region, particularly in Russia, is characterized by a growing demand for high-end products from renowned brands. Louis Vuitton, a global leader in this industry, caters to this escalating demand through their extensive range of luxury clothing, accessories, and luggage. The brand has significantly increased its presence in Russia by leveraging the power of Ecommerce, effectively… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Louis.Vuitton.Product.prices.Russia.nando_dt7clara_dt7MJRecap_Human_filtered
Dataset Card
Introduction
In recent studies, there has been a significant concern regarding data privacy. While one can scrape data online or create one’s own datasets, there is a major problem concerning the privacy of any individuals depicted in a dataset, their consent to be published, as well as dataset copyright.
Because of this, the present dataset aims to focus on fully human-free and copyright-free material based on the existing published datasets created… See the full description on the dataset page: https://huggingface.co/datasets/RunningInTheVoid/MJRecap_Human_filtered.rule34-webp-4Mpixel
Rule34 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/rule34_full. And all the resized images are maintained here.
There are 11336686 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-13 16:02:32 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34-webp-4Mpixel.ak47-acoustic-rul-simulated
AK-47 Acoustic Run-to-Failure (RUL) Simulation Dataset
A synthetic Run-to-Failure dataset for Remaining Useful Life (RUL) estimation of an
AK-47's recoil spring from gunshot audio. Because real run-to-failure recordings of a
wearing firearm are practically impossible to collect, this dataset is generated by a
physics-based Digital Twin that takes a small set of real, healthy gunshot recordings and
mathematically simulates the acoustic signature of mechanical wear over thousands… See the full description on the dataset page: https://huggingface.co/datasets/karankhatavkar/ak47-acoustic-rul-simulated.Rue.La.La.Product.prices.United.States
Rue La La web scraped data
About the website
Rue La La operates in the thriving Ecommerce industry in the United States. This online-driven marketplace sector is significantly fuelled by the continuous innovations in technology, the growing adoption of mobile devices, and the changing consumption patterns of consumers. In particular, Rue La La is recognized for its flash sales model, selling designer apparels, accessories, footwear, and home decor among other things. As… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Rue.La.La.Product.prices.United.States.Net.a.Porter.Product.prices.Russia
Net-a-Porter web scraped data
About the website
The EMEA fashion industry, particularly in Russia, has been experiencing substantial growth in online channels due to increased internet penetration and smartphone usage. A significant player in this advancement is Net-a-Porter. This platform belongs to the luxury ecommerce industry, offering a wide range of premium brands. With the shift towards digital platforms in the shopping behavior of consumers, Net-a-porter is making… See the full description on the dataset page: https://huggingface.co/datasets/DBQ/Net.a.Porter.Product.prices.Russia.unigenbench-assets
unigenbench-assets
Uploaded ZIP of generated faithful candidates for UnigenBench.
central-europe-rural-landscape-dataset-sample
SAMPLE VERSION (Preview Subset)
This repository contains the preview sample (45 images) of the Central European Rural Landscape Dataset (v1.0).
The full dataset (945 images, class-wise ZIP archives, full annotation set) is available separately under a commercial license.
Full version repository:
👉 https://huggingface.co/datasets/batris-data/central-europe-rural-landscape-dataset-full
For licensing inquiries:
batris.sro@gmail.com
Central European Rural Landscape Dataset… See the full description on the dataset page: https://huggingface.co/datasets/batris-data/central-europe-rural-landscape-dataset-sample.DATA-DIFFUSION-BASED-PAPER
DATA-DIFFUSION-BASED-PAPER
Time-lapse fluorescence-microscopy data released with the DATA-DIFFUSION-BASED-PAPER project. The archive contains source image files and accompanying metadata used by the DINO benchmark and downstream analyses.
Contents
Experimental-condition directories containing TIFF frames and ZIP archives.
2,946 TIFF files and 78 ZIP archives.
3,024 metadata files.
Conditions include susceptible mono-culture + DTPA, DTPA controls, and… See the full description on the dataset page: https://huggingface.co/datasets/rubentium/DATA-DIFFUSION-BASED-PAPER.
