datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rule-VLN
Rule-VLN Dataset
Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements.
This dataset accompanies the paper:
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.rule34_full
Rule34 Full Dataset
This is the full dataset of rule34.xxx. And all the original images are maintained here.
Information
Images
There are 11336807 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-10 21:23:24 JST.
These are the information of recent 50 images:
id
filename
width
height
mimetype
tags
file_size
file_url
13078768
13078768.jpeg
1024
1024
image/jpeg
1boy 1girls ai_generated ass bubble_butt… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34_full.rule34xyz
Dataset Card for rule34.xyz
Dataset Summary
This dataset contains information about image files from rule34.xyz, a booru-style imageboard. The dataset includes metadata for 590,983 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset metadata is… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34xyz.flippd-depop-verify
Flippd Depop benchmark — public verification subset
Anonymized subset for independently reproducing the Flippd Depop leave-X-out
recommendation benchmark without model weights or the full dataset.
metadata.jsonl — one row per listing: img_id, seller_id (anonymized),
gender, category, color0, brand, price.
embeddings.npz — precomputed vectors (model outputs, not weights) for
every encoder scored in the study: resnet, clip, fc_nohead, fc_head
(the study's private encoder)… See the full description on the dataset page: https://huggingface.co/datasets/rayna-rules/flippd-depop-verify.rule34lol-images-part2
Dataset Card for rule34lol-images-part2
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.image-aesthetic-scores
Rule34.nexus · Licence: Rule34.nexus Derived Dataset Licence 1.0
Rule34.nexus Image Aesthetic Scores
1. Overview
This dataset contains per-image aesthetic predictions for images in the Rule34.nexus corpus.
Predictions were generated using
discus0434/aesthetic-predictor-v2-5. Source images are not
included in this dataset — only opaque post identifiers, the source image's SHA-256 hash,
the post's content type, and the predicted score.… See the full description on the dataset page: https://huggingface.co/datasets/rule34nexus/image-aesthetic-scores.flippd-verify
Flippd benchmark — public verification subset
Anonymized subset for independently reproducing the Flippd leave-X-out
recommendation benchmark without model weights or the full dataset.
metadata.jsonl — one row per listing: img_id, seller_id (anonymized),
gender, category, color0, brand, price.
embeddings.npz — precomputed vectors (model outputs, not weights) for
every encoder scored in the study: resnet, clip, fc_nohead, fc_head
(the study's private encoder), aligned to ids.… See the full description on the dataset page: https://huggingface.co/datasets/rayna-rules/flippd-verify.rule34world
Dataset Card for rule34.world
Dataset Summary
This dataset contains information about image files from rule34.world, a booru-style imageboard. The dataset includes metadata for 580,977 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34world.rule34lol-images-part1
Dataset Card for rule34lol-images-part1
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.rule34-webp-4Mpixel
Rule34 4M Re-encoded Dataset
This is the re-encoded dataset of deepghs/rule34_full. And all the resized images are maintained here.
There are 11336686 images in total. The maximum ID of these images is 13078768. Last updated at 2025-04-13 16:02:32 JST.
How to Painlessly Use This
Use cheesechaser to quickly get images from this repository.
Before using this code, you have to grant the access from this gated repository. And then set your personal HuggingFace token into… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/rule34-webp-4Mpixel.ak47-acoustic-rul-simulated
AK-47 Acoustic Run-to-Failure (RUL) Simulation Dataset
A synthetic Run-to-Failure dataset for Remaining Useful Life (RUL) estimation of an
AK-47's recoil spring from gunshot audio. Because real run-to-failure recordings of a
wearing firearm are practically impossible to collect, this dataset is generated by a
physics-based Digital Twin that takes a small set of real, healthy gunshot recordings and
mathematically simulates the acoustic signature of mechanical wear over thousands… See the full description on the dataset page: https://huggingface.co/datasets/karankhatavkar/ak47-acoustic-rul-simulated.
