datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RuleMaze
RuleMaze Dataset
RuleMaze is a dataset for studying rule-compliant visual spatial planning in Multimodal Large Language Models (MLLMs).
Each task contains a visual maze together with natural-language rules. Models must understand the environment and rules and generate a valid multi-step trajectory.
Dataset Content
RuleMaze contains two types of visual environments:
Regular: grid-based visual maze environments
Quest: adventure-style environments with richer visual… See the full description on the dataset page: https://huggingface.co/datasets/Fish-03/RuleMaze.RULER-BenchRULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
📢 News
[2025-12-19] We have released the Evaluation Code !
[2025-12-03] We have released the Paper, Project Page, and Dataset !
📋 TODOs
Release paper
Release dataset
Release evaluation code
🧩Overview of RULER-Bench
We propose RULER-Bench, a comprehensive benchmark designed to evaluate the… See the full description on the dataset page: https://huggingface.co/datasets/hexmSeeU/RULER-Bench.rule34xyz
Dataset Card for rule34.xyz
Dataset Summary
This dataset contains information about image files from rule34.xyz, a booru-style imageboard. The dataset includes metadata for 590,983 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset metadata is… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34xyz.rule34lol-images-part2
Dataset Card for rule34lol-images-part2
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 77,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files (except the last archive). This is Part 2 of 2 for the complete rule34lol-images dataset. Part 1 can be found here.… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part2.image-aesthetic-scores
Rule34.nexus · Licence: Rule34.nexus Derived Dataset Licence 1.0
Rule34.nexus Image Aesthetic Scores
1. Overview
This dataset contains per-image aesthetic predictions for images in the Rule34.nexus corpus.
Predictions were generated using
discus0434/aesthetic-predictor-v2-5. Source images are not
included in this dataset — only opaque post identifiers, the source image's SHA-256 hash,
the post's content type, and the predicted score.… See the full description on the dataset page: https://huggingface.co/datasets/rule34nexus/image-aesthetic-scores.yeji-bazi-rules
██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗
██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝
██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗
██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║
██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║
╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝
⚡ INTERPRETATION RULEBOOK ⚡
> ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.repro-evidence-grokking-ca-local-rules
Evidence trail — Grokking phase transitions in learning local rules with gradient descent
Full evidence for an automated claim-by-claim audit of
Grokking phase transitions in learning local rules with gradient descent, produced by
Lemma, an AI-scientist pipeline built
for re:AGENT (Founders Inc, Aug 15–16 2026).
Verdict: 5 supported / 0 falsified / 1 inconclusive
of 6 extracted claims. Judge verdict: PASS (5/5).
Claim
Title
Verdict
C1
Critical exponent in 1D… See the full description on the dataset page: https://huggingface.co/datasets/Papajams/repro-evidence-grokking-ca-local-rules.eleusis-calibrated-rules
Eleusis Calibrated Rules — 100-turn reward calibration
A calibrated rule dataset for the single-player Eleusis inductive-reasoning
environment. It extends the 26-rule Hugging Face benchmark with controlled
static, transition, conditional, periodic, chunk, higher-order history, global
history, and compositional rule families.
Source benchmark: Hugging Face Eleusis.
Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11
The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.CodeAnything-1.835M
CodeAnything 1.835M SFT and evaluation release
Gated public release containing the clean 1,835,476-sample training set, the
800-sample/16-domain evaluation set, paper-model predictions and rendered
outputs, and raw per-sample rating records.
Layout
training/
manifest/
all_training_v5.jsonl
all_training_v5.jsonl.idx
all_training_v5.jsonl.true_lengths.u32
shards/<domain>/ exact media/code closure (tar shards)
evaluation/
benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.rule34world
Dataset Card for rule34.world
Dataset Summary
This dataset contains information about image files from rule34.world, a booru-style imageboard. The dataset includes metadata for 580,977 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. The data collection cutoff for this dataset is end of August/early September 2024.
Languages
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34world.rule34lol-images-part1
Dataset Card for rule34lol-images-part1
Dataset Summary
This dataset contains information about image files from rule34.lol, a booru-style imageboard. The dataset includes metadata for 196,000 image files, including URLs, tags, file information, and like counts. The actual image files are stored in zip archives, with each archive containing 1000 image files. This is Part 1 of 2 for the complete rule34lol-images dataset. Part 2 can be found here.
Languages… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/rule34lol-images-part1.RuleSmith-model-comparisonLongWrite-V-Rulerak47-acoustic-rul-simulated
AK-47 Acoustic Run-to-Failure (RUL) Simulation Dataset
A synthetic Run-to-Failure dataset for Remaining Useful Life (RUL) estimation of an
AK-47's recoil spring from gunshot audio. Because real run-to-failure recordings of a
wearing firearm are practically impossible to collect, this dataset is generated by a
physics-based Digital Twin that takes a small set of real, healthy gunshot recordings and
mathematically simulates the acoustic signature of mechanical wear over thousands… See the full description on the dataset page: https://huggingface.co/datasets/karankhatavkar/ak47-acoustic-rul-simulated.colpali_ice_hockey_rulebook_datasetvisual_rules_VQA_dataimnet1k_slide_rule_slipsticksft_ruleimnet1k_rule_rulerSolitaire-Rule-Reasoning-Benchmark
Solitaire Rule Reasoning Benchmark
Overview
The dataset aims to benchmark LLMs on how well they can reason about simple rules. We use Solitaire games as testbeds for creating rule questions. Despite the simplicity of rules, these games have numerous variations, each played very differently, making them ideal for our research questions.
Our findings show that most LLMs perform poorly on this task. However, this performance can be improved through fine-tuning. More… See the full description on the dataset page: https://huggingface.co/datasets/bbateni/Solitaire-Rule-Reasoning-Benchmark.Sthv2_500_3scope
Something-Something V2 (SSV2) 视频预测数据集子集构建文档
🚀 1. 项目目标与任务定义
本项目旨在从庞大的 Something-Something V2 训练集中,构建一个针对指令驱动型视频预测任务的、高质量、小规模的训练和验证子集。
核心任务定义
我们采用经典的视频预测任务定义:给定一个短序列 (F_1 到 F_20) 和一个文本指令,预测序列的下一帧 (F_21)。
元素
描述
索引
图像尺寸
输入序列 (I_Input)
20 张连续帧作为观测输入。
F_01 到 F_20
128 × 128
目标帧 (I_Target)
序列的第 21 帧作为模型预测的真值。
F_21
128 × 128
文本指令 (T)
视频的原始文本标签 (label 字段)。
SSV2 train.json
字符串
🔬 2. 数据集子集精准提取策略(创新点)
Something-Something V2 包含 174 种动作类型和… See the full description on the dataset page: https://huggingface.co/datasets/RuLan03/Sthv2_500_3scope.my_first_lora_v1-dataset
