datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Rule-VLN
Rule-VLN Dataset
Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements.
This dataset accompanies the paper:
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.BenchMAX_Rule-based
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios.
We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.RULER-500
RULER Dataset (Qwen3 Tokenizer)
Dataset generated from RULER for long context evaluation (qwen3 tokenizer).
This dataset is used in the paper Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE.
The official code is available at github.com/jet-ai-projects/jet-long.
RULER_50
RULER_50 Official-Code Qwen3 Subset
This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks.
It was generated from the official NVIDIA/RULER GitHub code, not from a
third-party pre-generated mirror.
Official generation source:
Repository: https://github.com/NVIDIA/RULER
Branch: main
Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7
Generation entrypoint: scripts/data/prepare.py
Benchmark config: scripts/synthetic.yaml
Generation settings:
tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.DHSA_RULER
RULER Evaluation Data
This dataset contains pre-generated JSONL files for the RULER long-context evaluation benchmark, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight). RULER is designed to evaluate effective context length and long-context behavior beyond simple retrieval, covering retrieval, multi-hop tracing, aggregation, and question answering style tasks.
The files are organized by target… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_RULER.ruler-niah-multilength-eval-benchmark
📌 Fixed Multi-Length RULER NIAH Benchmark (1K, 2K, 4K, 8K)
Deterministic synthetic Needle-In-A-Haystack (NIAH) benchmark splits for reproducible long-context evaluation.
Dataset Specifications:
Tasks (4):
niah_single_1: Repeat haystack, single word needle, number value.
niah_single_2: Essay haystack, single word needle, number value.
niah_single_3: Essay haystack, single word needle, UUID value.
niah_multikey_1: Essay haystack, 4 keys needle, number value.… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ruler-niah-multilength-eval-benchmark.rulm
Dataset for training Russian language models
Overall: 75G
Scripts: https://github.com/IlyaGusev/rulm/tree/master/data_processing
Website
Char count (M)
Word count (M)
pikabu
14938
2161
lenta
1008
135
stihi
2994
393
stackoverflow
1073
228
habr
5112
753
taiga_fontanka
419
55
librusec
10149
1573
buriy
2646
352
ods_tass
1908
255
wiki
3473
469
math987
177
sigmaforge-detection-rules
SigmaForge Detection Rules
SigmaForge is a structured, operational dataset for building and evaluating
systems that generate, validate, and translate Sigma
detection rules. Sigma is a vendor-agnostic YAML format that describes
detection logic so it can be shared across SIEM platforms.
The dataset is derived from the open-source SigmaHQ
rule corpus. Every rule is normalized and enriched with:
MITRE ATT&CK technique and tactic mappings extracted from rule tags.
Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.Rule2DRC
Rule2DRC
Rule2DRC is a benchmark for generating KLayout DRC Ruby runsets from natural-language design-rule specifications.
Paper
This dataset accompanies the Rule2DRC paper. See also the Hugging Face Papers page.
Usage
from datasets import load_dataset
tasks = load_dataset("jusjinuk/Rule2DRC", "tasks", split="test")
testcases = load_dataset("jusjinuk/Rule2DRC", "testcases", split="test")
Dataset Structure
tasks: 1000 problem rows… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/Rule2DRC.rules
rules
Natural-language LLM-agent rule files (AGENTS.md, CLAUDE.md, SKILL.md, .cursor/rules/*.mdc,
and friends) crawled from public GitHub repositories, with the content stored inline.
2,204,470 files from 45,014 repositories. Each row is one rule file, pinned to the commit SHA
it was read at, so link always resolves to the exact bytes in file.
Columns
column
type
description
file
string
full text of the rule file
content_sha256
string
SHA-256 of file… See the full description on the dataset page: https://huggingface.co/datasets/Andwwy/rules.private-letter-rulings
Private Letter Rulings
Text of IRS Private Letter Rulings (and other written determinations [TAMs, CCAs, etc.]), covering 1999 through August 2026. The IRS publishes these as PDF files each week; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed.
Dataset Structure
45,401 rows, one per ruling. Columns:
Column
Type
Description
wd_number
string
9-digit IRS written determination number: 4-digit year + 2-digit week… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/private-letter-rulings.touch-rugby-rules
Touch Rugby Rules Dataset
train.csv is comprised of a set of questions based on rules from the International Touch Website
For educational and non-commercial use only.
task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.eleusis-calibrated-rules
Eleusis Calibrated Rules — 100-turn reward calibration
A calibrated rule dataset for the single-player Eleusis inductive-reasoning
environment. It extends the 26-rule Hugging Face benchmark with controlled
static, transition, conditional, periodic, chunk, higher-order history, global
history, and compositional rule families.
Source benchmark: Hugging Face Eleusis.
Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11
The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.CodeAnything-1.835M
CodeAnything 1.835M SFT and evaluation release
Gated public release containing the clean 1,835,476-sample training set, the
800-sample/16-domain evaluation set, paper-model predictions and rendered
outputs, and raw per-sample rating records.
Layout
training/
manifest/
all_training_v5.jsonl
all_training_v5.jsonl.idx
all_training_v5.jsonl.true_lengths.u32
shards/<domain>/ exact media/code closure (tar shards)
evaluation/
benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.yeji-bazi-rules
██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗
██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝
██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗
██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║
██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║
╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝
⚡ INTERPRETATION RULEBOOK ⚡
> ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.revenue-rulings
Revenue Rulings
Text of IRS published guidance — Revenue Rulings, Revenue Procedures, Notices, Announcements, and a small number of Information Releases — sourced from the IRS's guidance drop folder, covering 2000 through August 2026. The IRS publishes these as PDF files; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed.
Dataset Structure
3,479 rows, one per document. Columns:
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/revenue-rulings.ruler-2m-niah-external
RULER-2M NIAH Eval — external handoff
50 × single_needle_uuid samples at ~2 M tokens per sample. One of four
length variants (1 M / 2 M / 6 M / 12 M) prepared for the external
long-context retrieval handoff.
field
value
samples
50
tasks
{single_needle_uuid: 50}
target tokens
2,000,000
seed
1344
negative_rate
0.0
eval/heldout/data.jsonl is the chat-templated form, ready for
model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-2m-niah-external.ru-libinpoc-11k11,5k russian books in txt format, divided by genres
11,5 тыщ книг русской литературы. датасет сделан из древнющего диска "lib in poc"
scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.ruler-1m-niah-external
RULER-1M NIAH Eval — external handoff
50 × single_needle_uuid samples at ~1 M tokens per sample. One of four
length variants (1 M / 2 M / 6 M / 12 M) prepared for the external
long-context retrieval handoff.
field
value
samples
50
tasks
{single_needle_uuid: 50}
target tokens
1,000,000
seed
1344
negative_rate
0.0
eval/heldout/data.jsonl is the chat-templated form, ready for
model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-1m-niah-external.ruleloopvit-sft-generated-rules-019200
RuleLoopViT SFT Generated ARC-AGI-1 Rules
This dataset contains one generated rule text for each of the 400 ARC-AGI-1
training tasks.
The rules were generated by the RuleLoopViT principal SFT language model:
omrisap/sft_lm_principal
Each task was prompted with 2-4 official ARC-AGI demonstration pairs and decoded
deterministically. The generated output follows the project's five-section rule
schema:
[CORE RULE]
[INPUT STRUCTURE]
[TARGET SELECTION]
[TRANSFORMATION]
[OUTPUT… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/ruleloopvit-sft-generated-rules-019200.memory_sft_data
memory_sft_data
SFT data that teaches an agent to manage its own context window while solving
software-engineering tasks: gather the right code, compress aggressively with an
edit_context tool (offloading stale output to a memory store and leaving a short
self-contained note), and reuse that offloaded memory like a retrieval datastore
(ls/grep/cat over /tmp/.unified_memory/), then write a precise, grounded fix
plan that recalls offloaded details.
Each example is a full… See the full description on the dataset page: https://huggingface.co/datasets/rulins/memory_sft_data.cbp-rulings-past-2012
AI-Extracted CBP Customs Rulings Dataset
Dataset Summary
This dataset contains itemized product classifications extracted from U.S. Customs and Border Protection (CBP) rulings published on the CROSS (Customs Rulings Online Search System) database. Text fields, descriptions, and Harmonized System (HS) codes were extracted and structured using Gemini AI models thanks to Google's generous free tier.
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/fklc/cbp-rulings-past-2012.ruler-12m-niah-external
RULER-12M NIAH Eval — external handoff
50 × single_needle_uuid samples at ~12 M tokens per sample. One of four
length variants (1 M / 2 M / 6 M / 12 M) prepared for the external
long-context retrieval handoff.
field
value
samples
50
tasks
{single_needle_uuid: 50}
target tokens
12,000,000
seed
1344
negative_rate
0.0
eval/heldout/data.jsonl is the chat-templated form, ready for
model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-12m-niah-external.eleusis-frontier-rules
Eleusis Frontier Rules
A simple rule dataset for the
nph4rd/eleusis
inductive-reasoning environment.
It contains 1,228 rules from eight rule families:
train: 907 rules
validation: 289 rules
test: 32 rules, with four rules from each family
Each row has exactly four fields:
rule_id: unique rule identifier
label: human-readable rule label
family: semantic rule family
code: executable hidden-rule predicate
Run it with the environment's default settings:
uv run eval… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-frontier-rules.gst-rulings-corpus
GST/Tax Regulatory Text Corpus
A narrow-domain corpus of Indian GST (Goods and Services Tax) regulatory
text, assembled for pretraining a small (~130M parameter) language model
from scratch, following Sebastian Raschka's Build a Large Language Model
From Scratch.
Contents
2150 training documents / 238 validation documents
~10,621,967 tokens (GPT-2 BPE)
Two source types:
circulars — CGST circulars from India Code (indiacode.nic.in)
aar_rulings — Authority for… See the full description on the dataset page: https://huggingface.co/datasets/Tharun007/gst-rulings-corpus.touch-rugby-rules-unsupervised
Touch Rugby Rules Dataset
train.csv is taken from the International Touch Website
All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible.
For educational and non-commercial use only.
rulefollower_results
RuleFollower Results
This dataset repository contains parser outputs and downstream annotation results for the RuleFollower experiments.
The repository is currently organized as a flat set of top-level folders.
Annotation result folders
These folders contain the final claim-level outputs for the three downstream tasks:
accuracy_outputs/
difficulty_outputs/
reasoning_outputs/
Folder meanings:
annotation_gpt-oss
parser: gpt-oss-120b
annotation model: gpt-oss-120b
input… See the full description on the dataset page: https://huggingface.co/datasets/RuleFollower/rulefollower_results.falcon-snort-cti-rule
FALCON SNORT CTI ↔ Ground-Truth Rule Dataset
Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders.
Schema
column
type
description
cti
string
CTI description
gold_rule
string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.
