datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BenchMAX_Rule-based
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios.
We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.sigmaforge-detection-rules
SigmaForge Detection Rules
SigmaForge is a structured, operational dataset for building and evaluating
systems that generate, validate, and translate Sigma
detection rules. Sigma is a vendor-agnostic YAML format that describes
detection logic so it can be shared across SIEM platforms.
The dataset is derived from the open-source SigmaHQ
rule corpus. Every rule is normalized and enriched with:
MITRE ATT&CK technique and tactic mappings extracted from rule tags.
Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.Rule2DRC
Rule2DRC
Rule2DRC is a benchmark for generating KLayout DRC Ruby runsets from natural-language design-rule specifications.
Paper
This dataset accompanies the Rule2DRC paper. See also the Hugging Face Papers page.
Usage
from datasets import load_dataset
tasks = load_dataset("jusjinuk/Rule2DRC", "tasks", split="test")
testcases = load_dataset("jusjinuk/Rule2DRC", "testcases", split="test")
Dataset Structure
tasks: 1000 problem rows… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/Rule2DRC.private-letter-rulings
Private Letter Rulings
Text of IRS Private Letter Rulings (and other written determinations [TAMs, CCAs, etc.]), covering 1999 through August 2026. The IRS publishes these as PDF files each week; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed.
Dataset Structure
45,401 rows, one per ruling. Columns:
Column
Type
Description
wd_number
string
9-digit IRS written determination number: 4-digit year + 2-digit week… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/private-letter-rulings.touch-rugby-rules
Touch Rugby Rules Dataset
train.csv is comprised of a set of questions based on rules from the International Touch Website
For educational and non-commercial use only.
task966_ruletaker_fact_checking_based_on_given_context
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.eleusis-calibrated-rules
Eleusis Calibrated Rules — 100-turn reward calibration
A calibrated rule dataset for the single-player Eleusis inductive-reasoning
environment. It extends the 26-rule Hugging Face benchmark with controlled
static, transition, conditional, periodic, chunk, higher-order history, global
history, and compositional rule families.
Source benchmark: Hugging Face Eleusis.
Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11
The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.yeji-bazi-rules
██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗
██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝
██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗
██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║
██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║
╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝
⚡ INTERPRETATION RULEBOOK ⚡
> ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.revenue-rulings
Revenue Rulings
Text of IRS published guidance — Revenue Rulings, Revenue Procedures, Notices, Announcements, and a small number of Information Releases — sourced from the IRS's guidance drop folder, covering 2000 through August 2026. The IRS publishes these as PDF files; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed.
Dataset Structure
3,479 rows, one per document. Columns:
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/revenue-rulings.scoutieDataset_russian_language_grammar_and_rules_vectorized
Description in English:
A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules.
The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link.
Dataset fields:
taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.cbp-rulings-past-2012
AI-Extracted CBP Customs Rulings Dataset
Dataset Summary
This dataset contains itemized product classifications extracted from U.S. Customs and Border Protection (CBP) rulings published on the CROSS (Customs Rulings Online Search System) database. Text fields, descriptions, and Harmonized System (HS) codes were extracted and structured using Gemini AI models thanks to Google's generous free tier.
Dataset Structure
Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/fklc/cbp-rulings-past-2012.eleusis-frontier-rules
Eleusis Frontier Rules
A simple rule dataset for the
nph4rd/eleusis
inductive-reasoning environment.
It contains 1,228 rules from eight rule families:
train: 907 rules
validation: 289 rules
test: 32 rules, with four rules from each family
Each row has exactly four fields:
rule_id: unique rule identifier
label: human-readable rule label
family: semantic rule family
code: executable hidden-rule predicate
Run it with the environment's default settings:
uv run eval… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-frontier-rules.gst-rulings-corpus
GST/Tax Regulatory Text Corpus
A narrow-domain corpus of Indian GST (Goods and Services Tax) regulatory
text, assembled for pretraining a small (~130M parameter) language model
from scratch, following Sebastian Raschka's Build a Large Language Model
From Scratch.
Contents
2150 training documents / 238 validation documents
~10,621,967 tokens (GPT-2 BPE)
Two source types:
circulars — CGST circulars from India Code (indiacode.nic.in)
aar_rulings — Authority for… See the full description on the dataset page: https://huggingface.co/datasets/Tharun007/gst-rulings-corpus.touch-rugby-rules-unsupervised
Touch Rugby Rules Dataset
train.csv is taken from the International Touch Website
All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible.
For educational and non-commercial use only.
falcon-snort-cti-rule
FALCON SNORT CTI ↔ Ground-Truth Rule Dataset
Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders.
Schema
column
type
description
cti
string
CTI description
gold_rule
string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.falcon-yara-cti-rule
FALCON YARA CTI ↔ Ground-Truth Rule Dataset
Cyber-threat-intelligence descriptions paired with their ground-truth YARA rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders.
Schema
column
type
description
cti
string
CTI description
gold_rule
string
ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-yara-cti-rule.touch-rugby-rules-embeddings
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
touch-rugby-rules-embeddings
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
polish-court-rulings-sample
Polish Court Rulings — Sample (korpus-pl)
A production-grade, PII-hardened corpus of Polish court rulings — free evaluation sample.
Full corpus: 505,611 rulings · ~3.18B tokens, licensed commercially.
Contact: licensing@aioil.ai · aioil.ai
What this is
This sample contains 500 Polish court rulings drawn from the full korpus-pl dataset — a cleaned, deduplicated and PII-audited corpus of Polish jurisprudence built for AI training, evaluation and legal RAG… See the full description on the dataset page: https://huggingface.co/datasets/aioil-ai/polish-court-rulings-sample.task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph.alimony-rules-by-state-2026
Alimony Rules By State 2026
Alimony/spousal support rules for 13 states with mortgage impact.
Details
Records: 13
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814
Publisher: Wendy Thompson Lending Team
Thompson Alpha Logic
State-by-state alimony duration and calculation methods mapped to mortgage qualification impact. Shows how alimony income qualifies (or disqualifies) for FHA, VA, and… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/alimony-rules-by-state-2026.touch-rugby-rules-embeddings
Touch Rugby Rules Dataset (for embeddings)
train.csv is taken from the International Touch Website
test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage.
All text is chunked to a length of 100 tokens with 50% overlap.
For educational and non-commercial use only.
LS_chatbaggageitems_rules_llama2ru-linux-sysadmin-dialogues
Russian Linux Sysadmin Dialogues (Датасет для обучения ИИ)
Высококачественный структурированный набор данных (датасет) на русском языке, содержащий профессиональные инструкции, разборы технических проблем и сценарии общения в сфере системного администрирования операционных систем семейства Linux.
Этот датасет разработан специально для тонкой настройки (fine-tuning) больших языковых моделей (LLM), обучения диалоговых агентов, умных помощников технической поддержки и наполнения… See the full description on the dataset page: https://huggingface.co/datasets/CBERX/ru-linux-sysadmin-dialogues.ru_llm_calibration
Ru LLM calibration
This dataset is created by Ivan Bondarenko for calibrating (importance matrix computation) and evaluating GGUF quantizations of large language models targeting Russian language, including but not limited to Meno-Lite-0.1-GGUF.
Purpose
Train split: calibration for llama.cpp quantization (any Russian-focused LLM).
Test split: quality evaluation (perplexity, etc.) via llama-perplexity or similar tools.
Dataset Composition
Train: Russian… See the full description on the dataset page: https://huggingface.co/datasets/bond005/ru_llm_calibration.lora-rules-dataset
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-dataset.lora-rules-qwen3-0.6b-r8-n180
LoRA Rules Dataset
Synthetic behavioral rules dataset for training a hypernetwork that generates
LoRA adapters on-the-fly from structured rule strings.
Format
Each record is a JSON line with fields:
rule_id — unique identifier
rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone
weight — float 0.0–1.0, importance of the rule
description — natural language rule description
raw — full rule string [RuleType|Weight] Description
training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-qwen3-0.6b-r8-n180.pi-of-ai-rules-sft
Pi-of-AI · Rules-Baker SFT dataset
Synthetic supervised-fine-tuning data for Pi-of-AI / Rules-Baker,
generated fully locally against an OpenAI-compatible teacher (Ollama running
Qwen2.5-Coder).
Each example is a chat-format messages record. Two kinds:
positive — an ordinary coding request → rule-compliant code.
revision — rule-violating code → corrected code (teaches self-repair).
The whole trick: the teacher saw the house-style rules when writing the
target code, but the… See the full description on the dataset page: https://huggingface.co/datasets/LutzMoran/pi-of-ai-rules-sft.
