CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes605 downloads2y agoHugging Face02alirezaaminzadeh /sigmaforge-detection-rules SigmaForge Detection Rules SigmaForge is a structured, operational dataset for building and evaluating systems that generate, validate, and translate Sigma detection rules. Sigma is a vendor-agnostic YAML format that describes detection logic so it can be shared across SIEM platforms. The dataset is derived from the open-source SigmaHQ rule corpus. Every rule is normalized and enriched with: MITRE ATT&CK technique and tactic mappings extracted from rule tags. Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.texttext-generation1K<n<10K0 likes97 downloads2mo agoHugging Face03jusjinuk /Rule2DRC Rule2DRC Rule2DRC is a benchmark for generating KLayout DRC Ruby runsets from natural-language design-rule specifications. Paper This dataset accompanies the Rule2DRC paper. See also the Hugging Face Papers page. Usage from datasets import load_dataset tasks = load_dataset("jusjinuk/Rule2DRC", "tasks", split="test") testcases = load_dataset("jusjinuk/Rule2DRC", "testcases", split="test") Dataset Structure tasks: 1000 problem rows… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/Rule2DRC.texttext-generation10K<n<100K0 likes95 downloads4mo agoHugging Face04andrew-mitchel /private-letter-rulings Private Letter Rulings Text of IRS Private Letter Rulings (and other written determinations [TAMs, CCAs, etc.]), covering 1999 through August 2026. The IRS publishes these as PDF files each week; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed. Dataset Structure 45,401 rows, one per ruling. Columns: Column Type Description wd_number string 9-digit IRS written determination number: 4-digit year + 2-digit week… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/private-letter-rulings.texttext-generation10K<n<100K0 likes92 downloads23d agoHugging Face05Trelis /touch-rugby-rules Touch Rugby Rules Dataset train.csv is comprised of a set of questions based on rules from the International Touch Website For educational and non-commercial use only. texttext-generationn<1K0 likes91 downloads3y agoHugging Face06Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face07nph4rd /eleusis-calibrated-rules Eleusis Calibrated Rules — 100-turn reward calibration A calibrated rule dataset for the single-player Eleusis inductive-reasoning environment. It extends the 26-rule Hugging Face benchmark with controlled static, transition, conditional, periodic, chunk, higher-order history, global history, and compositional rule families. Source benchmark: Hugging Face Eleusis. Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11 The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.imagereinforcement-learning1K<n<10K0 likes79 downloads1mo agoHugging Face08tellang /yeji-bazi-rules ██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗ ██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝ ██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗ ██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║ ██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║ ╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝ ⚡ INTERPRETATION RULEBOOK ⚡ > ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.imagetext-generationn<1K1 likes70 downloads8mo agoHugging Face09andrew-mitchel /revenue-rulings Revenue Rulings Text of IRS published guidance — Revenue Rulings, Revenue Procedures, Notices, Announcements, and a small number of Information Releases — sourced from the IRS's guidance drop folder, covering 2000 through August 2026. The IRS publishes these as PDF files; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed. Dataset Structure 3,479 rows, one per document. Columns: Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/revenue-rulings.texttext-generation1K<n<10K0 likes63 downloads22d agoHugging Face10ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes49 downloads2y agoHugging Face11fklc /cbp-rulings-past-2012 AI-Extracted CBP Customs Rulings Dataset Dataset Summary This dataset contains itemized product classifications extracted from U.S. Customs and Border Protection (CBP) rulings published on the CROSS (Customs Rulings Online Search System) database. Text fields, descriptions, and Harmonized System (HS) codes were extracted and structured using Gemini AI models thanks to Google's generous free tier. Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/fklc/cbp-rulings-past-2012.tabularother10K<n<100K0 likes29 downloads2mo agoHugging Face12nph4rd /eleusis-frontier-rules Eleusis Frontier Rules A simple rule dataset for the nph4rd/eleusis inductive-reasoning environment. It contains 1,228 rules from eight rule families: train: 907 rules validation: 289 rules test: 32 rules, with four rules from each family Each row has exactly four fields: rule_id: unique rule identifier label: human-readable rule label family: semantic rule family code: executable hidden-rule predicate Run it with the environment's default settings: uv run eval… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-frontier-rules.textreinforcement-learning1K<n<10K0 likes27 downloads1mo agoHugging Face13Tharun007 /gst-rulings-corpus GST/Tax Regulatory Text Corpus A narrow-domain corpus of Indian GST (Goods and Services Tax) regulatory text, assembled for pretraining a small (~130M parameter) language model from scratch, following Sebastian Raschka's Build a Large Language Model From Scratch. Contents 2150 training documents / 238 validation documents ~10,621,967 tokens (GPT-2 BPE) Two source types: circulars — CGST circulars from India Code (indiacode.nic.in) aar_rulings — Authority for… See the full description on the dataset page: https://huggingface.co/datasets/Tharun007/gst-rulings-corpus.texttext-generation1K<n<10K1 likes27 downloads1mo agoHugging Face14Trelis /touch-rugby-rules-unsupervised Touch Rugby Rules Dataset train.csv is taken from the International Touch Website All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible. For educational and non-commercial use only. texttext-generationn<1K0 likes24 downloads3y agoHugging Face15shaswatamitra /falcon-snort-cti-rule FALCON SNORT CTI ↔ Ground-Truth Rule Dataset Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders. Schema column type description cti string CTI description gold_rule string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.texttext-generationn<1K0 likes23 downloads4mo agoHugging Face16shaswatamitra /falcon-yara-cti-rule FALCON YARA CTI ↔ Ground-Truth Rule Dataset Cyber-threat-intelligence descriptions paired with their ground-truth YARA rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders. Schema column type description cti string CTI description gold_rule string ground-truth… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-yara-cti-rule.texttext-generationn<1K0 likes21 downloads4mo agoHugging Face17Trelis /touch-rugby-rules-embeddings Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes20 downloads3y agoHugging Face18acazau /touch-rugby-rules-embeddings Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes18 downloads3y agoHugging Face19aioil-ai /polish-court-rulings-sample Polish Court Rulings — Sample (korpus-pl) A production-grade, PII-hardened corpus of Polish court rulings — free evaluation sample. Full corpus: 505,611 rulings · ~3.18B tokens, licensed commercially. Contact: licensing@aioil.ai · aioil.ai What this is This sample contains 500 Polish court rulings drawn from the full korpus-pl dataset — a cleaned, deduplicated and PII-audited corpus of Polish jurisprudence built for AI training, evaluation and legal RAG… See the full description on the dataset page: https://huggingface.co/datasets/aioil-ai/polish-court-rulings-sample.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face20Lots-of-LoRAs /task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task967_ruletaker_incorrect_fact_generation_based_on_given_paragraph.texttext-generationn<1K0 likes16 downloads2y agoHugging Face21Wendy-Thompson-Lending-Team /alimony-rules-by-state-2026 Alimony Rules By State 2026 Alimony/spousal support rules for 13 states with mortgage impact. Details Records: 13 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814 Publisher: Wendy Thompson Lending Team Thompson Alpha Logic State-by-state alimony duration and calculation methods mapped to mortgage qualification impact. Shows how alimony income qualifies (or disqualifies) for FHA, VA, and… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/alimony-rules-by-state-2026.tabularquestion-answeringn<1K0 likes16 downloads6mo agoHugging Face22halilozturkci /touch-rugby-rules-embeddings Touch Rugby Rules Dataset (for embeddings) train.csv is taken from the International Touch Website test.csv is copy pasted from abbreviated rules on the UK Touch website. Note that I'm bypassing the pdf to text stage. All text is chunked to a length of 100 tokens with 50% overlap. For educational and non-commercial use only. texttext-generationn<1K0 likes14 downloads3y agoHugging Face23Rulga /LS_chattexttext-generationn<1K0 likes10 downloads2y agoHugging Face24ritulk /baggageitems_rules_llama2textquestion-answeringn<1K0 likes9 downloads1y agoHugging Face25CBERX /ru-linux-sysadmin-dialogues Russian Linux Sysadmin Dialogues (Датасет для обучения ИИ) Высококачественный структурированный набор данных (датасет) на русском языке, содержащий профессиональные инструкции, разборы технических проблем и сценарии общения в сфере системного администрирования операционных систем семейства Linux. Этот датасет разработан специально для тонкой настройки (fine-tuning) больших языковых моделей (LLM), обучения диалоговых агентов, умных помощников технической поддержки и наполнения… See the full description on the dataset page: https://huggingface.co/datasets/CBERX/ru-linux-sysadmin-dialogues.texttext-generationn<1K0 likes8 downloads3mo agoHugging Face26bond005 /ru_llm_calibration Ru LLM calibration This dataset is created by Ivan Bondarenko for calibrating (importance matrix computation) and evaluating GGUF quantizations of large language models targeting Russian language, including but not limited to Meno-Lite-0.1-GGUF. Purpose Train split: calibration for llama.cpp quantization (any Russian-focused LLM). Test split: quality evaluation (perplexity, etc.) via llama-perplexity or similar tools. Dataset Composition Train: Russian… See the full description on the dataset page: https://huggingface.co/datasets/bond005/ru_llm_calibration.texttext-generation1K<n<10K1 likes6 downloads5mo agoHugging Face27broadfield-dev /lora-rules-dataset LoRA Rules Dataset Synthetic behavioral rules dataset for training a hypernetwork that generates LoRA adapters on-the-fly from structured rule strings. Format Each record is a JSON line with fields: rule_id — unique identifier rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone weight — float 0.0–1.0, importance of the rule description — natural language rule description raw — full rule string [RuleType|Weight] Description training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-dataset.tabulartext-generationn<1K0 likes4 downloads7mo agoHugging Face28broadfield-dev /lora-rules-qwen3-0.6b-r8-n180 LoRA Rules Dataset Synthetic behavioral rules dataset for training a hypernetwork that generates LoRA adapters on-the-fly from structured rule strings. Format Each record is a JSON line with fields: rule_id — unique identifier rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone weight — float 0.0–1.0, importance of the rule description — natural language rule description raw — full rule string [RuleType|Weight] Description training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-qwen3-0.6b-r8-n180.tabulartext-generationn<1K0 likes3 downloads7mo agoHugging Face29LutzMoran /pi-of-ai-rules-sft Pi-of-AI · Rules-Baker SFT dataset Synthetic supervised-fine-tuning data for Pi-of-AI / Rules-Baker, generated fully locally against an OpenAI-compatible teacher (Ollama running Qwen2.5-Coder). Each example is a chat-format messages record. Two kinds: positive — an ordinary coding request → rule-compliant code. revision — rule-violating code → corrected code (teaches self-repair). The whole trick: the teacher saw the house-style rules when writing the target code, but the… See the full description on the dataset page: https://huggingface.co/datasets/LutzMoran/pi-of-ai-rules-sft.texttext-generationn<1K0 likes3 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.