CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jeffry77 /Rule-VLN Rule-VLN Dataset Rule-VLN is a rule-compliant outdoor vision-and-language navigation benchmark built on the Touchdown / StreetLearn urban navigation environment. It studies whether navigation agents can follow language instructions while also complying with semantic traffic rules, such as regulatory signs that prohibit otherwise reachable movements. This dataset accompanies the paper: Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric… See the full description on the dataset page: https://huggingface.co/datasets/jeffry77/Rule-VLN.text-generation1 likes1.1k downloads3mo agoHugging Face02LLaMAX /BenchMAX_Rule-based Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Rule-based is a dataset of BenchMAX, sourcing from IFEval, which is a rule-based benchmark for evaluating the instruction following capabilities in multilingual scenarios. We extend the original dataset to 16 non-English languages by first… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Rule-based.texttext-generation1K<n<10K1 likes605 downloads2y agoHugging Face03jet-ai /RULER-500 RULER Dataset (Qwen3 Tokenizer) Dataset generated from RULER for long context evaluation (qwen3 tokenizer). This dataset is used in the paper Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE. The official code is available at github.com/jet-ai-projects/jet-long. text-generation0 likes384 downloads2mo agoHugging Face04VenusChenyy /RULER_50 RULER_50 Official-Code Qwen3 Subset This dataset is a fixed 50-sample-per-group subset of RULER synthetic tasks. It was generated from the official NVIDIA/RULER GitHub code, not from a third-party pre-generated mirror. Official generation source: Repository: https://github.com/NVIDIA/RULER Branch: main Commit: 38da79d79519ef87aa46ae804f838e1eab7f86d7 Generation entrypoint: scripts/data/prepare.py Benchmark config: scripts/synthetic.yaml Generation settings: tokenizer:… See the full description on the dataset page: https://huggingface.co/datasets/VenusChenyy/RULER_50.text-generation1 likes354 downloads1mo agoHugging Face05sxiong /DHSA_RULER RULER Evaluation Data This dataset contains pre-generated JSONL files for the RULER long-context evaluation benchmark, used in Long-Context Modeling with Dynamic Hierarchical Sparse Attention for Memory-Constrained LLM Inference (ICML 2026 Spotlight). RULER is designed to evaluate effective context length and long-context behavior beyond simple retrieval, covering retrieval, multi-hop tracing, aggregation, and question answering style tasks. The files are organized by target… See the full description on the dataset page: https://huggingface.co/datasets/sxiong/DHSA_RULER.text-generation10K<n<100K1 likes301 downloads2mo agoHugging Face06gyung /ruler-niah-multilength-eval-benchmark 📌 Fixed Multi-Length RULER NIAH Benchmark (1K, 2K, 4K, 8K) Deterministic synthetic Needle-In-A-Haystack (NIAH) benchmark splits for reproducible long-context evaluation. Dataset Specifications: Tasks (4): niah_single_1: Repeat haystack, single word needle, number value. niah_single_2: Essay haystack, single word needle, number value. niah_single_3: Essay haystack, single word needle, UUID value. niah_multikey_1: Essay haystack, 4 keys needle, number value.… See the full description on the dataset page: https://huggingface.co/datasets/gyung/ruler-niah-multilength-eval-benchmark.question-answeringn<1K0 likes154 downloads1mo agoHugging Face07IlyaGusev /rulm Dataset for training Russian language models Overall: 75G Scripts: https://github.com/IlyaGusev/rulm/tree/master/data_processing Website Char count (M) Word count (M) pikabu 14938 2161 lenta 1008 135 stihi 2994 393 stackoverflow 1073 228 habr 5112 753 taiga_fontanka 419 55 librusec 10149 1573 buriy 2646 352 ods_tass 1908 255 wiki 3473 469 math987 177 text-generation10M<n<100M22 likes124 downloads4y agoHugging Face08alirezaaminzadeh /sigmaforge-detection-rules SigmaForge Detection Rules SigmaForge is a structured, operational dataset for building and evaluating systems that generate, validate, and translate Sigma detection rules. Sigma is a vendor-agnostic YAML format that describes detection logic so it can be shared across SIEM platforms. The dataset is derived from the open-source SigmaHQ rule corpus. Every rule is normalized and enriched with: MITRE ATT&CK technique and tactic mappings extracted from rule tags. Compiled SIEM… See the full description on the dataset page: https://huggingface.co/datasets/alirezaaminzadeh/sigmaforge-detection-rules.texttext-generation1K<n<10K0 likes97 downloads2mo agoHugging Face09jusjinuk /Rule2DRC Rule2DRC Rule2DRC is a benchmark for generating KLayout DRC Ruby runsets from natural-language design-rule specifications. Paper This dataset accompanies the Rule2DRC paper. See also the Hugging Face Papers page. Usage from datasets import load_dataset tasks = load_dataset("jusjinuk/Rule2DRC", "tasks", split="test") testcases = load_dataset("jusjinuk/Rule2DRC", "testcases", split="test") Dataset Structure tasks: 1000 problem rows… See the full description on the dataset page: https://huggingface.co/datasets/jusjinuk/Rule2DRC.texttext-generation10K<n<100K0 likes95 downloads4mo agoHugging Face10Andwwy /rules rules Natural-language LLM-agent rule files (AGENTS.md, CLAUDE.md, SKILL.md, .cursor/rules/*.mdc, and friends) crawled from public GitHub repositories, with the content stored inline. 2,204,470 files from 45,014 repositories. Each row is one rule file, pinned to the commit SHA it was read at, so link always resolves to the exact bytes in file. Columns column type description file string full text of the rule file content_sha256 string SHA-256 of file… See the full description on the dataset page: https://huggingface.co/datasets/Andwwy/rules.text-generation1M<n<10M0 likes92 downloads2mo agoHugging Face11andrew-mitchel /private-letter-rulings Private Letter Rulings Text of IRS Private Letter Rulings (and other written determinations [TAMs, CCAs, etc.]), covering 1999 through August 2026. The IRS publishes these as PDF files each week; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed. Dataset Structure 45,401 rows, one per ruling. Columns: Column Type Description wd_number string 9-digit IRS written determination number: 4-digit year + 2-digit week… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/private-letter-rulings.texttext-generation10K<n<100K0 likes92 downloads24d agoHugging Face12Trelis /touch-rugby-rules Touch Rugby Rules Dataset train.csv is comprised of a set of questions based on rules from the International Touch Website For educational and non-commercial use only. texttext-generationn<1K0 likes91 downloads3y agoHugging Face13Lots-of-LoRAs /task966_ruletaker_fact_checking_based_on_given_context Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task966_ruletaker_fact_checking_based_on_given_context Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task966_ruletaker_fact_checking_based_on_given_context.texttext-generationn<1K0 likes85 downloads2y agoHugging Face14nph4rd /eleusis-calibrated-rules Eleusis Calibrated Rules — 100-turn reward calibration A calibrated rule dataset for the single-player Eleusis inductive-reasoning environment. It extends the 26-rule Hugging Face benchmark with controlled static, transition, conditional, periodic, chunk, higher-order history, global history, and compositional rule families. Source benchmark: Hugging Face Eleusis. Dataset version: v2.1-frontier-calibrated-100turn-20260812Protocol: eleusis-100-v11 The structural, GPT Sol… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-calibrated-rules.imagereinforcement-learning1K<n<10K0 likes79 downloads2mo agoHugging Face15Ruler138 /CodeAnything-1.835Mgated CodeAnything 1.835M SFT and evaluation release Gated public release containing the clean 1,835,476-sample training set, the 800-sample/16-domain evaluation set, paper-model predictions and rendered outputs, and raw per-sample rating records. Layout training/ manifest/ all_training_v5.jsonl all_training_v5.jsonl.idx all_training_v5.jsonl.true_lengths.u32 shards/<domain>/ exact media/code closure (tar shards) evaluation/ benchmark/… See the full description on the dataset page: https://huggingface.co/datasets/Ruler138/CodeAnything-1.835M.imageimage-to-text10K<n<100K0 likes78 downloads10d agoHugging Face16tellang /yeji-bazi-rules ██████╗ █████╗ ███████╗██╗ ██████╗ ██╗ ██╗██╗ ███████╗███████╗ ██╔══██╗██╔══██╗╚══███╔╝██║ ██╔══██╗██║ ██║██║ ██╔════╝██╔════╝ ██████╔╝███████║ ███╔╝ ██║ ██████╔╝██║ ██║██║ █████╗ ███████╗ ██╔══██╗██╔══██║ ███╔╝ ██║ ██╔══██╗██║ ██║██║ ██╔══╝ ╚════██║ ██████╔╝██║ ██║███████╗██║ ██║ ██║╚██████╔╝███████╗███████╗███████║ ╚═════╝ ╚═╝ ╚═╝╚══════╝╚═╝ ╚═╝ ╚═╝ ╚═════╝ ╚══════╝╚══════╝╚══════╝ ⚡ INTERPRETATION RULEBOOK ⚡ > ACCESS… See the full description on the dataset page: https://huggingface.co/datasets/tellang/yeji-bazi-rules.imagetext-generationn<1K1 likes70 downloads8mo agoHugging Face17andrew-mitchel /revenue-rulings Revenue Rulings Text of IRS published guidance — Revenue Rulings, Revenue Procedures, Notices, Announcements, and a small number of Information Releases — sourced from the IRS's guidance drop folder, covering 2000 through August 2026. The IRS publishes these as PDF files; these were converted to text using pdfminer, falling back to OCR via pytesseract where needed. Dataset Structure 3,479 rows, one per document. Columns: Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/andrew-mitchel/revenue-rulings.texttext-generation1K<n<10K0 likes63 downloads23d agoHugging Face18ryansubq /ruler-2m-niah-external RULER-2M NIAH Eval — external handoff 50 × single_needle_uuid samples at ~2 M tokens per sample. One of four length variants (1 M / 2 M / 6 M / 12 M) prepared for the external long-context retrieval handoff. field value samples 50 tasks {single_needle_uuid: 50} target tokens 2,000,000 seed 1344 negative_rate 0.0 eval/heldout/data.jsonl is the chat-templated form, ready for model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-2m-niah-external.text-generationn<1K0 likes54 downloads3mo agoHugging Face194eJIoBek /ru-libinpoc-11k11,5k russian books in txt format, divided by genres 11,5 тыщ книг русской литературы. датасет сделан из древнющего диска "lib in poc" text-generation10K<n<100K2 likes52 downloads4y agoHugging Face20ScoutieAutoML /scoutieDataset_russian_language_grammar_and_rules_vectorized Description in English: A dataset collected from 30 Russian-language Telegram channels on the topic of learning the Russian language. This dataset contains grammar, syntax, spelling and punctuation rules. The dataset was collected and marked automatically using the Scoutie data collection and marking service.Try Scoutie and collect the same or another dataset using the link. Dataset fields: taskId - task identifier in the Scouti service. text - main text. url -… See the full description on the dataset page: https://huggingface.co/datasets/ScoutieAutoML/scoutieDataset_russian_language_grammar_and_rules_vectorized.tabulartext-classification10K<n<100K2 likes49 downloads2y agoHugging Face21ryansubq /ruler-1m-niah-external RULER-1M NIAH Eval — external handoff 50 × single_needle_uuid samples at ~1 M tokens per sample. One of four length variants (1 M / 2 M / 6 M / 12 M) prepared for the external long-context retrieval handoff. field value samples 50 tasks {single_needle_uuid: 50} target tokens 1,000,000 seed 1344 negative_rate 0.0 eval/heldout/data.jsonl is the chat-templated form, ready for model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-1m-niah-external.text-generationn<1K0 likes44 downloads3mo agoHugging Face22omrisap /ruleloopvit-sft-generated-rules-019200 RuleLoopViT SFT Generated ARC-AGI-1 Rules This dataset contains one generated rule text for each of the 400 ARC-AGI-1 training tasks. The rules were generated by the RuleLoopViT principal SFT language model: omrisap/sft_lm_principal Each task was prompted with 2-4 official ARC-AGI demonstration pairs and decoded deterministically. The generated output follows the project's five-section rule schema: [CORE RULE] [INPUT STRUCTURE] [TARGET SELECTION] [TRANSFORMATION] [OUTPUT… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/ruleloopvit-sft-generated-rules-019200.text-generation0 likes44 downloads20d agoHugging Face23rulins /memory_sft_data memory_sft_data SFT data that teaches an agent to manage its own context window while solving software-engineering tasks: gather the right code, compress aggressively with an edit_context tool (offloading stale output to a memory store and leaving a short self-contained note), and reuse that offloaded memory like a retrieval datastore (ls/grep/cat over /tmp/.unified_memory/), then write a precise, grounded fix plan that recalls offloaded details. Each example is a full… See the full description on the dataset page: https://huggingface.co/datasets/rulins/memory_sft_data.text-generationn<1K0 likes33 downloads3mo agoHugging Face24fklc /cbp-rulings-past-2012 AI-Extracted CBP Customs Rulings Dataset Dataset Summary This dataset contains itemized product classifications extracted from U.S. Customs and Border Protection (CBP) rulings published on the CROSS (Customs Rulings Online Search System) database. Text fields, descriptions, and Harmonized System (HS) codes were extracted and structured using Gemini AI models thanks to Google's generous free tier. Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/fklc/cbp-rulings-past-2012.tabularother10K<n<100K0 likes29 downloads2mo agoHugging Face25ryansubq /ruler-12m-niah-external RULER-12M NIAH Eval — external handoff 50 × single_needle_uuid samples at ~12 M tokens per sample. One of four length variants (1 M / 2 M / 6 M / 12 M) prepared for the external long-context retrieval handoff. field value samples 50 tasks {single_needle_uuid: 50} target tokens 12,000,000 seed 1344 negative_rate 0.0 eval/heldout/data.jsonl is the chat-templated form, ready for model.forward(); eval/heldout/raw.jsonl is the pre-template form… See the full description on the dataset page: https://huggingface.co/datasets/ryansubq/ruler-12m-niah-external.text-generationn<1K0 likes27 downloads3mo agoHugging Face26nph4rd /eleusis-frontier-rules Eleusis Frontier Rules A simple rule dataset for the nph4rd/eleusis inductive-reasoning environment. It contains 1,228 rules from eight rule families: train: 907 rules validation: 289 rules test: 32 rules, with four rules from each family Each row has exactly four fields: rule_id: unique rule identifier label: human-readable rule label family: semantic rule family code: executable hidden-rule predicate Run it with the environment's default settings: uv run eval… See the full description on the dataset page: https://huggingface.co/datasets/nph4rd/eleusis-frontier-rules.textreinforcement-learning1K<n<10K0 likes27 downloads1mo agoHugging Face27Tharun007 /gst-rulings-corpus GST/Tax Regulatory Text Corpus A narrow-domain corpus of Indian GST (Goods and Services Tax) regulatory text, assembled for pretraining a small (~130M parameter) language model from scratch, following Sebastian Raschka's Build a Large Language Model From Scratch. Contents 2150 training documents / 238 validation documents ~10,621,967 tokens (GPT-2 BPE) Two source types: circulars — CGST circulars from India Code (indiacode.nic.in) aar_rulings — Authority for… See the full description on the dataset page: https://huggingface.co/datasets/Tharun007/gst-rulings-corpus.texttext-generation1K<n<10K1 likes27 downloads1mo agoHugging Face28Trelis /touch-rugby-rules-unsupervised Touch Rugby Rules Dataset train.csv is taken from the International Touch Website All text is chunked to a length of 250 tokens, aiming to keep sentences whole where possible. For educational and non-commercial use only. texttext-generationn<1K0 likes24 downloads3y agoHugging Face29RuleFollower /rulefollower_results RuleFollower Results This dataset repository contains parser outputs and downstream annotation results for the RuleFollower experiments. The repository is currently organized as a flat set of top-level folders. Annotation result folders These folders contain the final claim-level outputs for the three downstream tasks: accuracy_outputs/ difficulty_outputs/ reasoning_outputs/ Folder meanings: annotation_gpt-oss parser: gpt-oss-120b annotation model: gpt-oss-120b input… See the full description on the dataset page: https://huggingface.co/datasets/RuleFollower/rulefollower_results.text-classification0 likes24 downloads6mo agoHugging Face30shaswatamitra /falcon-snort-cti-rule FALCON SNORT CTI ↔ Ground-Truth Rule Dataset Cyber-threat-intelligence descriptions paired with their ground-truth SNORT IDS rule and the LLM-generated decoy rules (hard-negative look-alikes) of that gold rule. This is the FALCON held-out test benchmark used both for LLM rule-generation evaluation and for retrieval testing of contrastively fine-tuned sentence encoders. Schema column type description cti string CTI description gold_rule string… See the full description on the dataset page: https://huggingface.co/datasets/shaswatamitra/falcon-snort-cti-rule.texttext-generationn<1K0 likes23 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.