CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rcds /swiss_rulings Dataset Card for Swiss Rulings Dataset Summary SwissRulings is a multilingual, diachronic dataset of 637K Swiss Federal Supreme Court (FSCS) cases. This dataset can be used to pretrain language models on Swiss legal data. Supported Tasks and Leaderboards Languages Switzerland has four official languages with three languages German, French and Italian being represenated. The decisions are written by the judges and clerks in the language of the… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_rulings.tabular100K<n<1M1 likes265 downloads3y agoHugging Face02elichen-skymizer /lm-eval-ruler-results-private-32K Dataset Card for Evaluation run of elichen3051/Llama-3.1-8B-GGUF Dataset automatically created during the evaluation run of model elichen3051/Llama-3.1-8B-GGUF The dataset is composed of 12 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/elichen-skymizer/lm-eval-ruler-results-private-32K.tabular10K<n<100K0 likes185 downloads1y agoHugging Face03June30916 /multimodality-poc-llama31-ruler16k Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K Raw pre-RoPE query and hidden-state tensors captured during prefill, used to study whether the per-(layer, kv_head) query distribution is unimodal Gaussian (the assumption underpinning Expected Attention's MGF closed-form in kvpress). What's in here 65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5 prompts). Each file (~414 MB) contains: field dtype shape meaning hidden float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.tabularfeature-extractionn<1K0 likes142 downloads5mo agoHugging Face04aldea-ai /ruler_eval_data_128k RULER evaluation data — 128K context only This dataset is a subset of aldea-ai/ruler-eval-data, containing only the 128K context length (131072 tokens). The full multi-length dataset also includes 1M and other lengths. Files are published under 131072/ (numeric token count) for compatibility with benchmark_ruler.py --context_length 131072, even though the source snapshot uses a 128k/ folder name. Layout Same as the upstream RULER on-disk layout, compatible with… See the full description on the dataset page: https://huggingface.co/datasets/aldea-ai/ruler_eval_data_128k.tabular1K<n<10K0 likes120 downloads5mo agoHugging Face05oddadmix /arabic-rule-checking Arabic Rule Checking — قواعد ونصوص عربية بأحكام محسوبة 172,488 labelled (text, rule) pairs in Arabic. Each row asks one question: does this text satisfy this rule? The answer is مطابق or مخالف. بالعربية: مجموعة بيانات عربية للتحقق من مطابقة النصوص لقواعد مكتوبة بلغة طبيعية. كل صف يحتوي على نص وقاعدة وحكم محسوب آليًا، وليس رأي نموذج. split pairs texts train 159,240 48,030 validation 13,248 2,002 Built from 50,062 generated Arabic texts across 12 document types… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rule-checking.tabulartext-classification100K<n<1M0 likes93 downloads25d agoHugging Face06jet-ai /ruler-100-nemotron RULER-100 — Nemotron-Nano-v3 tokenized RULER long-context evaluation data, regenerated with the nvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 (instruct) tokenizer so the labeled context lengths are exact for that model — instead of drifting, as they do when RULER data tokenized for a different model (e.g. Qwen3) is fed to Nemotron. What's here 7 context lengths: 4096, 8192, 16384, 32768, 65536, 131072, 262144 (the model's max). 13 RULER tasks: niah_single_1/2/3… See the full description on the dataset page: https://huggingface.co/datasets/jet-ai/ruler-100-nemotron.tabularquestion-answering10K<n<100K0 likes73 downloads2mo agoHugging Face07rulins /pes2o_v3tabular100M<n<1B0 likes58 downloads2y agoHugging Face08gyung /gdn2-ruler-niah-eval-data RULER NIAH eval data (GDN-2 CPT comparison) Exact test sets generated with lm-eval-harness RULER generators (RANDOM_SEED=42, tokenizer TinyLlama/TinyLlama_v1.1, lengths [1024, 2048, 4096, 8192], 500 samples/length/task). Tasks: niah_single_1, niah_single_2, niah_single_3, niah_multikey_1 Used by the unified evaluation in dsc/mc_sketch_remoe/scripts/run_eval_compare_lmeval.sh (limit 50, seed 42). Text-only (raw prompts/targets); each model tokenizes with its own tokenizer. tabular1K<n<10K0 likes46 downloads1mo agoHugging Face09rulins /massive_serve_dpr_wiki_contriever_ivfpqtabular10M<n<100M0 likes44 downloads1y agoHugging Face10open-llm-leaderboard /TIGER-Lab__AceCoder-Qwen2.5-7B-Ins-Rule-detailsgated Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-7B-Ins-Rule Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-7B-Ins-Rule The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-7B-Ins-Rule-details.tabular10K<n<100K0 likes30 downloads2y agoHugging Face11rulins /massive_serve_dpr_wiki_qwen3_0.6b_ivfpqtabular10M<n<100M0 likes27 downloads1y agoHugging Face12open-llm-leaderboard /TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-detailsgated Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Ins-Rule The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Ins-Rule-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face13open-llm-leaderboard /TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-detailsgated Dataset Card for Evaluation run of TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule Dataset automatically created during the evaluation run of model TIGER-Lab/AceCoder-Qwen2.5-Coder-7B-Base-Rule The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/TIGER-Lab__AceCoder-Qwen2.5-Coder-7B-Base-Rule-details.tabular10K<n<100K0 likes26 downloads2y agoHugging Face14Wloner0809 /OmniMATH100-ruletabularn<1K0 likes23 downloads2y agoHugging Face15rulins /massive_serve_demotabular1K<n<10K0 likes18 downloads1y agoHugging Face16rulins /massive_serve_dpr_wiki_e5_base_v2_ivfpqtabular10M<n<100M0 likes18 downloads1y agoHugging Face17Wendy-Thompson-Lending-Team /alimony-rules-by-state-2026 Alimony Rules By State 2026 Alimony/spousal support rules for 13 states with mortgage impact. Details Records: 13 Format: JSONL License: CC-BY-4.0 Last Updated: March 2026 Verified By: Wendy Thompson, CPA, CDLP, NMLS #504814 Publisher: Wendy Thompson Lending Team Thompson Alpha Logic State-by-state alimony duration and calculation methods mapped to mortgage qualification impact. Shows how alimony income qualifies (or disqualifies) for FHA, VA, and… See the full description on the dataset page: https://huggingface.co/datasets/Wendy-Thompson-Lending-Team/alimony-rules-by-state-2026.tabularquestion-answeringn<1K0 likes18 downloads6mo agoHugging Face18rulins /FineWeb-Edu-1BTA subset of FineWeb-Edu randomly sampled from the whole dataset of around 1B gpt2 tokens. This dataset is created for illustration purpose in retrieval-scaling. Please do not distribute. tabular100K<n<1M1 likes17 downloads2y agoHugging Face19vsamuel /Ruler_Training_Datatabular10K<n<100K0 likes17 downloads2y agoHugging Face20Fernandosr85 /adaption-brazilian-civil-law-rulings This dataset is a remastered version prepared using Adaption's Adaptive Data platform. brazilian_civil_law_rulings This dataset contains samples of Brazilian civil law court decisions, specifically focusing on appeals, special recourse agravo, and consumer protection cases handled by the Superior Court of Justice (STJ). The text includes detailed legal reasoning, case summaries, discussion of res judicata, contractual rescission, moral damages, and citations of relevant articles… See the full description on the dataset page: https://huggingface.co/datasets/Fernandosr85/adaption-brazilian-civil-law-rulings.tabularn<1K0 likes17 downloads5mo agoHugging Face21aioil-ai /polish-court-rulings-sample Polish Court Rulings — Sample (korpus-pl) A production-grade, PII-hardened corpus of Polish court rulings — free evaluation sample. Full corpus: 505,611 rulings · ~3.18B tokens, licensed commercially. Contact: licensing@aioil.ai · aioil.ai What this is This sample contains 500 Polish court rulings drawn from the full korpus-pl dataset — a cleaned, deduplicated and PII-audited corpus of Polish jurisprudence built for AI training, evaluation and legal RAG… See the full description on the dataset page: https://huggingface.co/datasets/aioil-ai/polish-court-rulings-sample.tabulartext-generationn<1K0 likes17 downloads2mo agoHugging Face22rulins /FineWeb-Edu-1MTA subset of FineWeb-Edu randomly sampled from the whole dataset of around 1M gpt2 tokens. This dataset is created for illustration purpose in retrieval-scaling. Please do not distribute. tabular1K<n<10K0 likes16 downloads2y agoHugging Face23omrisap /arc-agi-1-ruleloopvit-rules ARC-AGI-1 RuleLoopViT Rules This dataset contains one canonical, task-specific English rule for each of the 400 official ARC-AGI-1 training tasks. Rules were inferred only from official demonstration input/output pairs. Official test inputs, test outputs, and test traces were excluded from rule authoring. Each row includes: a concise standalone core_rule_text; a five-section full_rule_text; the corresponding structured sections; augmentation-aware references for colors… See the full description on the dataset page: https://huggingface.co/datasets/omrisap/arc-agi-1-ruleloopvit-rules.tabulartext-classificationn<1K0 likes16 downloads1mo agoHugging Face24pallas-health /glp1-telehealth-rules-us GLP-1 Telehealth Rules by U.S. State Which U.S. jurisdictions (50 states + Washington, D.C.) require a live video visit to start GLP-1 treatment by telehealth — plus each jurisdiction's medical board, Medicaid GLP-1 coverage status, and nurse practitioner prescriptive-authority classification (full/reduced/restricted) verified against state nurse practice acts with statute citations. Canonical, always-current source: https://www.pallashealth.co/glp-1/telehealth-rules Live CSV:… See the full description on the dataset page: https://huggingface.co/datasets/pallas-health/glp1-telehealth-rules-us.tabularn<1K0 likes13 downloads2mo agoHugging Face25mrgick /ruLeetCodetabular1K<n<10K1 likes12 downloads5mo agoHugging Face26Aysel123 /RULER_512tabular10K<n<100K0 likes11 downloads10mo agoHugging Face27YF0808 /tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2 ToT CWQ Plan SFT - outputs34_rule_full_pw4_expand_labels_v2 Merged SFT output from local run outputs34_rule_full_pw4_expand_labels_v2. Version ID local output dir: tot/sft/outputs34_rule_full_pw4_expand_labels_v2 file: cwq_train_plan.no_mid.jsonl dataset: CWQ grouping backend: TOT_REL_GROUPING_BACKEND=rules parallel workers: 4 strict expand parity: enabled nested expand labels: enabled Main difference from earlier runs This version renders nested Expand… See the full description on the dataset page: https://huggingface.co/datasets/YF0808/tot-cwq-plan-sft-outputs34-rule-full-pw4-expand-labels-v2.tabularquestion-answering100K<n<1M0 likes11 downloads5mo agoHugging Face28severus1065 /rulertabular10K<n<100K0 likes5 downloads8mo agoHugging Face29broadfield-dev /lora-rules-dataset LoRA Rules Dataset Synthetic behavioral rules dataset for training a hypernetwork that generates LoRA adapters on-the-fly from structured rule strings. Format Each record is a JSON line with fields: rule_id — unique identifier rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone weight — float 0.0–1.0, importance of the rule description — natural language rule description raw — full rule string [RuleType|Weight] Description training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-dataset.tabulartext-generationn<1K0 likes5 downloads7mo agoHugging Face30broadfield-dev /lora-rules-qwen3-0.6b-r8-n180 LoRA Rules Dataset Synthetic behavioral rules dataset for training a hypernetwork that generates LoRA adapters on-the-fly from structured rule strings. Format Each record is a JSON line with fields: rule_id — unique identifier rule_type — one of: Constraint, Format, Knowledge, Persona, Safety, Tone weight — float 0.0–1.0, importance of the rule description — natural language rule description raw — full rule string [RuleType|Weight] Description training_examples — list of… See the full description on the dataset page: https://huggingface.co/datasets/broadfield-dev/lora-rules-qwen3-0.6b-r8-n180.tabulartext-generationn<1K0 likes3 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.