datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RULER-8192-Qwen2.5-3B-tokenizerRULER-llama3-1M
RULER-Llama3-1M
A 1M token version of the RULER dataset based on the Llama-3 chat template.
It is automatically generated based on the scripts available in the RULER repository: https://github.com/NVIDIA/RULER. It is designed for evaluating the performance of Long Language Models (LLMs) on various tasks with varying sequence lengths.
How to Use
from datasets import load_dataset
LENGTH_IN_STRING = ['4k', '8k', '16k', '32k', '64k', '128k', '256k', '512k', '1M']
TASKS =… See the full description on the dataset page: https://huggingface.co/datasets/self-long/RULER-llama3-1M.ruler-300-seed42
Frozen RULER 300, seed 42
This dataset freezes the exact RULER inputs used by the
short-long-pretraining native evaluation suite.
Repository: bicycleman15/ruler-300-seed42
Rows: 6,300
Tasks: s-niah-1, s-niah-2, s-niah-3, mk1, mk2, mv, mq
Context lengths: 1024, 2048, 4096
Samples per task/length: 300
Seed: 42
Dataset SHA-256: 4d82df6f9b1f2d9c45c0a0bda8c734032e62f517b746c6351bf9c2f38335ab3d
Tokenizer SHA-256: 1f186971e25f7bda3dd6f93a100bb8fa2a6801cf8dc3807c8a8c4e45f296ab90… See the full description on the dataset page: https://huggingface.co/datasets/bicycleman15/ruler-300-seed42.RULER-32768-Qwen2.5-3B-tokenizerRUListening
RUListening: Building Perceptually-Aware Music-QA Benchmarks
Multimodal LLMs, particularly Large Audio Language Models (LALMs), have shown progress in music understanding tasks due to text-only LLM initialization. However, we find that seven of the top ten Music Question Answering (Music-QA) models are text-only models, suggesting these benchmarks rely on reasoning rather than audio perception. To address this limitation, we present RUListening: Robust Understanding through… See the full description on the dataset page: https://huggingface.co/datasets/yongyizang/RUListening.AIRBOT_MMK2_place_the_umbrella_and_the_ruler
AIRBOT_MMK2_place_the_umbrella_and_the_ruler
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_umbrella_and_the_ruler.Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-1024.statml-arxivRULER-luciole_tokenizer_128k-arab-regional_v2nasa-cmapss-rul
Modified CMAPSS Dataset (Turbofan Engine Degradation)
📘 Description
This dataset is a modified version of the NASA C-MAPSS (Commercial Modular Aero-Propulsion System Simulation) turbofan engine degradation simulation dataset. The modification was created by our team as part of a submission for RISTEK UI Datathon 2025, in conjunction with the predictive modeling work we developed.
Each entry in this dataset corresponds to one engine's operating cycle. Engines begin with… See the full description on the dataset page: https://huggingface.co/datasets/penikmatrumput/nasa-cmapss-rul.4B-predict.rule-r-1.0-k-256.L-1024.statml-arxivruler_qwenswiss_rulings
Dataset Card for Swiss Rulings
Dataset Summary
SwissRulings is a multilingual, diachronic dataset of 637K Swiss Federal Supreme Court (FSCS) cases. This dataset can be used to pretrain language models on Swiss legal data.
Supported Tasks and Leaderboards
Languages
Switzerland has four official languages with three languages German, French and Italian being represenated. The decisions are written by the judges and clerks in the language of the… See the full description on the dataset page: https://huggingface.co/datasets/rcds/swiss_rulings.Qwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-128.L-512.statml-arxivRULER-32768-llama-3.1-tokenizer-chat-templateQwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-64.L-256.statml-arxivQwen3-4B-Instruct-2507.rule-thoughtful-except-first.k-128.L-1024.statml-arxivamc-ruler-qwen35-32k
AMC RULER 32k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 32,768 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-32k.amc-ruler-qwen35-16k
AMC RULER 16k
This dataset contains frozen inputs for the RULER benchmark.
Generation metadata
Benchmark: RULER
Sequence length: 16,384 tokens
Tokenizer: Qwen/Qwen3.5-9B
Tokenizer revision: c202236
lm-eval version: 0.4.12
Task configurations: 13
Samples per configuration: 500
Deterministic generation: Yes. Each configuration resets Python, NumPy, and task random state to seed 42.
Task configurations
niah_single_1
niah_single_2
niah_single_3… See the full description on the dataset page: https://huggingface.co/datasets/khashazad/amc-ruler-qwen35-16k.RULER-262144-gemma3-instructlm-eval-ruler-results-private-32K
Dataset Card for Evaluation run of elichen3051/Llama-3.1-8B-GGUF
Dataset automatically created during the evaluation run of model elichen3051/Llama-3.1-8B-GGUF
The dataset is composed of 12 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/elichen-skymizer/lm-eval-ruler-results-private-32K.pause.rule-r-1.0-k-8.L-128.statml-arxivRULER-32768-Qwen-3-InstructRULER-16384-Qwen-34B-ranked-v7.rule-stride-train4-test32.k-8.L-4096.statml-arxivRULER-8192-Qwen-3RULERmultimodality-poc-llama31-ruler16k
Multimodality PoC corpus — Llama-3.1-8B-Instruct on RULER-16K
Raw pre-RoPE query and hidden-state tensors captured during prefill, used
to study whether the per-(layer, kv_head) query distribution is unimodal
Gaussian (the assumption underpinning Expected Attention's MGF closed-form
in kvpress).
What's in here
65 .npz files, one per (RULER task, prompt_index) pair (13 tasks × 5
prompts).
Each file (~414 MB) contains:
field
dtype
shape
meaning
hidden
float16… See the full description on the dataset page: https://huggingface.co/datasets/June30916/multimodality-poc-llama31-ruler16k.rulerThis is a synthetic dataset generated using 📏 RULER: What’s the Real Context Size of Your Long-Context Language Models?.
It can be used to evaluate long-context language models with configurable sequence length and task complexity.
Currently, It includes 4 tasks from RULER:
QA2 (hotpotqa after adding distracting information)
Multi-hop Tracing: Variable Tracking (VT)
Aggregation: Common Words (CWE)
Multi-keys Needle-in-a-haystack (NIAH)
For each of the task, two target sequence lengths are… See the full description on the dataset page: https://huggingface.co/datasets/rbiswasfc/ruler.RULER-16384-llama-3.2-tokenizerRULER-65536-Qwen-3
