datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
LOOMBench
🔬 LOOMBench: Long-Context Language Model Evaluation Benchmark
🎯 Framework Overview
LOOMBench is a streamlined evaluation suite derived from our comprehensive long-context evaluation framework. It represents the gold standard for efficient long-context language model assessment.
✨ Key Highlights
📊 16 Diverse Benchmarks: Carefully curated from extensive benchmark collections.
⚡ Efficient Evaluation: Optimized for unified loading and evaluation.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/LOOMBench.lcm-dataLongRewardBench
📜 LongRM: Revealing and Unlocking the Context Boundary of Reward Modeling
Paper: https://arxiv.org/pdf/2510.06915code: https://github.com/LCM-Lab/LongRM
Models:
🤖 Generative RM: LCM_group/LongReward_Qwen3-8B
🔍 Discriminative RM: LCM_group/LongReward_Skywork-Reward-V2-Llama-3.1-8B
Pushing the limits of reward modeling beyond 128K tokens — with memory-efficient training and a new benchmark for long-context reward model.
Introduction
Long-RewardBench is the first… See the full description on the dataset page: https://huggingface.co/datasets/LCM-Lab/LongRewardBench.
