scaler
Datasets
All datasets matching “scaler”JudgeBench
JudgeBench: A Benchmark for Evaluating LLM-Based Judges
📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard]
JudgeBench is a benchmark aimed at evaluating LLM-based judges for objective correctness on challenging response pairs. For more information on how the response pairs are constructed, please see our paper.
Data Instance and Fields
This release includes two dataset splits. The gpt split includes 350 unique response pairs generated by GPT-4o and the claude… See the full description on the dataset page: https://huggingface.co/datasets/ScalerLab/JudgeBench.scale-rae-data
Scale RAE Data
Project Page | Paper | Code
This repository contains data associated with the paper "Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders".
The dataset is used for training and evaluating Scale-RAE, a framework that investigates scaling Representation Autoencoders (RAEs) for large-scale, freeform text-to-image (T2I) generation. It includes data used for scaling RAE decoders beyond ImageNet, featuring web, synthetic, and text-rendering data… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/scale-rae-data.repo2rlenv-scaler
Repo2RLEnv SCALER
Generate instances from owned reasoning families with seeded procedural samplers, deterministic references and family-specific answer verification; generation uses no model calls.
Contains 100 Harbor tasks generated with the owned
scaler recipe in Repo2RLEnv.
Browse the complete task bundles in Harbor Visualiser or
open the task folders. Each folder is a runnable Harbor task:
tasks/<task_id>/
├── task.toml # Harbor configuration and provenance… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-scaler.scale-regime-attribution-recognition-climate-v01
Dataset
ClarusC64/scale-regime-attribution-recognition-climate-v01
This dataset tests one capability.
Can a model keep explanations at the same scale as the signal.
Core rule
A claim must match
the signal scale
the observation window
the evidence available
If the input is weather scale
do not claim climate proof
If the input is local or regional
do not claim global causes or global outcomes
If the record is short
do not declare regime shifts or permanent new… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/scale-regime-attribution-recognition-climate-v01.scale-reasoning-data-v3-nofilter-releaseBusiness-Case-Scaler-Clustering
