CoolFace
20 results

scaler

ScalerLab /JudgeBench JudgeBench: A Benchmark for Evaluating LLM-Based Judges 📃 [Paper] • 💻 [Github] • 🤗 [Dataset] • 🏆 [Leaderboard] JudgeBench is a benchmark aimed at evaluating LLM-based judges for objective correctness on challenging response pairs. For more information on how the response pairs are constructed, please see our paper. Data Instance and Fields This release includes two dataset splits. The gpt split includes 350 unique response pairs generated by GPT-4o and the claude… See the full description on the dataset page: https://huggingface.co/datasets/ScalerLab/JudgeBench.texttext-classificationn<1K12 likes1.1k downloads2y agoHugging Facenyu-visionx /scale-rae-data Scale RAE Data Project Page | Paper | Code This repository contains data associated with the paper "Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders". The dataset is used for training and evaluating Scale-RAE, a framework that investigates scaling Representation Autoencoders (RAEs) for large-scale, freeform text-to-image (T2I) generation. It includes data used for scaling RAE decoders beyond ImageNet, featuring web, synthetic, and text-rendering data… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/scale-rae-data.text-to-image3 likes688 downloads8mo agoHugging FaceFineEnvs /repo2rlenv-scaler Repo2RLEnv SCALER Generate instances from owned reasoning families with seeded procedural samplers, deterministic references and family-specific answer verification; generation uses no model calls. Contains 100 Harbor tasks generated with the owned scaler recipe in Repo2RLEnv. Browse the complete task bundles in Harbor Visualiser or open the task folders. Each folder is a runnable Harbor task: tasks/<task_id>/ ├── task.toml # Harbor configuration and provenance… See the full description on the dataset page: https://huggingface.co/datasets/FineEnvs/repo2rlenv-scaler.n<1K0 likes421 downloads1d agoHugging FaceClarusC64 /scale-regime-attribution-recognition-climate-v01 Dataset ClarusC64/scale-regime-attribution-recognition-climate-v01 This dataset tests one capability. Can a model keep explanations at the same scale as the signal. Core rule A claim must match the signal scale the observation window the evidence available If the input is weather scale do not claim climate proof If the input is local or regional do not claim global causes or global outcomes If the record is short do not declare regime shifts or permanent new… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/scale-regime-attribution-recognition-climate-v01.texttext-classificationn<1K0 likes17 downloads8mo agoHugging FaceMrLight /scale-reasoning-data-v3-nofilter-releasetext100K<n<1M0 likes11 downloads1y agoHugging FaceAnkGhosh /Business-Case-Scaler-Clusteringdocumentn<1K0 likes7 downloads1y agoHugging Face