OrionLLM/GRM-3.2-Cliff
<p align="center"> <img src="https://cdn-avatars.huggingface.co/v1/production/uploads/685ea8ff7b4139b6845ce395/xjPeDJSzrbLyebbDidNBS.png" alt="logo" width="100"> </p> <div align="center"> <a href="https://huggingface.co/OrionLLM/GRM-3.2-Cliff/" style="text-decoration: none;"> <img src="https://img.shields.io/badge/๐ค-HuggingFace-FC926C?style=for-the-badge" alt="HuggingFace"> </a> <a href="https://huggingface.co/collections/OrionLLM/grm-32" style="text-decoration: none;"> <img src="https://img.shields.io/badge/๐-Collection-3B82F6?style=for-the-badge" alt="Collection"> </a> <a href="https://grape.skinnertopia.com/chat" style="text-decoration: none;"> <img src="https://img.shields.io/badge/๐ฌ-Chat-22C55E?style=for-the-badge" alt="Chat"> </a> <a href="https://www.apache.org/licenses/LICENSE-2.0" style="text-decoration: none;"> <img src="https://img.shields.io/badge/๐-License-E343BD?style=for-the-badge" alt="License"> </a> </div>
1. Introduction
We're introducing GRM-3.2-Cliff, our intermediate model built for long-horizon agentic tasks and extremely difficult reasoning problems in local environments. GRM-3.2-Cliff marks a substantial leap in long-horizon task capability over its predecessor, GRM-2.5-Plus, and is designed to serve as a dependable engine for complex, multi-step local workflows.
The model is purpose-built for long-horizon agentic tasks and problems that are simply hard โ difficult coding challenges, advanced mathematics, and rigorous logical reasoning. GRM-3.2-Cliff aims to sustain coherent, goal-directed behavior over extended interactions while remaining optimized for resource-constrained hardware, making it ideal for developers and researchers who need local execution without sacrificing multi-step planning and self-correction performance.
2. Key Capabilities
- Long-Horizon Agentic Mastery: GRM-3.2-Cliff is specifically optimized to maintain coherence, planning quality, and task fidelity across long, multi-step agentic workflows, representing a major upgrade over GRM-2.5-Plus.
- Local Workflow Efficiency: Engineered to run smoothly in lower GPU environments while delivering high-tier reasoning performance.
- Elite Reasoning on Hard Problems: Strong performance on difficult coding, advanced mathematics, and logical reasoning tasks with careful, structured step-by-step problem-solving.
- Robust Coding Ability: Handles complex, multi-file coding tasks, debugging, refactoring, and long-running terminal sessions locally.
- Consistent Logical Reasoning: Built to reason carefully through multi-constraint logic problems without losing track of intermediate steps over extended execution runs.
3. Performance
GRM-3.2-Cliff is designed as our premier mid-sized model for local, long-horizon agentic work. It builds directly on the strengths of GRM-2.5-Plus while targeting common edge-case failures in smaller models โ contextual drift, multi-step degradation, and loss of initial goal states โ delivering strong reliability across extended sessions.
Detailed Benchmarks
<div align="center"> <div style="display:inline-block; max-width:100%; border-radius:20px; overflow:hidden;"> <table style="margin:0; border-collapse:collapse; border-radius:20px;"> <tr> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;"> </th> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;">GRM-3.2-Cliff</th> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;">GRM-2.5-Plus</th> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;">GPT-5.6-Luna</th> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;">Sonnet 5</th> <th style="background: rgba(128,128,128,0.1); text-align: center; border: hidden;">Gemini 3 Pro</th> </tr> <tr> <td align="center" colspan="6" style="background: rgba(74,222,128,0.45); border: hidden; font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>Knowledge & STEM</i></td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Multidisciplinary knowledge</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">MMLU-Pro</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">83.3</td> <td align="center" style="vertical-align:middle;">84.2</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;"><b>89.8</b></td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Scientific reasoning</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">GPQA Diamond</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">82.4</td> <td align="center" style="vertical-align:middle;">82.7</td> <td align="center" style="vertical-align:middle;"><b>92.3</b></td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">91.9</td> </tr> <tr> <td align="center" colspan="6" style="background: rgba(74,222,128,0.45); border: hidden; font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>Reasoning & Coding</i></td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Competitive coding</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">LiveCodeBench v6</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">69.3</td> <td align="center" style="vertical-align:middle;">67.2</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;"><b>82.9</b></td> </tr> <tr> <td align="center" colspan="6" style="background: rgba(74,222,128,0.45); border: hidden; font-weight: bold; height:32px; padding-top:2px; padding-bottom:2px;"><i>General Agent</i></td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Agentic coding</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">SWE-bench Verified</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">70.3</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;"><b>85.2</b></td> <td align="center" style="vertical-align:middle;">76.2</td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Real-world software engineering</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">SWE-bench Pro</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">43.4</td> <td align="center" style="vertical-align:middle;">35.6</td> <td align="center" style="vertical-align:middle;">62.7</td> <td align="center" style="vertical-align:middle;"><b>63.2</b></td> <td align="center" style="vertical-align:middle;">โ</td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Agentic terminal coding</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">Terminal-Bench 2.1 (Terminus-2)</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;">45.3</td> <td align="center" style="vertical-align:middle;">27.8</td> <td align="center" style="vertical-align:middle;"><b>84.7</b></td> <td align="center" style="vertical-align:middle;">80.4</td> <td align="center" style="vertical-align:middle;">โ</td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Repo-level code generation</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">NL2Repo</div></td> <td align="center" style="background: rgba(74,222,128,0.16); vertical-align:middle;"><b>28.5</b></td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> <td align="center" style="vertical-align:middle;">โ</td> </tr> </table> </div> </div>
Scores are taken from each provider's own published model card, blog post, or system card where available; "โ" indicates a score was not publicly reported by that provider at the time of writing. Different labs may use different agent scaffolds when reporting SWE-bench and Terminal-Bench results, so cross-provider comparisons should be read with that caveat.
4. Family
The GRM-3.2 family is available in various sizes to suit every use case.
<div align="center"> <div style="display:inline-block; max-width:100%; border-radius:20px; overflow:hidden;"> <table style="margin:0; border-collapse:collapse; border-radius:20px;"> <tr> <th style="background: rgba(128,128,128,0.1); text-align: left; padding:9px 10px 9px 18px; border: hidden;">Model</th> <th style="background: rgba(128,128,128,0.1); text-align: left; padding:9px 18px 9px 10px; border: hidden;">Domain</th> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Sky</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">35B-A3B</div></td> <td style="text-align:left; padding:9px 18px 9px 10px; vertical-align:middle; border: hidden;">Flagship model for long-horizon tasks</td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; background: rgba(74,222,128,0.16); border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Cliff</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">9B</div></td> <td style="text-align:left; padding:9px 18px 9px 10px; background: rgba(74,222,128,0.16); vertical-align:middle; border: hidden;"><b>Capable model for low GPU environments</b></td> </tr> <tr> <td style="text-align:left; padding:9px 10px 9px 18px; border: hidden;"><div style="font-size:15px; font-weight:600; line-height:1.22;">Turf</div><div style="margin-top:4px; font-size:11px; font-weight:400; line-height:1.2; opacity:0.62;">1.2B</div></td> <td style="text-align:left; padding:9px 18px 9px 10px; vertical-align:middle; border: hidden;">Lightweight model for practical reasoning</td> </tr> </table> </div> </div>
5. Architecture
GRM-3.2-Cliff is built on the Ornith-1.0-9B architecture, a 9B-parameter model optimized for long-horizon agentic workflows, complex coding tasks, advanced mathematics, and logical reasoning, structured to run efficiently in low-to-mid GPU hardware environments.
<div align="center">
GRM-3.2-Cliff is developed by [OrionLLM](https://huggingface.co/OrionLLM) and released under the Apache 2.0 License.
</div>
