CoolFace
Modelpublic

Momix-44/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

sourceHugging Faceapache-2.0updated 6mo agoView on Hugging Face
2likes985downloads
Model Card

voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning is a 9B-scale model based on Qwen/Qwen3.5-9B. This model is designed to improve reasoning-oriented multiple-choice performance while preserving strong general capability.

In our zero-shot evaluation, the model achieves the best overall aggregate performance among the following three models:

  • —Qwen/Qwen3.5-9B
  • —DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT
  • —voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

The largest gains appear on ARC-Challenge, ARC-Easy, and BoolQ. These results suggest that the model improves structured reasoning and calibrated answer selection.


Model Summary

  • —Base model: Qwen/Qwen3.5-9B
  • —Model type: Causal language model
  • —Primary focus: Reasoning, multiple-choice QA, and general zero-shot evaluation
  • —Strengths: ARC, BoolQ, aggregate benchmark performance
  • —Trade-offs: Slightly weaker than some baselines on HellaSwag and OpenBookQA

Evaluation Setup

We compare three models under the same zero-shot setting:

  • —0-shot
  • —No few-shot examples
  • —Same benchmark suite
  • —Reported with standard error

Compared Models

  1. 1.Qwen/Qwen3.5-9B
  2. 2.DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT
  3. 3.voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

Main Results

Representative 7-task Average

We use acc_norm when available, and acc otherwise.

ModelAvg. Score
Qwen/Qwen3.5-9B0.7041
DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT0.6927
voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning0.7133

Macro Average over All 12 Reported Metrics

ModelMacro Avg.
Qwen/Qwen3.5-9B0.6655
DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT0.6587
voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning0.6749

These results indicate that voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning is the strongest overall model in this comparison.


Benchmark Results

TaskMetricQwen3.5-9BDavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCTvoidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning
arc_challengeacc0.54270.52050.5631
arc_challengeacc_norm0.55550.54690.5836
arc_easyacc0.81400.80180.8354
arc_easyacc_norm0.74330.73480.7950
boolqacc0.89270.78780.8792
hellaswagacc0.58270.60620.5882
hellaswagacc_norm0.78060.79440.7856
openbookqaacc0.32800.33600.3240
openbookqaacc_norm0.42800.45200.4260
piqaacc0.79050.79050.7949
piqaacc_norm0.80140.80360.7992
winograndeacc0.72690.72930.7245

Key Observations

Strengths

  • —The model achieves the best overall average score across the compared models.
  • —The model shows clear improvements on ARC-Challenge and ARC-Easy.
  • —The model strongly outperforms DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT on BoolQ.
  • —The gains are especially visible on reasoning-oriented benchmarks.

Trade-offs

  • —The model is not the top model on every benchmark.
  • —The model is slightly behind DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT on HellaSwag and OpenBookQA.
  • —The model is close to the baselines on PIQA and Winogrande.

Overall, the model improves the reasoning profile of the base model without uniformly dominating all commonsense tasks.


Interpretation

The benchmark pattern suggests that this model improves:

  • —structured answer selection
  • —reasoning-oriented multiple-choice QA
  • —calibration on science and reading-style benchmarks

At the same time, the gains are smaller on tasks that rely more heavily on narrative continuation or broad commonsense completion priors.

This behavior is consistent with a model that is optimized more toward reasoning quality than pure completion fluency.


Limitations

  • —The evaluation here is limited to a small set of common zero-shot benchmarks.
  • —Some benchmark differences are small and may fall within the reported standard error.
  • —The model should not be described as universally better on every task.
  • —Additional evaluations on instruction following, long-context reasoning, coding, multilingual performance, and open-ended generation are still needed.

Conclusion

voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning is a strong 9B reasoning-oriented model built on top of Qwen/Qwen3.5-9B.

In this comparison, it delivers:

  • —the best overall aggregate benchmark score
  • —the strongest ARC performance
  • —strong BoolQ performance
  • —competitive general capability on other zero-shot commonsense tasks

This makes it a good choice for users who care about reasoning-oriented zero-shot performance in a compact 9B model.


Raw Results

Qwen/Qwen3.5-9B

TaskMetricValueStderr
arc_challengeacc0.54270.0146
arc_challengeacc_norm0.55550.0145
arc_easyacc0.81400.0080
arc_easyacc_norm0.74330.0090
boolqacc0.89270.0054
hellaswagacc0.58270.0049
hellaswagacc_norm0.78060.0041
openbookqaacc0.32800.0210
openbookqaacc_norm0.42800.0221
piqaacc0.79050.0095
piqaacc_norm0.80140.0093
winograndeacc0.72690.0125

DavidAU/Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT

TaskMetricValueStderr
arc_challengeacc0.52050.0146
arc_challengeacc_norm0.54690.0145
arc_easyacc0.80180.0082
arc_easyacc_norm0.73480.0091
boolqacc0.78780.0072
hellaswagacc0.60620.0049
hellaswagacc_norm0.79440.0040
openbookqaacc0.33600.0211
openbookqaacc_norm0.45200.0223
piqaacc0.79050.0095
piqaacc_norm0.80360.0093
winograndeacc0.72930.0125

voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

TaskMetricValueStderr
arc_challengeacc0.56310.0145
arc_challengeacc_norm0.58360.0144
arc_easyacc0.83540.0076
arc_easyacc_norm0.79500.0083
boolqacc0.87920.0057
hellaswagacc0.58820.0049
hellaswagacc_norm0.78560.0041
openbookqaacc0.32400.0210
openbookqaacc_norm0.42600.0221
piqaacc0.79490.0094
piqaacc_norm0.79920.0093
winograndeacc0.72450.0126

<img src="https://raw.githubusercontent.com/axolotl-ai-cloud/axolotl/main/image/axolotl-badge-web.png" alt="Built with Axolotl" width="200" height="32"/> <details><summary>See axolotl config</summary>

axolotl version: 0.16.0.dev0

yaml
# Example config for RCCA-TR A+ (Reliability-Calibrated Conflict-Aware Trust-Region) fine-tuning
# A+ variant: only 1 model in GPU memory (active model)
# Prior = offline cache, EMA = drift buffer

base_model: Qwen/Qwen3.5-9B

plugins:
  - axolotl.integrations.rcca_tr.RCCATRPlugin
  - axolotl.integrations.liger.LigerPlugin

liger_rms_norm: true
liger_glu_activation: true

# Enable RCCA-TR trainer
rcca_tr_trainer: true

# Conflict score hyperparameters
rcca_tr_conflict_lambda1: 1.0      # weight for surprisal in conflict score
rcca_tr_conflict_lambda2: 0.5      # weight for margin-based conflict
rcca_tr_conflict_tau: 1.0          # temperature for conflict sigmoid

# Reliability score hyperparameters
rcca_tr_reliability_beta: 0.5      # balance between stability and evidence
rcca_tr_reliability_tau: 1.0       # temperature for reliability sigmoid

# Trust-region hyperparameters
rcca_tr_epsilon_min: 0.01          # minimum trust-region radius
rcca_tr_epsilon_max: 1.0           # maximum trust-region radius
rcca_tr_kl_lambda: 1.0             # Lagrange multiplier for KL penalty
rcca_tr_use_smooth_objective: true  # smooth g(r_t)*KL vs hinge

# Drift buffer (replaces EMA model)
rcca_tr_ema_decay: 0.999           # decay rate for drift buffer
rcca_tr_drift_gamma: 1.0           # drift → reliability scaling

# Prior cache (optional; omit to use fallback mode)
# rcca_tr_prior_cache_path: ./prior_cache/prior_cache.pt

# Dataset
datasets:
  - path: voidful/gemini-3.1-opus-4.6-reasoning-merged
    type: chat_template
    split: train

dataset_prepared_path: ./prepared_data/rcca_tr

chat_template: qwen3_5

# Training settings
sequence_len: 16384
sample_packing: true
pad_to_sequence_len: true

gradient_accumulation_steps: 4
micro_batch_size: 1
num_epochs: 3
optimizer: adamw_torch_fused
lr_scheduler: cosine
learning_rate: 2e-5

bf16: true
gradient_checkpointing: true
flash_attention: true

dataloader_num_workers: 0

deepspeed: deepspeed_configs/zero2.json

val_set_size: 0.05

save_strategy: epoch

output_dir: ./outputs/rcca-tr-fft

hub_model_id: voidful/Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning
push_to_hub: true
hub_strategy: end

log_on_each_node: false
logging_steps: 1

</details><br>

Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning

This model is a fine-tuned version of Qwen/Qwen3.5-9B on the voidful/gemini-3.1-opus-4.6-reasoning-merged dataset.

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • —learning_rate: 2e-05
  • —trainbatchsize: 1
  • —evalbatchsize: 1
  • —seed: 42
  • —distributed_type: multi-GPU
  • —num_devices: 80
  • —gradientaccumulationsteps: 4
  • —totaltrainbatch_size: 320
  • —totalevalbatch_size: 80
  • —optimizer: Use OptimizerNames.ADAMWTORCHFUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • —lrschedulertype: cosine
  • —training_steps: 6

Training results

Framework versions

  • —Transformers 5.3.0
  • —Pytorch 2.10.0+cu128
  • —Datasets 4.5.0
  • —Tokenizers 0.22.2