CoolFace
Modelpublic

RexTRO111/Qwen3-4B-MegaR3ASONER-v1

sourceHugging Faceotherupdated 3mo agoView on Hugging Face
1likes27downloads
Model Card

Qwen3-4B-MegaR3ASONER-v1

A full merged reasoning model created by merging the MegaR3ASONER LoRA into `Qwen/Qwen3-4B-Thinking-2507`.

Model type

This repository contains the complete merged Transformers model, not only the PEFT adapter. It can be loaded directly without PeftModel.

  • —Base model: Qwen/Qwen3-4B-Thinking-2507
  • —Merge dtype: BF16
  • —Adapter source: RexTRO111/Qwen3-4B-MegaR3ASONER-LoRA-v1
  • —Training hardware: NVIDIA A10G on Modal
  • —Merge hardware: NVIDIA A10G on Modal

Preliminary evaluation

On the first 100 examples selected by EleutherAI's gsm8k_cot task:

  • —Flexible extraction exact match: 88%
  • —Strict match: 83%

This was a limited 100-question run, not a full GSM8K score and not a controlled base-versus-fine-tune comparison.

Loading

python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "RexTRO111/Qwen3-4B-MegaR3ASONER-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
model.eval()

Limitations

  • —It may produce an incorrect intermediate thought before correcting itself.
  • —It can overthink simple prompts.
  • —Long reasoning traces increase latency and cost.
  • —Benchmark contamination has not been exhaustively ruled out.
  • —Verify answers before high-stakes use.

Licensing note

The Qwen base model and every training dataset retain their own licenses and upstream terms. Review all applicable terms before redistribution or commercial use.