CoolFace
Modelpublic

vilarin/Llama-Qwen3-4B-RPG-gguf

sourceHugging Faceapache-2.0updated 9mo agoView on Hugging Face
1likes17downloads
Model Card

Llama-qwen3-4b-RPG

Llama-qwen3-4b-RPG is a fine-tuned variant of Qwen3-4B with facebook/research-plan-gen datasets Llama4-maverick generated, optimized for Research Plan Generation (RPG). The model is designed to generate structured, high-quality research plans for complex scientific and technical tasks across multiple domains.

It is trained by using Unsloth notebook


Key Features

  • —Research-aware generation Produces clear, structured research plans with goals, methodologies, evaluation criteria, and constraints.
  • —Two-stage training
  • —SFT warm-up for instruction following
  • —GRPO refinement using custom reward functions
  • —Multi-domain coverage
  • —Machine Learning
  • —ArXiv research
  • —PubMed / biomedical research
  • —Custom chat template Tailored specifically for research-planning tasks rather than generic chat.
  • —Long-form optimized Tuned for long context windows and coherent multi-section outputs.

Training Overview

Base Model

  • —Qwen3-4B

Dataset

  • —Research Plan Generation Dataset
  • —Source: facebook/research-plan-gen
  • —Structure:
  • —Goal
  • —Rubric
  • —Reference Solution

Training Strategy

  1. 1.Stage 1: Supervised Fine-Tuning (SFT)
  2. 2.Learns structured research-plan formatting
  3. 3.Aligns outputs with rubric-based expectations
  1. 1.Stage 2: GRPO Reinforcement Learning
  2. 2.Improves plan quality using reward functions
  3. 3.Encourages:
  4. 4.Completeness
  5. 5.Logical structure
  6. 6.Methodological rigor
  7. 7.Faithfulness to constraints

Intended Use Cases

  • —Automated research planning
  • —Scientific assistant systems
  • —Academic proposal drafting
  • —R&D ideation and experiment design
  • —LLM-based research agents

Limitations

  • —Not intended for factual verification or citation generation
  • —Outputs should be reviewed by domain experts
  • —Optimized for planning, not final paper writing

License

This model follows the license of its base model Qwen3-4B and the dataset used for training. Please review upstream licenses before commercial use.


Acknowledgements

  • —Qwen model family
  • —Facebook Research Plan Generation dataset
  • —Open-source RLHF / GRPO tooling

Citation

If you use this model in research or products, please cite appropriately.