fjmgAI/b1-R1-Zero-3B-GGUF
Fine-Tuned Model
`fjmgAI/b1-R1-Zero-3B-GGUF`
Base Model
`unsloth/qwen2.5-3b-instruct-unsloth-bnb-4bit`
Fine-Tuning Method
Fine-tuning was performed using [`unsloth`](https://github.com/unslothai/unsloth), an efficient fine-tuning framework optimized for low-resource environments and Huggingface's TRL library.
Dataset
[`Kukedlc/dpo-orpo-spanish-15k`](https://huggingface.co/datasets/Kukedlc/dpo-orpo-spanish-15k)
Description
A Spanish-language dataset containing 15,000 examples, designed for Direct Preference Optimization (DPO) or Outcome-Regularized Preference Optimization (ORPO).
Adaptation
The dataset was adapted to a reasoning-based format for GPRO, enhancing its ability to guide preference-based decision-making during fine-tuning. This adaptation ensures better alignment with instruction-following tasks in Spanish.
Fine-Tuning Details
- The model was trained using the GPRO algorithm, leveraging structured preference data to refine its response generation.
- The model was fine-tuned to maintain its 4-bit quantization (`bnb-4bit`) for memory efficiency while aligning its outputs with the characteristics of the Spanish dataset.
- The focus was on retaining the model's instructional abilities while improving its understanding and generation of Spanish text.
Purpose
This fine-tuned model is intended for Spanish-language applications that require efficient AI that follows instructions using a lightweight reasoning process.
- Developed by: fjmgAI
- License: apache-2.0
<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/> <img src="https://camo.githubusercontent.com/9585eb3e70c8138cbc0f73de7e970be4c668e957e45d16fc3ee6687fcc1da905/68747470733a2f2f68756767696e67666163652e636f2f64617461736574732f74726c2d6c69622f646f63756d656e746174696f6e2d696d616765732f7265736f6c76652f6d61696e2f74726c5f62616e6e65725f6461726b2e706e67" width="200"/>
