CoolFace
Modelpublic

mmnga/Mixtral-Fusion-4x7B-Instruct-v0.1

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
18likes77downloads
Model Card

Model Card for Mixtral-Fusion-4x7B-Instruct-v0.1

This model is an experimental model created by merging mistralai/Mixtral-8x7B-Instruct-v0.1 experts.

How we merged experts

Changed to merge using slerp. Discussion

old merge version ~~We simply take the average of every two experts.weight.~~ ~~The same goes for gate.weight.~~

How To Convert

use colab cpu-high-memory. convert_mixtral_8x7b_to_4x7b.ipynb

OtherModels

mmnga/Mixtral-Extraction-4x7B-Instruct-v0.1

Usage

~~~python pip install git+https://github.com/huggingface/transformers --upgrade pip install torch accelerate bitsandbytes flash_attn ~~~

~~~python from transformers import AutoTokenizer, AutoModelForCausalLM, MixtralForCausalLM import torch

modelnameor_path = "mmnga/Mixtral-Fusion-4x7B-Instruct-v0.1"

tokenizer = AutoTokenizer.frompretrained(modelnameorpath) model = MixtralForCausalLM.frompretrained(modelnameorpath, loadin8bit=True)

text = "[INST] What was John Holt's vision on education? [/INST] " inputs = tokenizer(text, return_tensors="pt")

outputs = model.generate(**inputs, maxnewtokens=128) print(tokenizer.decode(outputs[0], skipspecialtokens=True))

~~~