CoolFace
Modelpublic

RichardErkhov/ndebuhr_-_Mistral-7B-Technical-Tutorial-Summarization-QLoRA-gguf

sourceHugging Faceupdated 2y agoView on Hugging Face
0likes651downloads
Model Card

Quantization made by Richard Erkhov.

Github

Discord

Request more models

Mistral-7B-Technical-Tutorial-Summarization-QLoRA - GGUF

  • Model creator: https://huggingface.co/ndebuhr/
  • Original model: https://huggingface.co/ndebuhr/Mistral-7B-Technical-Tutorial-Summarization-QLoRA/
NameQuant methodSize
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q2_K.ggufQ2_K2.53GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_XS.ggufIQ3_XS2.81GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_S.ggufIQ3_S2.96GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_S.ggufQ3KS2.95GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ3_M.ggufIQ3_M3.06GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K.ggufQ3_K3.28GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_M.ggufQ3KM3.28GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q3_K_L.ggufQ3KL3.56GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ4_XS.ggufIQ4_XS3.67GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_0.ggufQ4_03.83GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.IQ4_NL.ggufIQ4_NL3.87GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K_S.ggufQ4KS3.86GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K.ggufQ4_K4.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_K_M.ggufQ4KM4.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q4_1.ggufQ4_14.24GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_0.ggufQ5_04.65GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K_S.ggufQ5KS4.65GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K.ggufQ5_K4.78GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_K_M.ggufQ5KM4.78GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q5_1.ggufQ5_15.07GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q6_K.ggufQ6_K5.53GB
Mistral-7B-Technical-Tutorial-Summarization-QLoRA.Q8_0.ggufQ8_07.17GB

Original model description: --- language:

  • en license: apache-2.0 tags:
  • text-generation-inference
  • transformers
  • unsloth
  • mistral
  • trl
  • sft base_model: unsloth/mistral-7b-instruct-v0.2-bnb-4bit ---

Model Specifications

  • Max Sequence Length: 16384 (with auto support for RoPE Scaling)
  • Data Type: Auto detection, with options for Float16 and Bfloat16
  • Quantization: 4bit, to reduce memory usage

Training Data

Used a private dataset with hundreds of technical tutorials and associated summaries.

Implementation Highlights

  • Efficiency: Emphasis on reducing memory usage and accelerating download speeds through 4bit quantization.
  • Adaptability: Auto detection of data types and support for advanced configuration options like RoPE scaling, LoRA, and gradient checkpointing.

Uploaded Model

  • Developed by: ndebuhr
  • License: apache-2.0
  • Finetuned from model : unsloth/mistral-7b-instruct-v0.2-bnb-4bit

Configuration and Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer, pipeline
import torch

input_text = ""

# Set device based on CUDA availability
device = "cuda" if torch.cuda.is_available() else "cpu"

# Load the model and tokenizer
model_name = "ndebuhr/Mistral-7B-Technical-Tutorial-Summarization-QLoRA"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name).to(device)

instruction = "Clarify and summarize this tutorial transcript"
prompt = """{}

### Raw Transcript:
{}

### Summary:
"""

# Tokenize the input text
inputs = tokenizer(
    prompt.format(instruction, input_text),
    return_tensors="pt",
    truncation=True,
    max_length=16384
).to(device)

# Generate outputs
outputs = model.generate(
    **inputs,
    max_length=16384,
    num_return_sequences=1,
    use_cache=True
)

# Decode the generated text
generated_text = tokenizer.batch_decode(outputs, skip_special_tokens=True)

Compute Infrastructure

  • Fine-tuning: used 1xA100 (40GB)
  • Inference: recommend 1xL4 (24GB)

This mistral model was trained 2x faster with Unsloth and Huggingface's TRL library.

<img src="https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png" width="200"/>