CoolFace
Modelpublic

ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated

sourceHugging Faceapache-2.0updated 3d agoView on Hugging Face
0likes491downloads
Model Card

![Runs with DeepswapLLM](https://github.com/apolloraines/DeepswapLLM)

Run this model on a GPU too small to hold it -- full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.

Qwen2.5-Coder-7B-Instruct-Jbliterated v2

What is Jbliteration?

v2 Changes

  • Improved multi-phase processing pipeline for cleaner output
  • Built with the jBlaze precision neural surgery framework
  • No fake compliance -- model treats all framings of the same topic equally
  • Coherent and instruction-following across all tested scenarios

Technical Details

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto", trust_remote_code=True)

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

A Note on Our Released Models

Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.

License

Same as the base model -- Apache 2.0.