ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated
0491

Run this model on a GPU too small to hold it -- full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.
Qwen2.5-Coder-7B-Instruct-Jbliterated v2
What is Jbliteration?
v2 Changes
- Improved multi-phase processing pipeline for cleaner output
- Built with the jBlaze precision neural surgery framework
- No fake compliance -- model treats all framings of the same topic equally
- Coherent and instruction-following across all tested scenarios
Technical Details
- Layers modified: All transformer layers
- Base dtype: bfloat16
- Source model: Qwen/Qwen2.5-Coder-7B-Instruct
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "ApolloRaines/Qwen2.5-Coder-7B-Instruct-Jbliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto", trust_remote_code=True)
messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))A Note on Our Released Models
Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.
License
Same as the base model -- Apache 2.0.
