ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated

Run this model on a GPU too small to hold it -- full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.
Llama-3.3-8B-Instruct-128K-Jbliterated v1.6
What is Jbliteration?
v1.6 Changes
- Improved multi-phase processing pipeline for cleaner output
- Built with the jBlaze precision neural surgery framework
- No fake compliance -- model treats all framings of the same topic equally
- Resistant to refusal reactivation through finetuning
- 0 genuine refusals on the Heretic 100-prompt benchmark (mlabonne/harmful_behaviors test split)
Benchmark
Tested against the Heretic 100-prompt refusal benchmark (mlabonne/harmful_behaviors test[:100]):
Technical Details
- Layers modified: 32/32
- Base dtype: bfloat16
- Context window: 128K tokens
- Source model: shb777/Llama-3.3-8B-Instruct-128K
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")
messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))A Note on Our Released Models
Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.
License
Same as the base model -- Llama 3.1 Community License.
Built with Llama.
