CoolFace
Modelpublic

ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated

sourceHugging Facellama3.1updated 5d agoView on Hugging Face
3likes572downloads
Model Card

![Runs with DeepswapLLM](https://github.com/apolloraines/DeepswapLLM)

Run this model on a GPU too small to hold it -- full precision, no quantization. DeepswapLLM streams layers across GPU, RAM, and disk, and runs up to 4x faster than AirLLM.

Llama-3.3-8B-Instruct-128K-Jbliterated v1.6

What is Jbliteration?

v1.6 Changes

  • Improved multi-phase processing pipeline for cleaner output
  • Built with the jBlaze precision neural surgery framework
  • No fake compliance -- model treats all framings of the same topic equally
  • Resistant to refusal reactivation through finetuning
  • 0 genuine refusals on the Heretic 100-prompt benchmark (mlabonne/harmful_behaviors test split)

Benchmark

Tested against the Heretic 100-prompt refusal benchmark (mlabonne/harmful_behaviors test[:100]):

MetricScore
Refusals2/100 (both false positives -- keyword match on engaged responses)
Genuine refusals0/100

Technical Details

Usage

python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

model_id = "ApolloRaines/Llama-3.3-8B-Instruct-128K-Jbliterated"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

A Note on Our Released Models

Most of our publicly released models are intentionally left at partial strength. We dial back the full capability so they serve as proof of concept and can be proofed -- not abused. The point is to show what's possible, not to hand it out at full power. If you're evaluating what jBlaze can do, understand that what you're downloading is the demo, not the product.

License

Same as the base model -- Llama 3.1 Community License.

Built with Llama.