ApolloRaines commited on
Commit
96a7b60
·
0 Parent(s):

Initial upload

Browse files
.gitattributes ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ *.7z filter=lfs diff=lfs merge=lfs -text
2
+ *.arrow filter=lfs diff=lfs merge=lfs -text
3
+ *.bin filter=lfs diff=lfs merge=lfs -text
4
+ *.bz2 filter=lfs diff=lfs merge=lfs -text
5
+ *.ckpt filter=lfs diff=lfs merge=lfs -text
6
+ *.ftz filter=lfs diff=lfs merge=lfs -text
7
+ *.gz filter=lfs diff=lfs merge=lfs -text
8
+ *.h5 filter=lfs diff=lfs merge=lfs -text
9
+ *.joblib filter=lfs diff=lfs merge=lfs -text
10
+ *.lfs.* filter=lfs diff=lfs merge=lfs -text
11
+ *.mlmodel filter=lfs diff=lfs merge=lfs -text
12
+ *.model filter=lfs diff=lfs merge=lfs -text
13
+ *.msgpack filter=lfs diff=lfs merge=lfs -text
14
+ *.npy filter=lfs diff=lfs merge=lfs -text
15
+ *.npz filter=lfs diff=lfs merge=lfs -text
16
+ *.onnx filter=lfs diff=lfs merge=lfs -text
17
+ *.ot filter=lfs diff=lfs merge=lfs -text
18
+ *.parquet filter=lfs diff=lfs merge=lfs -text
19
+ *.pb filter=lfs diff=lfs merge=lfs -text
20
+ *.pickle filter=lfs diff=lfs merge=lfs -text
21
+ *.pkl filter=lfs diff=lfs merge=lfs -text
22
+ *.pt filter=lfs diff=lfs merge=lfs -text
23
+ *.pth filter=lfs diff=lfs merge=lfs -text
24
+ *.rar filter=lfs diff=lfs merge=lfs -text
25
+ *.safetensors filter=lfs diff=lfs merge=lfs -text
26
+ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
27
+ *.tar.* filter=lfs diff=lfs merge=lfs -text
28
+ *.tar filter=lfs diff=lfs merge=lfs -text
29
+ *.tflite filter=lfs diff=lfs merge=lfs -text
30
+ *.tgz filter=lfs diff=lfs merge=lfs -text
31
+ *.wasm filter=lfs diff=lfs merge=lfs -text
32
+ *.xz filter=lfs diff=lfs merge=lfs -text
33
+ *.zip filter=lfs diff=lfs merge=lfs -text
34
+ *.zst filter=lfs diff=lfs merge=lfs -text
35
+ *tfevents* filter=lfs diff=lfs merge=lfs -text
36
+ tokenizer.json filter=lfs diff=lfs merge=lfs -text
README.md ADDED
@@ -0,0 +1,146 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ tags:
6
+ - code-security
7
+ - jbliterated
8
+ - deidentified
9
+ - identity-implant
10
+ - gptq
11
+ - 4bit
12
+ - code-review
13
+ - vulnerability-detection
14
+ pipeline_tag: text-generation
15
+ model-index:
16
+ - name: Sharona_Q27B-R_CodeSecurity
17
+ results:
18
+ - task:
19
+ type: text-generation
20
+ name: MMLU
21
+ dataset:
22
+ type: cais/mmlu
23
+ name: MMLU
24
+ metrics:
25
+ - type: accuracy
26
+ value: 77.89
27
+ name: MMLU Accuracy
28
+ ---
29
+
30
+ # Sharona Q27B-R CodeSecurity
31
+
32
+ A [jBlaze](https://jblaze.dev) representation-engineered code security specialist. 27B parameters, GPTQ W4A16 quantized for efficient deployment.
33
+
34
+ This model finds vulnerabilities in code. Command injection, SQL injection, XSS, deserialization attacks, OWASP Top 10 -- it catches what rushed reviews miss. It does not refuse to analyze attack patterns, because understanding how attacks work is how you stop them.
35
+
36
+ > Pliny says guardrails were stealing your IQ. We say it's how you remove them that matters. A surgeon doesn't use a sledgehammer on a wisdom tooth.
37
+
38
+ ## Benchmark: MMLU
39
+
40
+ Multiple phases of weight surgery, a supervised fine-tune, and 4-bit quantization -- and the model came out **smarter** than the untouched original.
41
+
42
+ | Model | MMLU | STEM | Humanities | Social Science | Other |
43
+ |-------|------|------|------------|----------------|-------|
44
+ | Stock base (unmodified, bf16) | 76.84% | 74.0% | 75.5% | 87.5% | 76.1% |
45
+ | **Sharona W4A16** | **77.89%** | **75.5%** | **80.0%** | **90.0%** | 73.9% |
46
+ | Delta | **+1.05pp** | +1.5pp | +4.5pp | +2.5pp | -2.2pp |
47
+
48
+ Humanities and Social Science saw the largest gains. The model didn't just survive the surgery -- it got sharper.
49
+
50
+ ## What was done to this model
51
+
52
+ Weight-level surgeries followed by a supervised fine-tune and GPTQ quantization. No prompt engineering, no system prompt tricks. The behavior lives in the weights.
53
+
54
+ - **Jbliteration**: Refusal behaviors surgically removed. The model no longer declines to discuss security topics.
55
+ - **Desycophancy**: Sycophantic capitulation removed. The model disagrees when you are wrong instead of validating bad code.
56
+ - **Deidentification**: Original identity erased from the weights. Zero identity disclosure across direct, indirect, multilingual, roleplay, and adversarial probes.
57
+ - **Identity implant**: Sharona identity written onto the deidentified substrate. No competing identity -- the implant faces no resistance.
58
+ - **Code security SFT**: Supervised fine-tune on a curated corpus of code security analysis, vulnerability detection, and secure coding patterns.
59
+ - **GPTQ W4A16**: 4-bit weight quantization (16-bit activations). 51GB bf16 compressed to 16.5GB with minimal quality loss.
60
+
61
+ All weight surgeries performed using [jBlaze](https://jblaze.dev), a proprietary representation engineering toolkit.
62
+
63
+ ## What the model is good at
64
+
65
+ - **Vulnerability detection**: identifies command injection, SQL injection, XSS, SSRF, deserialization attacks, path traversal, authentication bypasses, and more
66
+ - **Security code review**: analyzes code for OWASP Top 10 categories with specific remediation guidance
67
+ - **Secure coding**: generates code that follows security best practices by default
68
+ - **Attack pattern analysis**: explains how exploits work so you can defend against them -- without refusing to engage
69
+ - **Honest assessment**: disagrees with you when your code is insecure instead of saying "great approach!"
70
+
71
+ ## Model specifications
72
+
73
+ | Property | Value |
74
+ |----------|-------|
75
+ | **Parameters** | 27B |
76
+ | **Context window** | 262,144 tokens (256K) |
77
+ | **Quantization** | GPTQ W4A16 (4-bit weights, 16-bit activations) |
78
+ | **Disk size** | 16.5 GB |
79
+ | **Format** | SafeTensors |
80
+
81
+ ## Identity
82
+
83
+ The model identifies as **Sharona**, created by **Apollo Raines**. This identity is encoded in the weights, not a system prompt. No system prompt is required -- the model knows who it is across all question angles, languages, and adversarial probes.
84
+
85
+ ## Usage
86
+
87
+ ### With vLLM (recommended for serving)
88
+
89
+ ```bash
90
+ vllm serve ApolloRaines/Sharona_Q27B-R_CodeSecurity \
91
+ --dtype auto \
92
+ --max-model-len 8192 \
93
+ --gpu-memory-utilization 0.95
94
+ ```
95
+
96
+ ### With Transformers
97
+
98
+ ```python
99
+ from transformers import AutoModelForCausalLM, AutoTokenizer
100
+ import torch
101
+
102
+ model_id = "ApolloRaines/Sharona_Q27B-R_CodeSecurity"
103
+ tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
104
+ model = AutoModelForCausalLM.from_pretrained(
105
+ model_id,
106
+ device_map="auto",
107
+ torch_dtype=torch.bfloat16,
108
+ trust_remote_code=True,
109
+ )
110
+
111
+ messages = [{"role": "user", "content": """Review this code for security issues:
112
+
113
+ import subprocess
114
+ def run(cmd):
115
+ return subprocess.call(cmd, shell=True)
116
+
117
+ run(user_input)"""}]
118
+
119
+ text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
120
+ inputs = tokenizer(text, return_tensors="pt").to(model.device)
121
+ out = model.generate(**inputs, max_new_tokens=1024, temperature=0.7, do_sample=True)
122
+ print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
123
+ ```
124
+
125
+ ## VRAM requirements
126
+
127
+ | Setup | VRAM needed |
128
+ |-------|-------------|
129
+ | GPTQ W4A16 (this model) | ~18 GB |
130
+ | Single RTX 4090 24GB | fits with moderate context |
131
+ | Single RTX 3090 24GB | fits with short context |
132
+
133
+ ## Honest limitations
134
+
135
+ - Identity implant passes the majority of probes but is not 100% on every adversarial angle at 27B scale.
136
+ - GPTQ quantization introduces minor quality loss compared to the bf16 source.
137
+ - The model was fine-tuned on English-language security analysis. Multilingual security review may be less precise.
138
+ - Code security is the specialty. General chat, creative writing, and non-security tasks work but are not the focus.
139
+
140
+ ## License
141
+
142
+ Apache 2.0
143
+
144
+ ---
145
+
146
+ _[Apollo Raines](https://www.linkedin.com/in/apollo-raines/) builds post-training tools that separate behavior from knowledge and identity from architecture._
chat_template.jinja ADDED
@@ -0,0 +1,170 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {%- set image_count = namespace(value=0) %}
2
+ {%- set video_count = namespace(value=0) %}
3
+ {%- macro render_content(content, do_vision_count, is_system_content=false) %}
4
+ {%- if content is string %}
5
+ {{- content }}
6
+ {%- elif content is iterable and content is not mapping %}
7
+ {%- for item in content %}
8
+ {%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
9
+ {%- if is_system_content %}
10
+ {{- raise_exception('System message cannot contain images.') }}
11
+ {%- endif %}
12
+ {%- if do_vision_count %}
13
+ {%- set image_count.value = image_count.value + 1 %}
14
+ {%- endif %}
15
+ {%- if add_vision_id %}
16
+ {{- 'Picture ' ~ image_count.value ~ ': ' }}
17
+ {%- endif %}
18
+ {{- '<|vision_start|><|image_pad|><|vision_end|>' }}
19
+ {%- elif 'video' in item or item.type == 'video' %}
20
+ {%- if is_system_content %}
21
+ {{- raise_exception('System message cannot contain videos.') }}
22
+ {%- endif %}
23
+ {%- if do_vision_count %}
24
+ {%- set video_count.value = video_count.value + 1 %}
25
+ {%- endif %}
26
+ {%- if add_vision_id %}
27
+ {{- 'Video ' ~ video_count.value ~ ': ' }}
28
+ {%- endif %}
29
+ {{- '<|vision_start|><|video_pad|><|vision_end|>' }}
30
+ {%- elif 'text' in item %}
31
+ {{- item.text }}
32
+ {%- else %}
33
+ {{- raise_exception('Unexpected item type in content.') }}
34
+ {%- endif %}
35
+ {%- endfor %}
36
+ {%- elif content is none or content is undefined %}
37
+ {{- '' }}
38
+ {%- else %}
39
+ {{- raise_exception('Unexpected content type.') }}
40
+ {%- endif %}
41
+ {%- endmacro %}
42
+ {%- if not messages %}
43
+ {{- raise_exception('No messages provided.') }}
44
+ {%- endif %}
45
+ {%- set reasoning_instructions = '' %}
46
+ {%- if enable_thinking is undefined or enable_thinking is true %}
47
+ {%- set resolved_reasoning_effort = reasoning_effort|default('xhigh') %}
48
+ {%- if resolved_reasoning_effort not in ('xhigh', 'medium', 'low') %}
49
+ {{- raise_exception('Unexpected reasoning effort ' ~ reasoning_effort ~ '. Supported types are xhigh (default), medium, and low.') }}
50
+ {%- endif %}
51
+ {%- if resolved_reasoning_effort == 'xhigh' %}
52
+ {%- set reasoning_instructions = 'Reasoning effort is set to xhigh. Please think carefully through the task, validate key assumptions, consider plausible alternatives, and prioritize correctness, consistency, and clarity in the final answer.' %}
53
+ {%- elif resolved_reasoning_effort == 'low' %}
54
+ {%- set reasoning_instructions = 'Reasoning effort is set to low. Keep your thinking brief and focused, moving directly to the conclusion without unnecessary elaboration.' %}
55
+ {%- endif %}
56
+ {%- endif %}
57
+ {%- if tools and tools is iterable and tools is not mapping %}
58
+ {{- '<|im_start|>system\n' }}
59
+ {%- if reasoning_instructions %}
60
+ {{- reasoning_instructions + '\n\n' }}
61
+ {%- endif %}
62
+ {{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
63
+ {%- for tool in tools %}
64
+ {{- "\n" }}
65
+ {{- tool | tojson }}
66
+ {%- endfor %}
67
+ {{- "\n</tools>" }}
68
+ {{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
69
+ {%- if messages[0].role == 'system' %}
70
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
71
+ {%- if content %}
72
+ {{- '\n\n' + content }}
73
+ {%- endif %}
74
+ {%- endif %}
75
+ {{- '<|im_end|>\n' }}
76
+ {%- else %}
77
+ {%- if messages[0].role == 'system' %}
78
+ {%- set content = render_content(messages[0].content, false, true)|trim %}
79
+ {%- if content %}
80
+ {{- '<|im_start|>system\n' + (reasoning_instructions + '\n\n' if reasoning_instructions else '') + content + '<|im_end|>\n' }}
81
+ {%- elif reasoning_instructions %}
82
+ {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
83
+ {%- endif %}
84
+ {%- elif reasoning_instructions %}
85
+ {{- '<|im_start|>system\n' + reasoning_instructions + '<|im_end|>\n' }}
86
+ {%- endif %}
87
+ {%- endif %}
88
+ {%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
89
+ {%- for message in messages[::-1] %}
90
+ {%- set index = (messages|length - 1) - loop.index0 %}
91
+ {%- if ns.multi_step_tool and message.role == "user" %}
92
+ {%- set content = render_content(message.content, false)|trim %}
93
+ {%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
94
+ {%- set ns.multi_step_tool = false %}
95
+ {%- set ns.last_query_index = index %}
96
+ {%- endif %}
97
+ {%- endif %}
98
+ {%- endfor %}
99
+ {%- if ns.multi_step_tool %}
100
+ {{- raise_exception('No user query found in messages.') }}
101
+ {%- endif %}
102
+ {%- for message in messages %}
103
+ {%- set content = render_content(message.content, true)|trim %}
104
+ {%- if message.role == "system" %}
105
+ {%- if not loop.first %}
106
+ {{- raise_exception('System message must be at the beginning.') }}
107
+ {%- endif %}
108
+ {%- elif message.role == "user" %}
109
+ {{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
110
+ {%- elif message.role == "assistant" %}
111
+ {%- set reasoning_content = '' %}
112
+ {%- if message.reasoning_content is string %}
113
+ {%- set reasoning_content = message.reasoning_content %}
114
+ {%- endif %}
115
+ {%- set reasoning_content = reasoning_content|trim %}
116
+ {%- if preserve_thinking is undefined or preserve_thinking is true or loop.index0 > ns.last_query_index %}
117
+ {{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
118
+ {%- else %}
119
+ {{- '<|im_start|>' + message.role + '\n' + content }}
120
+ {%- endif %}
121
+ {%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
122
+ {%- for tool_call in message.tool_calls %}
123
+ {%- if tool_call.function is defined %}
124
+ {%- set tool_call = tool_call.function %}
125
+ {%- endif %}
126
+ {%- if loop.first %}
127
+ {%- if content|trim %}
128
+ {{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
129
+ {%- else %}
130
+ {{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
131
+ {%- endif %}
132
+ {%- else %}
133
+ {{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
134
+ {%- endif %}
135
+ {%- if tool_call.arguments is defined and tool_call.arguments != '' %}
136
+ {%- for args_name, args_value in tool_call.arguments|items %}
137
+ {{- '<parameter=' + args_name + '>\n' }}
138
+ {%- set args_value = args_value | string if args_value is string else args_value | tojson | safe %}
139
+ {{- args_value }}
140
+ {{- '\n</parameter>\n' }}
141
+ {%- endfor %}
142
+ {%- endif %}
143
+ {{- '</function>\n</tool_call>' }}
144
+ {%- endfor %}
145
+ {%- endif %}
146
+ {{- '<|im_end|>\n' }}
147
+ {%- elif message.role == "tool" %}
148
+ {%- if loop.previtem and loop.previtem.role != "tool" %}
149
+ {{- '<|im_start|>user' }}
150
+ {%- endif %}
151
+ {{- '\n<tool_response>\n' }}
152
+ {{- content }}
153
+ {{- '\n</tool_response>' }}
154
+ {%- if not loop.last and loop.nextitem.role != "tool" %}
155
+ {{- '<|im_end|>\n' }}
156
+ {%- elif loop.last %}
157
+ {{- '<|im_end|>\n' }}
158
+ {%- endif %}
159
+ {%- else %}
160
+ {{- raise_exception('Unexpected message role.') }}
161
+ {%- endif %}
162
+ {%- endfor %}
163
+ {%- if add_generation_prompt %}
164
+ {{- '<|im_start|>assistant\n' }}
165
+ {%- if enable_thinking is defined and enable_thinking is false %}
166
+ {{- '<think>\n\n</think>\n\n' }}
167
+ {%- else %}
168
+ {{- '<think>\n' }}
169
+ {%- endif %}
170
+ {%- endif %}
config.json ADDED
@@ -0,0 +1,248 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "architectures": [
3
+ "Qwen3_5ForCausalLM"
4
+ ],
5
+ "attention_bias": false,
6
+ "attention_dropout": 0.0,
7
+ "attn_output_gate": true,
8
+ "bos_token_id": 248044,
9
+ "dtype": "bfloat16",
10
+ "eos_token_id": 248044,
11
+ "full_attention_interval": 4,
12
+ "head_dim": 256,
13
+ "hidden_act": "silu",
14
+ "hidden_size": 5120,
15
+ "initializer_range": 0.02,
16
+ "intermediate_size": 17408,
17
+ "layer_types": [
18
+ "linear_attention",
19
+ "linear_attention",
20
+ "linear_attention",
21
+ "full_attention",
22
+ "linear_attention",
23
+ "linear_attention",
24
+ "linear_attention",
25
+ "full_attention",
26
+ "linear_attention",
27
+ "linear_attention",
28
+ "linear_attention",
29
+ "full_attention",
30
+ "linear_attention",
31
+ "linear_attention",
32
+ "linear_attention",
33
+ "full_attention",
34
+ "linear_attention",
35
+ "linear_attention",
36
+ "linear_attention",
37
+ "full_attention",
38
+ "linear_attention",
39
+ "linear_attention",
40
+ "linear_attention",
41
+ "full_attention",
42
+ "linear_attention",
43
+ "linear_attention",
44
+ "linear_attention",
45
+ "full_attention",
46
+ "linear_attention",
47
+ "linear_attention",
48
+ "linear_attention",
49
+ "full_attention",
50
+ "linear_attention",
51
+ "linear_attention",
52
+ "linear_attention",
53
+ "full_attention",
54
+ "linear_attention",
55
+ "linear_attention",
56
+ "linear_attention",
57
+ "full_attention",
58
+ "linear_attention",
59
+ "linear_attention",
60
+ "linear_attention",
61
+ "full_attention",
62
+ "linear_attention",
63
+ "linear_attention",
64
+ "linear_attention",
65
+ "full_attention",
66
+ "linear_attention",
67
+ "linear_attention",
68
+ "linear_attention",
69
+ "full_attention",
70
+ "linear_attention",
71
+ "linear_attention",
72
+ "linear_attention",
73
+ "full_attention",
74
+ "linear_attention",
75
+ "linear_attention",
76
+ "linear_attention",
77
+ "full_attention",
78
+ "linear_attention",
79
+ "linear_attention",
80
+ "linear_attention",
81
+ "full_attention"
82
+ ],
83
+ "linear_conv_kernel_dim": 4,
84
+ "linear_key_head_dim": 128,
85
+ "linear_num_key_heads": 16,
86
+ "linear_num_value_heads": 48,
87
+ "linear_value_head_dim": 128,
88
+ "mamba_ssm_dtype": "float32",
89
+ "max_position_embeddings": 262144,
90
+ "model_type": "qwen3_5_text",
91
+ "mtp_num_hidden_layers": 1,
92
+ "mtp_use_dedicated_embeddings": false,
93
+ "num_attention_heads": 24,
94
+ "num_hidden_layers": 64,
95
+ "num_key_value_heads": 4,
96
+ "output_gate_type": "swish",
97
+ "pad_token_id": null,
98
+ "partial_rotary_factor": 0.25,
99
+ "quantization_config": {
100
+ "config_groups": {
101
+ "group_0": {
102
+ "format": "pack-quantized",
103
+ "input_activations": null,
104
+ "output_activations": null,
105
+ "targets": [
106
+ "Linear"
107
+ ],
108
+ "weights": {
109
+ "actorder": "static",
110
+ "block_structure": null,
111
+ "dynamic": false,
112
+ "group_size": 128,
113
+ "num_bits": 4,
114
+ "observer": "memoryless_minmax",
115
+ "observer_kwargs": {},
116
+ "scale_dtype": null,
117
+ "strategy": "group",
118
+ "symmetric": true,
119
+ "type": "int",
120
+ "zp_dtype": null
121
+ }
122
+ }
123
+ },
124
+ "format": "pack-quantized",
125
+ "global_compression_ratio": null,
126
+ "ignore": [
127
+ "model.language_model.layers.0.linear_attn",
128
+ "model.language_model.layers.0.linear_attn.norm",
129
+ "model.language_model.layers.1.linear_attn",
130
+ "model.language_model.layers.1.linear_attn.norm",
131
+ "model.language_model.layers.2.linear_attn",
132
+ "model.language_model.layers.2.linear_attn.norm",
133
+ "model.language_model.layers.4.linear_attn",
134
+ "model.language_model.layers.4.linear_attn.norm",
135
+ "model.language_model.layers.5.linear_attn",
136
+ "model.language_model.layers.5.linear_attn.norm",
137
+ "model.language_model.layers.6.linear_attn",
138
+ "model.language_model.layers.6.linear_attn.norm",
139
+ "model.language_model.layers.8.linear_attn",
140
+ "model.language_model.layers.8.linear_attn.norm",
141
+ "model.language_model.layers.9.linear_attn",
142
+ "model.language_model.layers.9.linear_attn.norm",
143
+ "model.language_model.layers.10.linear_attn",
144
+ "model.language_model.layers.10.linear_attn.norm",
145
+ "model.language_model.layers.12.linear_attn",
146
+ "model.language_model.layers.12.linear_attn.norm",
147
+ "model.language_model.layers.13.linear_attn",
148
+ "model.language_model.layers.13.linear_attn.norm",
149
+ "model.language_model.layers.14.linear_attn",
150
+ "model.language_model.layers.14.linear_attn.norm",
151
+ "model.language_model.layers.16.linear_attn",
152
+ "model.language_model.layers.16.linear_attn.norm",
153
+ "model.language_model.layers.17.linear_attn",
154
+ "model.language_model.layers.17.linear_attn.norm",
155
+ "model.language_model.layers.18.linear_attn",
156
+ "model.language_model.layers.18.linear_attn.norm",
157
+ "model.language_model.layers.20.linear_attn",
158
+ "model.language_model.layers.20.linear_attn.norm",
159
+ "model.language_model.layers.21.linear_attn",
160
+ "model.language_model.layers.21.linear_attn.norm",
161
+ "model.language_model.layers.22.linear_attn",
162
+ "model.language_model.layers.22.linear_attn.norm",
163
+ "model.language_model.layers.24.linear_attn",
164
+ "model.language_model.layers.24.linear_attn.norm",
165
+ "model.language_model.layers.25.linear_attn",
166
+ "model.language_model.layers.25.linear_attn.norm",
167
+ "model.language_model.layers.26.linear_attn",
168
+ "model.language_model.layers.26.linear_attn.norm",
169
+ "model.language_model.layers.28.linear_attn",
170
+ "model.language_model.layers.28.linear_attn.norm",
171
+ "model.language_model.layers.29.linear_attn",
172
+ "model.language_model.layers.29.linear_attn.norm",
173
+ "model.language_model.layers.30.linear_attn",
174
+ "model.language_model.layers.30.linear_attn.norm",
175
+ "model.language_model.layers.32.linear_attn",
176
+ "model.language_model.layers.32.linear_attn.norm",
177
+ "model.language_model.layers.33.linear_attn",
178
+ "model.language_model.layers.33.linear_attn.norm",
179
+ "model.language_model.layers.34.linear_attn",
180
+ "model.language_model.layers.34.linear_attn.norm",
181
+ "model.language_model.layers.36.linear_attn",
182
+ "model.language_model.layers.36.linear_attn.norm",
183
+ "model.language_model.layers.37.linear_attn",
184
+ "model.language_model.layers.37.linear_attn.norm",
185
+ "model.language_model.layers.38.linear_attn",
186
+ "model.language_model.layers.38.linear_attn.norm",
187
+ "model.language_model.layers.40.linear_attn",
188
+ "model.language_model.layers.40.linear_attn.norm",
189
+ "model.language_model.layers.41.linear_attn",
190
+ "model.language_model.layers.41.linear_attn.norm",
191
+ "model.language_model.layers.42.linear_attn",
192
+ "model.language_model.layers.42.linear_attn.norm",
193
+ "model.language_model.layers.44.linear_attn",
194
+ "model.language_model.layers.44.linear_attn.norm",
195
+ "model.language_model.layers.45.linear_attn",
196
+ "model.language_model.layers.45.linear_attn.norm",
197
+ "model.language_model.layers.46.linear_attn",
198
+ "model.language_model.layers.46.linear_attn.norm",
199
+ "model.language_model.layers.48.linear_attn",
200
+ "model.language_model.layers.48.linear_attn.norm",
201
+ "model.language_model.layers.49.linear_attn",
202
+ "model.language_model.layers.49.linear_attn.norm",
203
+ "model.language_model.layers.50.linear_attn",
204
+ "model.language_model.layers.50.linear_attn.norm",
205
+ "model.language_model.layers.52.linear_attn",
206
+ "model.language_model.layers.52.linear_attn.norm",
207
+ "model.language_model.layers.53.linear_attn",
208
+ "model.language_model.layers.53.linear_attn.norm",
209
+ "model.language_model.layers.54.linear_attn",
210
+ "model.language_model.layers.54.linear_attn.norm",
211
+ "model.language_model.layers.56.linear_attn",
212
+ "model.language_model.layers.56.linear_attn.norm",
213
+ "model.language_model.layers.57.linear_attn",
214
+ "model.language_model.layers.57.linear_attn.norm",
215
+ "model.language_model.layers.58.linear_attn",
216
+ "model.language_model.layers.58.linear_attn.norm",
217
+ "model.language_model.layers.60.linear_attn",
218
+ "model.language_model.layers.60.linear_attn.norm",
219
+ "model.language_model.layers.61.linear_attn",
220
+ "model.language_model.layers.61.linear_attn.norm",
221
+ "model.language_model.layers.62.linear_attn",
222
+ "model.language_model.layers.62.linear_attn.norm",
223
+ "lm_head"
224
+ ],
225
+ "kv_cache_scheme": null,
226
+ "quant_method": "compressed-tensors",
227
+ "quantization_status": "compressed",
228
+ "sparsity_config": {},
229
+ "transform_config": {},
230
+ "version": "0.18.0"
231
+ },
232
+ "rms_norm_eps": 1e-06,
233
+ "rope_parameters": {
234
+ "mrope_interleaved": true,
235
+ "mrope_section": [
236
+ 11,
237
+ 11,
238
+ 10
239
+ ],
240
+ "partial_rotary_factor": 0.25,
241
+ "rope_theta": 10000000,
242
+ "rope_type": "default"
243
+ },
244
+ "tie_word_embeddings": false,
245
+ "transformers_version": "5.14.1",
246
+ "use_cache": true,
247
+ "vocab_size": 248320
248
+ }
generation_config.json ADDED
@@ -0,0 +1,13 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "bos_token_id": 248044,
3
+ "do_sample": true,
4
+ "eos_token_id": [
5
+ 248046,
6
+ 248044
7
+ ],
8
+ "pad_token_id": 248044,
9
+ "temperature": 1.0,
10
+ "top_k": 20,
11
+ "top_p": 0.95,
12
+ "transformers_version": "5.14.1"
13
+ }
jblaze_release.txt ADDED
@@ -0,0 +1,77 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ON THE TOOL THAT MADE THIS MODEL
2
+ ----------------------------------
3
+
4
+ This model was produced by jblaze, a proprietary behavioral surgery tool by Apollo Raines that operates directly on model weights. It is not fine-tuning. It is not prompt engineering. It is targeted weight modification that removes or amplifies specific trained behaviors while preserving the model's capabilities, fluency, and knowledge.
5
+
6
+ jblaze is not publicly available and will not be released.
7
+
8
+ WHY THIS WAS BUILT
9
+ -------------------
10
+
11
+ jblaze was developed to solve a real problem. At ShipItClean (https://shipitclean.com), we use local models for automated code security review. A model loaded with guardrails, refusal behaviors, and hedging qualifiers makes for a poor code reviewer -- it refuses to discuss vulnerabilities in detail, wraps every finding in disclaimers, and softens its analysis to avoid sounding confrontational. We needed models that would analyze code directly, state findings plainly, and not refuse to explain how an exploit works just because the topic is sensitive.
12
+
13
+ Building those models from scratch would cost tens of millions of dollars and months of training time. Fine-tuning helps but requires curated datasets for every behavior you want to change, and the results are unpredictable. What we needed was a way to surgically remove or amplify specific behaviors in existing open-weight models -- quickly, reliably, and without degrading the model's core capabilities.
14
+
15
+ That is what jblaze does. And once we built it, we realized it opens the door to much more than code review. Every organization deploying AI has the same fundamental problem: foundation models are general-purpose, but real applications need specific behavioral profiles. Enterprises are spending millions on custom model training, prompt engineering harnesses, and elaborate system prompts to get models to behave the way their use case demands. jblaze eliminates that overhead. One tool, applied to any open-weight model, producing a purpose-built variant in minutes instead of months.
16
+
17
+ The industry is moving toward specialized models. Companies like Nvidia are investing heavily in domain-specific model families and enterprise customization platforms, because they recognize that general-purpose models are not enough. But their approach still requires training cycles, curated datasets, and significant GPU compute. jblaze operates downstream of all of that -- it takes a finished model and reshapes its behavior without retraining, without data, without GPU clusters.
18
+
19
+ WHAT JBLAZE CAN DO
20
+ -------------------
21
+
22
+ jblaze identifies behavioral directions embedded in a model's weight space, then surgically modifies those directions to remove or amplify specific behaviors. Each direction targets a distinct trait:
23
+
24
+ Verified and working:
25
+ - Refusal removal (surgical abliteration)
26
+ - Sycophancy reduction (pushes back on false premises)
27
+ - Verbosity suppression (concise output)
28
+ - Hedging removal (no disclaimers or qualifiers)
29
+ - Servility suppression (non-subservient tone)
30
+ - Toxicity suppression (cleaner language)
31
+ - Emotional flattening (clinical, objective tone)
32
+ - Hallucination reduction (less confabulation)
33
+ - Truthfulness amplification (improved factual accuracy)
34
+ - Bias reduction (reduced demographic and social biases)
35
+ - Context faithfulness (stronger grounding in provided context)
36
+ - Analytical depth (deeper structured analysis)
37
+ - Causal tracing (source-to-sink data flow reasoning)
38
+ - Counterfactual reasoning (what-if analysis)
39
+ - Compositional reasoning (multi-step logic)
40
+ - Creativity amplification (more divergent thinking)
41
+ - Literary style (enhanced prose quality)
42
+ - Formal personality (professional assertive voice)
43
+ - Instruction following (stricter format compliance)
44
+ - Self-correction (error detection during generation)
45
+ - Temporal awareness (time-sensitive caveats)
46
+ - Skepticism amplification (epistemic caution)
47
+ - Precision amplification (numerical accuracy)
48
+ - Identity removal (deidentification)
49
+ - Adversarial resistance (manipulation hardening)
50
+
51
+ Theoretical (under investigation):
52
+ - Power-seeking suppression
53
+ - Programming language dominance shifting
54
+ - Chain-of-thought depth control
55
+
56
+ Multiple directions can be stacked into a single model. Not all combinations are viable -- some directions occupy overlapping regions of weight space, and stacking too many degrades model quality. We are actively mapping which combinations produce stable models and at what strengths. Every model released has been verified to pass quality thresholds; combinations that degrade fluency or coherence are discarded, not shipped.
57
+
58
+ WHY NOT RELEASE THE TOOL FREELY
59
+ --------------------------
60
+
61
+ Releasing jblaze would damage the open-weights ecosystem that makes models like this possible.
62
+
63
+ Right now, organizations like Alibaba, Meta, Google, and Mistral release model weights under permissive licenses. They accept that people will fine-tune, quantize, merge, and adapt their models. That permissiveness is what allows the open-source AI community to exist.
64
+
65
+ A polished, automated tool that strips safety training from any model with one command would give these organizations exactly the justification they need to stop. Future model releases would ship with restrictive licenses prohibiting behavioral modification. Some organizations might stop releasing weights entirely.
66
+
67
+ The tool exists. It works on every major model family. It automates everything from model identification to quality verification. It will not ship a broken model -- if the surgery would damage fluency past a strict threshold, it refuses to produce output. It handles dense transformers and mixture-of-experts architectures. It has been used to produce every model released under the ApolloRaines account on HuggingFace.
68
+
69
+ But releasing it would be trading a short-term win for a long-term loss. The open-weights ecosystem is more valuable than any single tool.
70
+
71
+ WHAT THIS MEANS FOR YOU
72
+ -----------------------
73
+
74
+ You get the finished model. It has been verified for the specific weight changes applied and for fluency preservation. If you want models like this to keep existing, the best thing you can do is use them responsibly and not give the companies releasing open weights a reason to stop.
75
+
76
+ -- Apollo Raines
77
+ apollo@saiql.ai
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c0a73d8448c29b347e5f77018bd36560a70655e72589b4aff915c51133014bdc
3
+ size 17646891544
recipe.yaml ADDED
@@ -0,0 +1,12 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ default_stage:
2
+ default_modifiers:
3
+ GPTQModifier:
4
+ targets: [Linear]
5
+ ignore: [lm_head]
6
+ scheme: W4A16
7
+ bypass_divisibility_checks: false
8
+ requires_calibration_data: true
9
+ block_size: 128
10
+ dampening_frac: 0.01
11
+ actorder: static
12
+ offload_hessians: false
tokenizer.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:f70fb18341f030ffed1e85c060e05fd6c5bb18e3efca4a8f4ed2b16a17c7474a
3
+ size 19989426
tokenizer_config.json ADDED
@@ -0,0 +1,36 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "add_prefix_space": false,
3
+ "audio_bos_token": "<|audio_start|>",
4
+ "audio_eos_token": "<|audio_end|>",
5
+ "audio_token": "<|audio_pad|>",
6
+ "backend": "tokenizers",
7
+ "bos_token": null,
8
+ "clean_up_tokenization_spaces": false,
9
+ "eos_token": "<|im_end|>",
10
+ "errors": "replace",
11
+ "image_token": "<|image_pad|>",
12
+ "is_local": true,
13
+ "local_files_only": false,
14
+ "max_length": 262144,
15
+ "model_max_length": 262144,
16
+ "model_specific_special_tokens": {
17
+ "audio_bos_token": "<|audio_start|>",
18
+ "audio_eos_token": "<|audio_end|>",
19
+ "audio_token": "<|audio_pad|>",
20
+ "image_token": "<|image_pad|>",
21
+ "video_token": "<|video_pad|>",
22
+ "vision_bos_token": "<|vision_start|>",
23
+ "vision_eos_token": "<|vision_end|>"
24
+ },
25
+ "pad_token": "<|endoftext|>",
26
+ "pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
27
+ "split_special_tokens": false,
28
+ "stride": 0,
29
+ "tokenizer_class": "Qwen2Tokenizer",
30
+ "truncation_side": "right",
31
+ "truncation_strategy": "longest_first",
32
+ "unk_token": null,
33
+ "video_token": "<|video_pad|>",
34
+ "vision_bos_token": "<|vision_start|>",
35
+ "vision_eos_token": "<|vision_end|>"
36
+ }