Deepnar commited on
Commit
b69b2d4
·
verified ·
1 Parent(s): f2e02db

Release retained frozen ICE v2 paper checkpoint and inference metadata

Browse files
Files changed (8) hide show
  1. LICENSE +202 -0
  2. NOTICE +16 -0
  3. README.md +132 -0
  4. classifier_inference.py +47 -0
  5. config.json +56 -0
  6. frozen_classifier.py +202 -0
  7. ice_classifier_v3_qwen_ft3.pt +3 -0
  8. model.py +19 -0
LICENSE ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+
2
+ Apache License
3
+ Version 2.0, January 2004
4
+ http://www.apache.org/licenses/
5
+
6
+ TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
7
+
8
+ 1. Definitions.
9
+
10
+ "License" shall mean the terms and conditions for use, reproduction,
11
+ and distribution as defined by Sections 1 through 9 of this document.
12
+
13
+ "Licensor" shall mean the copyright owner or entity authorized by
14
+ the copyright owner that is granting the License.
15
+
16
+ "Legal Entity" shall mean the union of the acting entity and all
17
+ other entities that control, are controlled by, or are under common
18
+ control with that entity. For the purposes of this definition,
19
+ "control" means (i) the power, direct or indirect, to cause the
20
+ direction or management of such entity, whether by contract or
21
+ otherwise, or (ii) ownership of fifty percent (50%) or more of the
22
+ outstanding shares, or (iii) beneficial ownership of such entity.
23
+
24
+ "You" (or "Your") shall mean an individual or Legal Entity
25
+ exercising permissions granted by this License.
26
+
27
+ "Source" form shall mean the preferred form for making modifications,
28
+ including but not limited to software source code, documentation
29
+ source, and configuration files.
30
+
31
+ "Object" form shall mean any form resulting from mechanical
32
+ transformation or translation of a Source form, including but
33
+ not limited to compiled object code, generated documentation,
34
+ and conversions to other media types.
35
+
36
+ "Work" shall mean the work of authorship, whether in Source or
37
+ Object form, made available under the License, as indicated by a
38
+ copyright notice that is included in or attached to the work
39
+ (an example is provided in the Appendix below).
40
+
41
+ "Derivative Works" shall mean any work, whether in Source or Object
42
+ form, that is based on (or derived from) the Work and for which the
43
+ editorial revisions, annotations, elaborations, or other modifications
44
+ represent, as a whole, an original work of authorship. For the purposes
45
+ of this License, Derivative Works shall not include works that remain
46
+ separable from, or merely link (or bind by name) to the interfaces of,
47
+ the Work and Derivative Works thereof.
48
+
49
+ "Contribution" shall mean any work of authorship, including
50
+ the original version of the Work and any modifications or additions
51
+ to that Work or Derivative Works thereof, that is intentionally
52
+ submitted to Licensor for inclusion in the Work by the copyright owner
53
+ or by an individual or Legal Entity authorized to submit on behalf of
54
+ the copyright owner. For the purposes of this definition, "submitted"
55
+ means any form of electronic, verbal, or written communication sent
56
+ to the Licensor or its representatives, including but not limited to
57
+ communication on electronic mailing lists, source code control systems,
58
+ and issue tracking systems that are managed by, or on behalf of, the
59
+ Licensor for the purpose of discussing and improving the Work, but
60
+ excluding communication that is conspicuously marked or otherwise
61
+ designated in writing by the copyright owner as "Not a Contribution."
62
+
63
+ "Contributor" shall mean Licensor and any individual or Legal Entity
64
+ on behalf of whom a Contribution has been received by Licensor and
65
+ subsequently incorporated within the Work.
66
+
67
+ 2. Grant of Copyright License. Subject to the terms and conditions of
68
+ this License, each Contributor hereby grants to You a perpetual,
69
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
70
+ copyright license to reproduce, prepare Derivative Works of,
71
+ publicly display, publicly perform, sublicense, and distribute the
72
+ Work and such Derivative Works in Source or Object form.
73
+
74
+ 3. Grant of Patent License. Subject to the terms and conditions of
75
+ this License, each Contributor hereby grants to You a perpetual,
76
+ worldwide, non-exclusive, no-charge, royalty-free, irrevocable
77
+ (except as stated in this section) patent license to make, have made,
78
+ use, offer to sell, sell, import, and otherwise transfer the Work,
79
+ where such license applies only to those patent claims licensable
80
+ by such Contributor that are necessarily infringed by their
81
+ Contribution(s) alone or by combination of their Contribution(s)
82
+ with the Work to which such Contribution(s) was submitted. If You
83
+ institute patent litigation against any entity (including a
84
+ cross-claim or counterclaim in a lawsuit) alleging that the Work
85
+ or a Contribution incorporated within the Work constitutes direct
86
+ or contributory patent infringement, then any patent licenses
87
+ granted to You under this License for that Work shall terminate
88
+ as of the date such litigation is filed.
89
+
90
+ 4. Redistribution. You may reproduce and distribute copies of the
91
+ Work or Derivative Works thereof in any medium, with or without
92
+ modifications, and in Source or Object form, provided that You
93
+ meet the following conditions:
94
+
95
+ (a) You must give any other recipients of the Work or
96
+ Derivative Works a copy of this License; and
97
+
98
+ (b) You must cause any modified files to carry prominent notices
99
+ stating that You changed the files; and
100
+
101
+ (c) You must retain, in the Source form of any Derivative Works
102
+ that You distribute, all copyright, patent, trademark, and
103
+ attribution notices from the Source form of the Work,
104
+ excluding those notices that do not pertain to any part of
105
+ the Derivative Works; and
106
+
107
+ (d) If the Work includes a "NOTICE" text file as part of its
108
+ distribution, then any Derivative Works that You distribute must
109
+ include a readable copy of the attribution notices contained
110
+ within such NOTICE file, excluding those notices that do not
111
+ pertain to any part of the Derivative Works, in at least one
112
+ of the following places: within a NOTICE text file distributed
113
+ as part of the Derivative Works; within the Source form or
114
+ documentation, if provided along with the Derivative Works; or,
115
+ within a display generated by the Derivative Works, if and
116
+ wherever such third-party notices normally appear. The contents
117
+ of the NOTICE file are for informational purposes only and
118
+ do not modify the License. You may add Your own attribution
119
+ notices within Derivative Works that You distribute, alongside
120
+ or as an addendum to the NOTICE text from the Work, provided
121
+ that such additional attribution notices cannot be construed
122
+ as modifying the License.
123
+
124
+ You may add Your own copyright statement to Your modifications and
125
+ may provide additional or different license terms and conditions
126
+ for use, reproduction, or distribution of Your modifications, or
127
+ for any such Derivative Works as a whole, provided Your use,
128
+ reproduction, and distribution of the Work otherwise complies with
129
+ the conditions stated in this License.
130
+
131
+ 5. Submission of Contributions. Unless You explicitly state otherwise,
132
+ any Contribution intentionally submitted for inclusion in the Work
133
+ by You to the Licensor shall be under the terms and conditions of
134
+ this License, without any additional terms or conditions.
135
+ Notwithstanding the above, nothing herein shall supersede or modify
136
+ the terms of any separate license agreement you may have executed
137
+ with Licensor regarding such Contributions.
138
+
139
+ 6. Trademarks. This License does not grant permission to use the trade
140
+ names, trademarks, service marks, or product names of the Licensor,
141
+ except as required for reasonable and customary use in describing the
142
+ origin of the Work and reproducing the content of the NOTICE file.
143
+
144
+ 7. Disclaimer of Warranty. Unless required by applicable law or
145
+ agreed to in writing, Licensor provides the Work (and each
146
+ Contributor provides its Contributions) on an "AS IS" BASIS,
147
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
148
+ implied, including, without limitation, any warranties or conditions
149
+ of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
150
+ PARTICULAR PURPOSE. You are solely responsible for determining the
151
+ appropriateness of using or redistributing the Work and assume any
152
+ risks associated with Your exercise of permissions under this License.
153
+
154
+ 8. Limitation of Liability. In no event and under no legal theory,
155
+ whether in tort (including negligence), contract, or otherwise,
156
+ unless required by applicable law (such as deliberate and grossly
157
+ negligent acts) or agreed to in writing, shall any Contributor be
158
+ liable to You for damages, including any direct, indirect, special,
159
+ incidental, or consequential damages of any character arising as a
160
+ result of this License or out of the use or inability to use the
161
+ Work (including but not limited to damages for loss of goodwill,
162
+ work stoppage, computer failure or malfunction, or any and all
163
+ other commercial damages or losses), even if such Contributor
164
+ has been advised of the possibility of such damages.
165
+
166
+ 9. Accepting Warranty or Additional Liability. While redistributing
167
+ the Work or Derivative Works thereof, You may choose to offer,
168
+ and charge a fee for, acceptance of support, warranty, indemnity,
169
+ or other liability obligations and/or rights consistent with this
170
+ License. However, in accepting such obligations, You may act only
171
+ on Your own behalf and on Your sole responsibility, not on behalf
172
+ of any other Contributor, and only if You agree to indemnify,
173
+ defend, and hold each Contributor harmless for any liability
174
+ incurred by, or claims asserted against, such Contributor by reason
175
+ of your accepting any such warranty or additional liability.
176
+
177
+ END OF TERMS AND CONDITIONS
178
+
179
+ APPENDIX: How to apply the Apache License to your work.
180
+
181
+ To apply the Apache License to your work, attach the following
182
+ boilerplate notice, with the fields enclosed by brackets "[]"
183
+ replaced with your own identifying information. (Don't include
184
+ the brackets!) The text should be enclosed in the appropriate
185
+ comment syntax for the file format. We also recommend that a
186
+ file or class name and description of purpose be included on the
187
+ same "printed page" as the copyright notice for easier
188
+ identification within third-party archives.
189
+
190
+ Copyright 2026 Deepesh Sonar
191
+
192
+ Licensed under the Apache License, Version 2.0 (the "License");
193
+ you may not use this file except in compliance with the License.
194
+ You may obtain a copy of the License at
195
+
196
+ http://www.apache.org/licenses/LICENSE-2.0
197
+
198
+ Unless required by applicable law or agreed to in writing, software
199
+ distributed under the License is distributed on an "AS IS" BASIS,
200
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
201
+ See the License for the specific language governing permissions and
202
+ limitations under the License.
NOTICE ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ICE — Infinite Context Engine
2
+ Copyright 2026 Deepesh Sonar
3
+
4
+ This product includes software developed by Deepesh Sonar.
5
+
6
+ Licensed under the Apache License, Version 2.0 (the "License");
7
+ you may not use this file except in compliance with the License.
8
+ You may obtain a copy of the License at
9
+
10
+ http://www.apache.org/licenses/LICENSE-2.0
11
+
12
+ Unless required by applicable law or agreed to in writing, software
13
+ distributed under the License is distributed on an "AS IS" BASIS,
14
+ WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
15
+ See the License for the specific language governing permissions and
16
+ limitations under the License.
README.md ADDED
@@ -0,0 +1,132 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ language:
4
+ - en
5
+ base_model: Qwen/Qwen3-Embedding-0.6B
6
+ library_name: pytorch
7
+ tags:
8
+ - conversational-memory
9
+ - lsrep
10
+ - ice-v2
11
+ - reproducibility
12
+ - arxiv:2609.16730
13
+ ---
14
+
15
+ # ICE v2 / LSREP experiment classifier
16
+
17
+ This releases the retained original classifier checkpoint selected by frozen
18
+ **ICE v2**, the system evaluated in the [LSREP paper](https://arxiv.org/abs/2609.16730).
19
+ It is **not the current ICE v3 classifier**. The `v3` in the historical
20
+ checkpoint filename denotes a classifier training generation, not ICE's system version.
21
+
22
+ - Repository: [Deepnar/ice](https://github.com/Deepnar/ice).
23
+ - Evaluated source tag: [`v2-paper-eval`](https://github.com/Deepnar/ice/tree/v2-paper-eval).
24
+ - Evaluated commit: `0521df9171b4a7d69f82d12d70497138c77b2678`.
25
+ - Original path: `models/classifier/ice_classifier_v3_qwen_ft3.pt`.
26
+ - Original file size: 212,827 bytes; bare PyTorch state dict.
27
+ - SHA-256: `25c758b6a7e5cf449f3e4c8bb250db759d37cb4f0ab7dd8e0c1acd8afbf05831`.
28
+
29
+ The checkpoint is copied byte for byte, without retraining or re-export.
30
+ The frozen Git tag records code and the selected path, but does not contain
31
+ checkpoint blobs or a historical checkpoint checksum. This release records the
32
+ checksum of the retained original, rather than claiming a checksum existed at evaluation time.
33
+
34
+ ## Architecture and label order
35
+
36
+ `Linear(384,128) -> ReLU -> Dropout(0.3) -> Linear(128,25)`;
37
+ 52,505 trainable parameters. Call `eval()` to disable dropout.
38
+ The frozen `model.py` is included verbatim. Output coordinates, in order:
39
+
40
+ ```text
41
+ 0:11 topics:
42
+ Software_&_Tech, STEM_&_Academics, Business_&_Finance,
43
+ Creative_&_Media, Admin_&_Productivity, Lifestyle_&_Health,
44
+ Social_&_Relationships, World_&_Current_Events, Meta_AI,
45
+ Null_Noise, General_Reference_&_Trivia
46
+
47
+ 11:22 intents:
48
+ Factual_Retrieval, Troubleshooting, Generation, Ideation,
49
+ Analysis_&_Summarization, Strategic_Planning, Decision_Making,
50
+ Emotional_Processing, Utility_Formatting, Casual_Banter, Open_Exploration
51
+
52
+ 22:25 context:
53
+ Zero_Shot, Long_Term_Memory, Real_Time_Search
54
+ ```
55
+
56
+ Topic and intent use sigmoid, a strict `> 0.3` threshold, and an argmax
57
+ fallback when a block has no selected label. Context uses softmax and argmax,
58
+ not independent sigmoids. `max_confidence` is the largest of the 25 decoded
59
+ probabilities. Exact order and settings are also in `config.json`.
60
+
61
+ ## Embeddings and preprocessing
62
+
63
+ Use frozen `Qwen/Qwen3-Embedding-0.6B`, revision
64
+ `97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3`, with
65
+ `SentenceTransformer(..., device="cpu", truncate_dim=384)`.
66
+ The upstream native width is 1024; ICE v2 takes its first 384 coordinates.
67
+ Use the snapshot's built-in pooling and normalization. Do **not** add
68
+ `normalize_embeddings=True` to `encode()`, renormalize the truncated prefix,
69
+ add a query instruction, or use the current ICE v3 native-width embedding path.
70
+
71
+ Embed the exact instructional prefix produced by `build_input()` in the included
72
+ `classifier_inference.py`, not the bare user prompt. With context, ICE v2
73
+ selects the last three episodic rows in timestamp order, prefers each summary,
74
+ otherwise uses raw text capped at 150 whitespace words, and caps the combined
75
+ context at 500 words. The frozen `frozen_classifier.py` preserves the exact
76
+ context selection and truncation behavior, including ellipses and the
77
+ context-specific prefix. An empty context uses the no-context prefix.
78
+
79
+ ## Minimal learned-head inference
80
+
81
+ The frozen tag's dependency versions are `torch==2.11.0`,
82
+ `sentence-transformers==5.5.1`, and `transformers==5.9.0`.
83
+ Use Python 3.11 and `huggingface_hub` to fetch the release:
84
+
85
+ ```python
86
+ import hashlib
87
+ import sys
88
+ from pathlib import Path
89
+ import torch
90
+ from huggingface_hub import snapshot_download
91
+ from sentence_transformers import SentenceTransformer
92
+
93
+ folder = Path(snapshot_download("Deepnar/ice-v2-classifier"))
94
+ sys.path.insert(0, str(folder))
95
+ from model import ICEClassifier
96
+ from classifier_inference import predict_head
97
+
98
+ checkpoint = folder / "ice_classifier_v3_qwen_ft3.pt"
99
+ assert hashlib.sha256(checkpoint.read_bytes()).hexdigest() == (
100
+ "25c758b6a7e5cf449f3e4c8bb250db759d37cb4f0ab7dd8e0c1acd8afbf05831"
101
+ )
102
+ head = ICEClassifier()
103
+ head.load_state_dict(torch.load(checkpoint, map_location="cpu", weights_only=True))
104
+ head.eval()
105
+ embedder = SentenceTransformer(
106
+ "Qwen/Qwen3-Embedding-0.6B",
107
+ revision="97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3",
108
+ device="cpu", truncate_dim=384,
109
+ )
110
+ print(predict_head(head, embedder, "Explain how a database index works."))
111
+ ```
112
+
113
+ For archival reuse, pin `snapshot_download(..., revision=<release commit>)`
114
+ to the upload commit recorded in the GitHub release verification report.
115
+ This is a custom PyTorch head, not a Transformers `AutoModel` or hosted pipeline.
116
+ The example returns **learned-head predictions**. The complete ICE v2 classifier
117
+ also runs the DI3 pre-classifier, hard overrides, and API-level memory policy;
118
+ reproduce those through the frozen Git repository. Neither this helper nor the
119
+ weights alone reproduce full-system routing or the paper's end-to-end scores.
120
+
121
+ ## Limitations and license
122
+
123
+ No new classifier accuracy claim is made by this release. The paper's fidelity
124
+ audit and negative results remain applicable. Context, preprocessing, rule
125
+ overrides and workload affect behavior. Training data and private conversational
126
+ corpora are not released, so this is checkpoint/inference reproducibility,
127
+ not a claim that private-data training can be independently regenerated.
128
+
129
+ The ICE head and accompanying ICE code are released under the repository's
130
+ Apache-2.0 license; `LICENSE` and `NOTICE` are included. The separately fetched
131
+ Qwen base model is also Apache-2.0 according to its pinned upstream model card.
132
+ No Qwen base weights, private corpora, caches, secrets, or ICE v3 artifacts are included.
classifier_inference.py ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """ICE v2 learned-head input and decoding, without database or DI3 rules."""
2
+
3
+ import json
4
+ from pathlib import Path
5
+
6
+ import torch
7
+
8
+
9
+ def build_input(prompt: str, context_text: str | None = None) -> str:
10
+ if context_text:
11
+ return (
12
+ f"Conversation context (summarized):\n{context_text}\n\n"
13
+ "Given the above conversation and the user's latest prompt, predict:\n"
14
+ "1. TOPIC: what is the subject (Software_&_Tech, Creative_&_Media, etc.)\n"
15
+ "2. INTENT: what is the user trying to do (Factual_Retrieval, Troubleshooting, etc.)\n"
16
+ "3. CONTEXT RELIANCE: does the user need memory (Zero_Shot, Long_Term_Memory, Real_Time_Search)\n\n"
17
+ f"User prompt: {prompt}"
18
+ )
19
+ return (
20
+ "Given a user prompt, predict:\n"
21
+ "1. TOPIC: what is the subject (Software_&_Tech, Creative_&_Media, etc.)\n"
22
+ "2. INTENT: what is the user trying to do (Factual_Retrieval, Troubleshooting, etc.)\n"
23
+ "3. CONTEXT RELIANCE: does the user need memory (Zero_Shot, Long_Term_Memory, Real_Time_Search)\n\n"
24
+ f"User prompt: {prompt}"
25
+ )
26
+
27
+
28
+ def predict_head(model, embedder, prompt: str, context_text: str | None = None):
29
+ config = json.loads(Path(__file__).with_name("config.json").read_text())
30
+ embedding = embedder.encode(build_input(prompt, context_text), convert_to_tensor=True)
31
+ if embedding.shape != (384,):
32
+ raise ValueError("ICE v2 requires exactly 384 embedding coordinates")
33
+ with torch.no_grad():
34
+ outputs = model(embedding.unsqueeze(0).float())
35
+ topic = torch.sigmoid(outputs[0, :11])
36
+ intent = torch.sigmoid(outputs[0, 11:22])
37
+ context = torch.softmax(outputs[0, 22:], dim=0)
38
+ def tags(probs, labels):
39
+ return [labels[i] for i, p in enumerate(probs) if p > 0.3] or [labels[probs.argmax().item()]]
40
+ probabilities = topic.tolist() + intent.tolist() + context.tolist()
41
+ return {
42
+ "topic_tags": tags(topic, config["topic_labels"]),
43
+ "intent_tags": tags(intent, config["intent_labels"]),
44
+ "context_reliance": config["context_labels"][context.argmax().item()],
45
+ "raw_probs": probabilities,
46
+ "max_confidence": max(probabilities),
47
+ }
config.json ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "system_version": "ICE v2",
3
+ "source_tag": "v2-paper-eval",
4
+ "source_commit": "0521df9171b4a7d69f82d12d70497138c77b2678",
5
+ "checkpoint": "ice_classifier_v3_qwen_ft3.pt",
6
+ "sha256": "25c758b6a7e5cf449f3e4c8bb250db759d37cb4f0ab7dd8e0c1acd8afbf05831",
7
+ "base_model": "Qwen/Qwen3-Embedding-0.6B",
8
+ "base_revision": "97b0c614be4d77ee51c0cef4e5f07c00f9eb65b3",
9
+ "native_embedding_width": 1024,
10
+ "embedding_width": 384,
11
+ "normalize_embeddings": false,
12
+ "license": "apache-2.0",
13
+ "architecture": "ICEClassifier",
14
+ "layers": [
15
+ 384,
16
+ 128,
17
+ 25
18
+ ],
19
+ "dropout": 0.3,
20
+ "topic_labels": [
21
+ "Software_&_Tech",
22
+ "STEM_&_Academics",
23
+ "Business_&_Finance",
24
+ "Creative_&_Media",
25
+ "Admin_&_Productivity",
26
+ "Lifestyle_&_Health",
27
+ "Social_&_Relationships",
28
+ "World_&_Current_Events",
29
+ "Meta_AI",
30
+ "Null_Noise",
31
+ "General_Reference_&_Trivia"
32
+ ],
33
+ "intent_labels": [
34
+ "Factual_Retrieval",
35
+ "Troubleshooting",
36
+ "Generation",
37
+ "Ideation",
38
+ "Analysis_&_Summarization",
39
+ "Strategic_Planning",
40
+ "Decision_Making",
41
+ "Emotional_Processing",
42
+ "Utility_Formatting",
43
+ "Casual_Banter",
44
+ "Open_Exploration"
45
+ ],
46
+ "context_labels": [
47
+ "Zero_Shot",
48
+ "Long_Term_Memory",
49
+ "Real_Time_Search"
50
+ ],
51
+ "topic_intent_threshold": 0.3,
52
+ "threshold_comparison": ">",
53
+ "empty_tag_fallback": "argmax",
54
+ "context_decode": "softmax then argmax",
55
+ "scope": "learned head; full classifier also uses DI3, hard overrides and API policy"
56
+ }
frozen_classifier.py ADDED
@@ -0,0 +1,202 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import torch
2
+ from sentence_transformers import SentenceTransformer
3
+ from typing import List, Optional
4
+ from .model import ICEClassifier
5
+ from .di3 import run_di3
6
+ from .schemas import ClassificationResult
7
+ from sqlalchemy.orm import Session
8
+
9
+
10
+ class PyTorchClassifier:
11
+ def __init__(self, model_path="models/classifier/ice_classifier.pt",
12
+ schema_path="data/labeled/label_schema.json"):
13
+ # These lists are fixed – the order must match training
14
+ self.TOPIC_LABELS = [
15
+ "Software_&_Tech", "STEM_&_Academics", "Business_&_Finance",
16
+ "Creative_&_Media", "Admin_&_Productivity", "Lifestyle_&_Health",
17
+ "Social_&_Relationships", "World_&_Current_Events", "Meta_AI",
18
+ "Null_Noise", "General_Reference_&_Trivia"
19
+ ]
20
+ self.INTENT_LABELS = [
21
+ "Factual_Retrieval", "Troubleshooting", "Generation", "Ideation",
22
+ "Analysis_&_Summarization", "Strategic_Planning", "Decision_Making",
23
+ "Emotional_Processing", "Utility_Formatting", "Casual_Banter",
24
+ "Open_Exploration"
25
+ ]
26
+ self.CONTEXT_RELIANCE_LABELS = [
27
+ "Zero_Shot", "Long_Term_Memory", "Real_Time_Search"
28
+ ]
29
+
30
+ # Load model on CPU
31
+ self.model = ICEClassifier()
32
+ self.model.load_state_dict(torch.load(model_path, map_location=torch.device('cpu')))
33
+ self.model.eval()
34
+
35
+ # Embedder also on CPU
36
+ # Embedder also on CPU – Qwen3-Embedding truncated to 384 dim for compatibility
37
+ self.embedder = SentenceTransformer(
38
+ "Qwen/Qwen3-Embedding-0.6B",
39
+ device="cpu",
40
+ truncate_dim=384
41
+ )
42
+
43
+ def _get_context_turns(self, conversation_id: str, n: int = 3, max_total_words: int = 500) -> str:
44
+ """Return a truncated, summary‑preferring context string from the last *n* turns."""
45
+ # Local import to avoid circular dependency at module level
46
+ from src.api.db import SessionLocal
47
+ db = SessionLocal()
48
+ try:
49
+ from src.memory.models import EpisodicMemory
50
+ turns = (
51
+ db.query(EpisodicMemory)
52
+ .filter_by(conversation_id=conversation_id)
53
+ .order_by(EpisodicMemory.timestamp.desc())
54
+ .limit(n)
55
+ .all()
56
+ )
57
+ turns.reverse()
58
+ parts = []
59
+ total_words = 0
60
+ for t in turns:
61
+ # Prefer summary, fall back to raw text (truncated)
62
+ text = t.summary_text or ""
63
+ if not text and t.raw_text:
64
+ words = t.raw_text.split()
65
+ text = " ".join(words[:150]) + "…" if len(words) > 150 else t.raw_text
66
+ if not text:
67
+ continue
68
+ word_count = len(text.split())
69
+ if total_words + word_count > max_total_words:
70
+ remaining = max_total_words - total_words
71
+ if remaining > 20:
72
+ w = text.split()
73
+ text = " ".join(w[:remaining]) + "…"
74
+ parts.append(text)
75
+ break
76
+ parts.append(text)
77
+ total_words += word_count
78
+ return "\n".join(parts)
79
+ finally:
80
+ db.close()
81
+
82
+ # ------------------------------------------------------------------
83
+ # Main entry point
84
+ # ------------------------------------------------------------------
85
+ def classify(
86
+ self,
87
+ prompt: str,
88
+ conversation_history: Optional[List[str]] = None,
89
+ conversation_length: int = 0,
90
+ conversation_id: Optional[str] = None,
91
+ ) -> ClassificationResult:
92
+ """Public entry point. Runs DI3 first, falls back to ML.
93
+ When *conversation_id* is given, the last 3 turns are used as context
94
+ (auto‑truncated) to improve the ML classifier's accuracy.
95
+ """
96
+ if conversation_history is None:
97
+ conversation_history = []
98
+ di3_result = run_di3(prompt, conversation_length, conversation_history)
99
+ if di3_result is not None:
100
+ # If DI3 forced LTM but left topic/intent blank, let the ML
101
+ # classifier provide the actual tags while keeping the LTM decision.
102
+ if not di3_result.topic_tags or not di3_result.intent_tags:
103
+ ml_result = self._run_ml_classifier(prompt, conversation_id)
104
+ if di3_result.context_reliance == "Long_Term_Memory":
105
+ ml_result.context_reliance = "Long_Term_Memory"
106
+ return self._apply_hard_overrides(ml_result, prompt)
107
+ else:
108
+ return self._apply_hard_overrides(di3_result, prompt)
109
+
110
+ return self._run_ml_classifier(prompt, conversation_id)
111
+
112
+ def _run_ml_classifier(self, prompt: str, conversation_id: Optional[str] = None) -> ClassificationResult:
113
+ """Original ML classification path (now private)."""
114
+ with torch.no_grad():
115
+ # Build context text if conversation_id is available
116
+ context_text = None
117
+ if conversation_id:
118
+ try:
119
+ context_text = self._get_context_turns(conversation_id)
120
+ except Exception:
121
+ context_text = None
122
+
123
+ if context_text:
124
+ prefixed_prompt = (
125
+ f"Conversation context (summarized):\n{context_text}\n\n"
126
+ f"Given the above conversation and the user's latest prompt, "
127
+ f"predict:\n"
128
+ f"1. TOPIC: what is the subject (Software_&_Tech, Creative_&_Media, etc.)\n"
129
+ f"2. INTENT: what is the user trying to do (Factual_Retrieval, Troubleshooting, etc.)\n"
130
+ f"3. CONTEXT RELIANCE: does the user need memory (Zero_Shot, Long_Term_Memory, Real_Time_Search)\n\n"
131
+ f"User prompt: {prompt}"
132
+ )
133
+ else:
134
+ prefixed_prompt = (
135
+ f"Given a user prompt, predict:\n"
136
+ f"1. TOPIC: what is the subject (Software_&_Tech, Creative_&_Media, etc.)\n"
137
+ f"2. INTENT: what is the user trying to do (Factual_Retrieval, Troubleshooting, etc.)\n"
138
+ f"3. CONTEXT RELIANCE: does the user need memory (Zero_Shot, Long_Term_Memory, Real_Time_Search)\n\n"
139
+ f"User prompt: {prompt}"
140
+ )
141
+ embedding = self.embedder.encode(prefixed_prompt, convert_to_tensor=True).unsqueeze(0).float()
142
+ outputs = self.model(embedding) # (1, 25)
143
+
144
+ topic_out = outputs[:, :11] # (1, 11)
145
+ intent_out = outputs[:, 11:22] # (1, 11)
146
+ ctx_out = outputs[:, 22:] # (1, 3)
147
+
148
+ topic_probs = torch.sigmoid(topic_out).squeeze(0) # (11,)
149
+ intent_probs = torch.sigmoid(intent_out).squeeze(0) # (11,)
150
+ ctx_probs = torch.softmax(ctx_out, dim=1).squeeze(0) # (3,)
151
+
152
+ # Build tag lists
153
+ topic_tags = [self.TOPIC_LABELS[i] for i in range(len(self.TOPIC_LABELS))
154
+ if topic_probs[i] > 0.3]
155
+ intent_tags = [self.INTENT_LABELS[i] for i in range(len(self.INTENT_LABELS))
156
+ if intent_probs[i] > 0.3]
157
+ if not topic_tags:
158
+ topic_tags = [self.TOPIC_LABELS[torch.argmax(topic_probs).item()]]
159
+ if not intent_tags:
160
+ intent_tags = [self.INTENT_LABELS[torch.argmax(intent_probs).item()]]
161
+ context_reliance = self.CONTEXT_RELIANCE_LABELS[torch.argmax(ctx_probs).item()]
162
+
163
+ # Combine probabilities
164
+ raw_probs = topic_probs.tolist() + intent_probs.tolist() + ctx_probs.tolist()
165
+ max_confidence = max(raw_probs)
166
+
167
+ result = ClassificationResult(
168
+ topic_tags=topic_tags,
169
+ intent_tags=intent_tags,
170
+ context_reliance=context_reliance,
171
+ raw_probs=raw_probs,
172
+ max_confidence=max_confidence,
173
+ prompt=prompt,
174
+ )
175
+ return self._apply_hard_overrides(result, prompt)
176
+
177
+ def _apply_hard_overrides(
178
+ self, result: ClassificationResult, prompt: str
179
+ ) -> ClassificationResult:
180
+ """Apply creative/software LTM overrides, but never downgrade an
181
+ existing Long_Term_Memory decision (e.g. from DI3 or LTM bias)."""
182
+
183
+ # If LTM has already been enforced (by DI3 or API‑level bias), keep it
184
+ if result.context_reliance == "Long_Term_Memory":
185
+ return result
186
+
187
+ if "Creative_&_Media" in result.topic_tags:
188
+ result.context_reliance = "Long_Term_Memory"
189
+
190
+ if "Software_&_Tech" in result.topic_tags:
191
+ referential_words = [
192
+ "my", "our", "mine", "ours", "we", "us",
193
+ "this", "that", "these", "those", "the",
194
+ "it", "they", "them", "their",
195
+ "previous", "last", "before", "yesterday", "earlier",
196
+ "again", "still", "same",
197
+ ]
198
+ prompt_lower = prompt.lower()
199
+ if any(word in prompt_lower for word in referential_words):
200
+ result.context_reliance = "Long_Term_Memory"
201
+
202
+ return result
ice_classifier_v3_qwen_ft3.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:25c758b6a7e5cf449f3e4c8bb250db759d37cb4f0ab7dd8e0c1acd8afbf05831
3
+ size 212827
model.py ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # src/classifier/model.py
2
+
3
+ import torch
4
+ import torch.nn as nn
5
+
6
+ class ICEClassifier(nn.Module):
7
+ def __init__(self):
8
+ super(ICEClassifier, self).__init__()
9
+ self.fc1 = nn.Linear(384, 128)
10
+ self.relu = nn.ReLU()
11
+ self.dropout = nn.Dropout(0.3)
12
+ self.fc2 = nn.Linear(128, 25)
13
+
14
+ def forward(self, x):
15
+ x = self.fc1(x)
16
+ x = self.relu(x)
17
+ x = self.dropout(x)
18
+ x = self.fc2(x)
19
+ return x