File size: 10,944 Bytes
0f5dd4e
 
 
 
 
 
 
 
e44d5cd
 
 
0f5dd4e
 
 
c1c1afb
 
 
0f5dd4e
f5e7902
0f5dd4e
f5e7902
0f5dd4e
 
 
 
f5e7902
 
0f5dd4e
 
 
 
e44d5cd
c1c1afb
f5e7902
 
 
0f5dd4e
 
 
 
f5e7902
e44d5cd
f5e7902
e44d5cd
f5e7902
 
 
 
 
 
 
e44d5cd
 
 
f5e7902
 
 
0f5dd4e
f5e7902
0f5dd4e
f5e7902
0f5dd4e
f5e7902
 
 
 
 
 
 
0f5dd4e
 
 
f5e7902
 
 
 
 
 
 
0f5dd4e
 
f5e7902
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0f5dd4e
f5e7902
0f5dd4e
f5e7902
0f5dd4e
f5e7902
c1c1afb
f5e7902
0f5dd4e
f5e7902
0f5dd4e
 
f5e7902
 
0f5dd4e
f5e7902
 
0f5dd4e
f5e7902
0f5dd4e
f5e7902
0f5dd4e
f5e7902
 
 
 
77c13ac
 
 
f5e7902
77c13ac
f5e7902
0f5dd4e
f5e7902
 
 
 
0f5dd4e
f5e7902
0f5dd4e
 
f5e7902
7bdae5f
f5e7902
7bdae5f
f5e7902
7bdae5f
f5e7902
7bdae5f
f5e7902
7bdae5f
c1c1afb
f5e7902
0f5dd4e
f5e7902
 
 
7bdae5f
f5e7902
 
 
c1c1afb
f5e7902
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0f5dd4e
 
f5e7902
 
 
7bdae5f
f5e7902
 
 
 
 
 
 
 
 
 
 
 
7bdae5f
0f5dd4e
 
f5e7902
 
 
0f5dd4e
 
f5e7902
 
 
 
 
 
 
0f5dd4e
 
 
f5e7902
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
---
language:
- en
license: cc-by-nc-nd-4.0
license_name: cc-by-nc-nd-4.0
license_link: LICENSE
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
tags:
- aazhi
- aazhi-coder
- aazhi-1.5b
- isom
- isom-r2
- qwen2.5-coder
- 1m-context
- 1048576-tokens
- million-context
- bounded-memory
- code-intelligence
- repository-level
- on-premise
- edge-ai
pipeline_tag: text-generation
---

# ๐ŸŒŠ Aazhi-Coder-1.5B (เฎ†เฎดเฎฟ): 1,048,576-Token Bounded-Memory Code Intelligence
### 1M Codebase Streaming • 3.24 GB Peak VRAM • 100% Retrieval Precision • Powered by ISOM-R2

<p align="center">
  <a href="https://doi.org/10.5281/zenodo.14925828"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.14925828.svg" alt="DOI"></a>
  <a href="https://www.linkedin.com/in/prannesshkva/"><img src="https://img.shields.io/badge/LinkedIn-Prannesh_K._V._A.-blue?logo=linkedin" alt="LinkedIn"></a>
  <img src="https://img.shields.io/badge/Model-Aazhi--Coder--1.5B-cyan.svg" alt="Aazhi">
  <img src="https://img.shields.io/badge/Context-1%2C048%2C576_Tokens_(1M)-blue.svg" alt="Context">
  <img src="https://img.shields.io/badge/Peak_VRAM-3.24_GB-brightgreen.svg" alt="Peak VRAM">
  <img src="https://img.shields.io/badge/Ingestion_Speed-8%2C526_tok%2Fs-orange.svg" alt="Ingestion Speed">
  <img src="https://img.shields.io/badge/Audited_Run-1%2C055%2C402_Tokens-purple.svg" alt="Audited Run">
</p>

---

## Executive Summary

**Aazhi-Coder-1.5B** (named after *Aazhi* [เฎ†เฎดเฎฟ] โ€” representing the boundless ocean) is a production-grade 1,048,576-token (1 Million) long-context coding model built on Alibaba's exceptional `Qwen2.5-Coder-1.5B-Instruct` foundation and powered by the **ISOM-R2 (Isometric State Space / Virtual SVD)** memory architecture.

Standard Transformer models suffer catastrophic memory bottlenecks at million-token scales. For a 1.05M-token sequence, a conventional FP16 Key-Value cache demands **28.2 GiB of VRAM** before allocating weights, instantly crashing single GPUs.

**Aazhi-Coder-1.5B completely eliminates the 28.2 GiB KV cache explosion:**
* **3.24 GB Peak VRAM:** Streams and processes **1,055,402 continuous tokens** within a strictly bounded memory pool ($< 3.5\text{ GB}$).
* **123.78s Ingestion Time:** Reaches an effective throughput of **~8,526 tokens/second**, digesting 181 production repository files in approximately 2 minutes.
* **100% Macro-Retrieval Precision:** Correctly isolates target files across 516 candidate chunks with zero distractor noise.
* **Reproducible Execution Proof:** Verifiable directly via the audited benchmark notebook [`ISOM_R2.ipynb`](./ISOM_R2.ipynb).

---

## โšก The Memory Bottleneck: Conventional Attention vs. ISOM-R2

In standard Grouped-Query Attention (GQA) architectures (28 layers, 2 KV heads, head dimension 128), KV cache memory scales linearly:

$$\text{KV Cache Size} = 2 \times L \times H_{KV} \times D_{\text{head}} \times N_{\text{tokens}} \times 2 \text{ bytes}$$

$$\text{At } 1,055,402 \text{ tokens: } 2 \times 28 \times 2 \times 128 \times 1,055,402 \times 2 \approx \mathbf{28.18\text{ GiB}}$$

| Parameter | Standard Transformer GQA | Aazhi-Coder-1.5B (ISOM-R2) | Advantage |
| :--- | :---: | :---: | :---: |
| **KV Cache Footprint (1.05M tokens)** | **28.18 GiB** (Linear $O(N)$) | **4,160 Active Tokens (~0.11 GiB)** | **99.6% Reduction** |
| **Prefill VRAM Behavior** | Explodes to OOM | **Flat 2.97 GB Invariant** | Bounded State Space |
| **Peak Execution VRAM** | $> 33.5\text{ GB}$ (A100 required) | **3.24 GB Peak** | **8.7ร— Total Compression** |
| **Target Hardware** | 40GB / 80GB Data Center GPUs | **Consumer 4GB / 6GB / 8GB GPUs** | On-premise / Laptop deployment |
| **Throughput (1.05M tokens)** | Quadratic deceleration | **123.78 seconds (~8,526 tok/s)** | Constant-time streaming |

---

## ๐Ÿ“Š Audited Production Benchmark: 1,055,402 Real Tokens

Benchmark executed on real Python source files cloned from the official **Hugging Face `transformers`** repository (**zero synthetic tokens, zero artificial padding**):

* **Corpus Scope:** 181 production source files (covering `models/`, `pipelines/`, `generation/`, `trainer/`)
* **Prompt Assembly:** 1,055,402 total sequence tokens across 516 discrete 2,048-token chunks
* **Target Multi-Hop Dependency:** Cross-file inquiry spanning `configuration_llama.py` (Chunk 218) and `modeling_llama.py` (Chunk 217)

```text
[ISOM-R2 Engine] Streaming 1,055,192 codebase context tokens across 516 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill  210,944 / 1,055,192 tokens (20.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  421,888 / 1,055,192 tokens (40.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  632,832 / 1,055,192 tokens (60.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  843,776 / 1,055,192 tokens (80.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 1,054,720 / 1,055,192 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2 Engine] Micro-window sliced: Chunk 218, anchor=1030, range=[518:1542] (1024 tokens)
[ISOM-R2 Engine] Micro-window sliced: Chunk 217, anchor=1010, range=[498:1522] (1024 tokens)
[ISOM-R2] Retrieved salient context pages: [218, 217] | Active KV: 4160 tokens

============================================================
1,000,000-TOKEN MULTI-HOP GENERATION RESULTS:
============================================================
Input Sequence Length : 1,055,402 tokens
Total Time Taken      : 123.78s
Peak GPU VRAM         : 3.24 GB
============================================================
```

### Retrieval & Memory Provenance
* **Macro-Retrieval Precision:** **100%**. Out of 516 candidate chunks across 181 files, only **Chunk 217** and **Chunk 218** were paged. Zero distractor chunks were retrieved.
* **Micro-Window Anchoring:** Dynamic anchor-centered slicing restricted active attention to 1,024 tokens per salient chunk, pinning active KV cache memory to **4,160 tokens**.

---

## ๐Ÿ—๏ธ Architecture: ISOM-R2 Under the Hood

```text
1,048,576 Token Stream (Codebase)
       โ”‚
       โ”œโ”€โ”€โ–บ [ 1. Streaming Bounded State Ingestion ] โ”€โ”€โ”€โ”€โ”€โ”€โ–บ 2,048-token chunk prefill @ flat 2.97 GB VRAM
       โ”‚
       โ”œโ”€โ”€โ–บ [ 2. Sub-Harmonic Lie Frequency Calibration ] โ”€โ–บ Zero phase aliasing across 1M tokens (ฯ‰_min = 5.99e-6)
       โ”‚
       โ”œโ”€โ”€โ–บ [ 3. Hybrid Salient Context Retrieval ] โ”€โ”€โ”€โ”€โ”€โ”€โ–บ Filters 516 chunks down to top-k relevant files
       โ”‚
       โ”œโ”€โ”€โ–บ [ 4. Anchor-Centered Micro-Window Slicing ] โ”€โ”€โ–บ Paged attention restricted to 1,024 tokens/chunk
       โ”‚
       โ””โ”€โ”€โ–บ [ 5. Focused Multi-Hop Decoding ] โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ High-fidelity generation with < 3.24 GB peak VRAM
```

### 1. Bounded State Space Ingestion
Rather than storing every intermediate key and value tensor in video memory, ISOM-R2 processes input sequences through bounded state projections. Token activations are compressed into isometric manifold representations, keeping GPU VRAM strictly flat at **2.97 GB** throughout prefill.

### 2. Sub-Harmonic Lie Calibration ($\omega_{\min}$)
Standard rotational embeddings experience severe phase wrap and aliasing past 32k tokens. ISOM-R2 establishes a calibrated sub-harmonic frequency floor:

$$\omega_{\min} < \frac{2\pi}{1,048,576} \approx 5.9921 \times 10^{-6} \text{ rad/token}$$

This guarantees that the slowest coordinate manifold rotates strictly less than one complete revolution across the entire 1,048,576-token sequence, preserving stable temporal geometry.

### 3. Salient Micro-Window Paging
During generation, ISOM-R2 performs two-tiered context activation:
1. **Macro Level:** Filters the repository down to the most relevant code modules.
2. **Micro Level:** Automatically isolates anchor-centered micro-windows (default 1,024 tokens) containing critical function implementations, enabling fast generation without attending over full megatoken buffers.

---

## ๐Ÿš€ Quickstart: Drop-In Usage

`Aazhi-Coder-1.5B` integrates natively with Hugging Face `transformers` using `trust_remote_code=True`.

### Installation
```bash
pip install -q transformers>=4.49.0 accelerate torch
```

### Loading the Model
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "Prannesshkva/Aazhi-Coder-1.5B"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    torch_dtype=torch.float16,
    device_map="cuda",
)
model.eval()

# Configure bounded retrieval policy
model.config.isom_r2_num_retrieved_chunks = 2
model.config.isom_r2_micro_window_size = 1024

print(f"Model loaded on: {next(model.parameters()).device}")
print(f"Base VRAM: {torch.cuda.memory_allocated(0)/(1024**3):.2f} GB")
```

### Ingesting & Querying a Massive Codebase
```python
# prompt containing 100k to 1,000,000+ tokens of repository code
inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    output = model.generate(
        inputs["input_ids"],
        max_new_tokens=512,
        num_retrieved_chunks=2,
        micro_window_size=1024,
        do_sample=False,
        tokenizer=tokenizer,
    )

response = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)
```

---

## ๐Ÿ“ Repository Structure

```text
โ”œโ”€โ”€ config.json                             # Model hyperparameters & ISOM-R2 defaults
โ”œโ”€โ”€ modeling_isom_qwen25_coder.py           # Core ISOM-R2 engine with Aazhi architectural classes
โ”œโ”€โ”€ configuration_isom_qwen25_coder.py      # Configuration wrapper
โ”œโ”€โ”€ isom_r2_module.py                       # Virtual SVD & Lie-algebra manifold operators
โ”œโ”€โ”€ ISOM_R2.ipynb                           # Audited 1,055,402-token Colab benchmark execution proof
โ”œโ”€โ”€ benchmarks/
โ”‚   โ”œโ”€โ”€ ISOM_R2_1M_Benchmark.ipynb          # Verified Colab benchmark notebook
โ”‚   โ”œโ”€โ”€ audited_systems_benchmark_qwen25_coder.json
โ”‚   โ””โ”€โ”€ isom_r2_1m_real_repo_benchmark_results.json
โ””โ”€โ”€ model.safetensors                       # Model weights (Native FP16, ~3.09 GB)
```

---

## ๐Ÿ“œ Citation & Attribution

If you use `Aazhi-Coder-1.5B` or the ISOM-R2 architecture in your research or applications, please cite:

```bibtex
@software{aazhi_coder_2026,
  author       = {Prannesh K. V. A.},
  title        = {Aazhi-Coder-1.5B: 1,048,576-Token Bounded-Memory Recurrent Code Intelligence},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.5281/zenodo.14925828},
  url          = {https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B}
}
```

### Acknowledgements
`Aazhi-Coder-1.5B` is built upon the `Qwen2.5-Coder-1.5B-Instruct` model developed by the Qwen team at Alibaba Cloud, released under the Apache 2.0 license.