File size: 10,944 Bytes
0f5dd4e e44d5cd 0f5dd4e c1c1afb 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e e44d5cd c1c1afb f5e7902 0f5dd4e f5e7902 e44d5cd f5e7902 e44d5cd f5e7902 e44d5cd f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 c1c1afb f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 77c13ac f5e7902 77c13ac f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 7bdae5f f5e7902 7bdae5f f5e7902 7bdae5f f5e7902 7bdae5f f5e7902 7bdae5f c1c1afb f5e7902 0f5dd4e f5e7902 7bdae5f f5e7902 c1c1afb f5e7902 0f5dd4e f5e7902 7bdae5f f5e7902 7bdae5f 0f5dd4e f5e7902 0f5dd4e f5e7902 0f5dd4e f5e7902 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 | ---
language:
- en
license: cc-by-nc-nd-4.0
license_name: cc-by-nc-nd-4.0
license_link: LICENSE
base_model: Qwen/Qwen2.5-Coder-1.5B-Instruct
tags:
- aazhi
- aazhi-coder
- aazhi-1.5b
- isom
- isom-r2
- qwen2.5-coder
- 1m-context
- 1048576-tokens
- million-context
- bounded-memory
- code-intelligence
- repository-level
- on-premise
- edge-ai
pipeline_tag: text-generation
---
# ๐ Aazhi-Coder-1.5B (เฎเฎดเฎฟ): 1,048,576-Token Bounded-Memory Code Intelligence
### 1M Codebase Streaming • 3.24 GB Peak VRAM • 100% Retrieval Precision • Powered by ISOM-R2
<p align="center">
<a href="https://doi.org/10.5281/zenodo.14925828"><img src="https://zenodo.org/badge/DOI/10.5281/zenodo.14925828.svg" alt="DOI"></a>
<a href="https://www.linkedin.com/in/prannesshkva/"><img src="https://img.shields.io/badge/LinkedIn-Prannesh_K._V._A.-blue?logo=linkedin" alt="LinkedIn"></a>
<img src="https://img.shields.io/badge/Model-Aazhi--Coder--1.5B-cyan.svg" alt="Aazhi">
<img src="https://img.shields.io/badge/Context-1%2C048%2C576_Tokens_(1M)-blue.svg" alt="Context">
<img src="https://img.shields.io/badge/Peak_VRAM-3.24_GB-brightgreen.svg" alt="Peak VRAM">
<img src="https://img.shields.io/badge/Ingestion_Speed-8%2C526_tok%2Fs-orange.svg" alt="Ingestion Speed">
<img src="https://img.shields.io/badge/Audited_Run-1%2C055%2C402_Tokens-purple.svg" alt="Audited Run">
</p>
---
## Executive Summary
**Aazhi-Coder-1.5B** (named after *Aazhi* [เฎเฎดเฎฟ] โ representing the boundless ocean) is a production-grade 1,048,576-token (1 Million) long-context coding model built on Alibaba's exceptional `Qwen2.5-Coder-1.5B-Instruct` foundation and powered by the **ISOM-R2 (Isometric State Space / Virtual SVD)** memory architecture.
Standard Transformer models suffer catastrophic memory bottlenecks at million-token scales. For a 1.05M-token sequence, a conventional FP16 Key-Value cache demands **28.2 GiB of VRAM** before allocating weights, instantly crashing single GPUs.
**Aazhi-Coder-1.5B completely eliminates the 28.2 GiB KV cache explosion:**
* **3.24 GB Peak VRAM:** Streams and processes **1,055,402 continuous tokens** within a strictly bounded memory pool ($< 3.5\text{ GB}$).
* **123.78s Ingestion Time:** Reaches an effective throughput of **~8,526 tokens/second**, digesting 181 production repository files in approximately 2 minutes.
* **100% Macro-Retrieval Precision:** Correctly isolates target files across 516 candidate chunks with zero distractor noise.
* **Reproducible Execution Proof:** Verifiable directly via the audited benchmark notebook [`ISOM_R2.ipynb`](./ISOM_R2.ipynb).
---
## โก The Memory Bottleneck: Conventional Attention vs. ISOM-R2
In standard Grouped-Query Attention (GQA) architectures (28 layers, 2 KV heads, head dimension 128), KV cache memory scales linearly:
$$\text{KV Cache Size} = 2 \times L \times H_{KV} \times D_{\text{head}} \times N_{\text{tokens}} \times 2 \text{ bytes}$$
$$\text{At } 1,055,402 \text{ tokens: } 2 \times 28 \times 2 \times 128 \times 1,055,402 \times 2 \approx \mathbf{28.18\text{ GiB}}$$
| Parameter | Standard Transformer GQA | Aazhi-Coder-1.5B (ISOM-R2) | Advantage |
| :--- | :---: | :---: | :---: |
| **KV Cache Footprint (1.05M tokens)** | **28.18 GiB** (Linear $O(N)$) | **4,160 Active Tokens (~0.11 GiB)** | **99.6% Reduction** |
| **Prefill VRAM Behavior** | Explodes to OOM | **Flat 2.97 GB Invariant** | Bounded State Space |
| **Peak Execution VRAM** | $> 33.5\text{ GB}$ (A100 required) | **3.24 GB Peak** | **8.7ร Total Compression** |
| **Target Hardware** | 40GB / 80GB Data Center GPUs | **Consumer 4GB / 6GB / 8GB GPUs** | On-premise / Laptop deployment |
| **Throughput (1.05M tokens)** | Quadratic deceleration | **123.78 seconds (~8,526 tok/s)** | Constant-time streaming |
---
## ๐ Audited Production Benchmark: 1,055,402 Real Tokens
Benchmark executed on real Python source files cloned from the official **Hugging Face `transformers`** repository (**zero synthetic tokens, zero artificial padding**):
* **Corpus Scope:** 181 production source files (covering `models/`, `pipelines/`, `generation/`, `trainer/`)
* **Prompt Assembly:** 1,055,402 total sequence tokens across 516 discrete 2,048-token chunks
* **Target Multi-Hop Dependency:** Cross-file inquiry spanning `configuration_llama.py` (Chunk 218) and `modeling_llama.py` (Chunk 217)
```text
[ISOM-R2 Engine] Streaming 1,055,192 codebase context tokens across 516 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 210,944 / 1,055,192 tokens (20.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 421,888 / 1,055,192 tokens (40.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 632,832 / 1,055,192 tokens (60.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 843,776 / 1,055,192 tokens (80.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 1,054,720 / 1,055,192 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2 Engine] Micro-window sliced: Chunk 218, anchor=1030, range=[518:1542] (1024 tokens)
[ISOM-R2 Engine] Micro-window sliced: Chunk 217, anchor=1010, range=[498:1522] (1024 tokens)
[ISOM-R2] Retrieved salient context pages: [218, 217] | Active KV: 4160 tokens
============================================================
1,000,000-TOKEN MULTI-HOP GENERATION RESULTS:
============================================================
Input Sequence Length : 1,055,402 tokens
Total Time Taken : 123.78s
Peak GPU VRAM : 3.24 GB
============================================================
```
### Retrieval & Memory Provenance
* **Macro-Retrieval Precision:** **100%**. Out of 516 candidate chunks across 181 files, only **Chunk 217** and **Chunk 218** were paged. Zero distractor chunks were retrieved.
* **Micro-Window Anchoring:** Dynamic anchor-centered slicing restricted active attention to 1,024 tokens per salient chunk, pinning active KV cache memory to **4,160 tokens**.
---
## ๐๏ธ Architecture: ISOM-R2 Under the Hood
```text
1,048,576 Token Stream (Codebase)
โ
โโโโบ [ 1. Streaming Bounded State Ingestion ] โโโโโโโบ 2,048-token chunk prefill @ flat 2.97 GB VRAM
โ
โโโโบ [ 2. Sub-Harmonic Lie Frequency Calibration ] โโบ Zero phase aliasing across 1M tokens (ฯ_min = 5.99e-6)
โ
โโโโบ [ 3. Hybrid Salient Context Retrieval ] โโโโโโโบ Filters 516 chunks down to top-k relevant files
โ
โโโโบ [ 4. Anchor-Centered Micro-Window Slicing ] โโโบ Paged attention restricted to 1,024 tokens/chunk
โ
โโโโบ [ 5. Focused Multi-Hop Decoding ] โโโโโโโโโโโโโบ High-fidelity generation with < 3.24 GB peak VRAM
```
### 1. Bounded State Space Ingestion
Rather than storing every intermediate key and value tensor in video memory, ISOM-R2 processes input sequences through bounded state projections. Token activations are compressed into isometric manifold representations, keeping GPU VRAM strictly flat at **2.97 GB** throughout prefill.
### 2. Sub-Harmonic Lie Calibration ($\omega_{\min}$)
Standard rotational embeddings experience severe phase wrap and aliasing past 32k tokens. ISOM-R2 establishes a calibrated sub-harmonic frequency floor:
$$\omega_{\min} < \frac{2\pi}{1,048,576} \approx 5.9921 \times 10^{-6} \text{ rad/token}$$
This guarantees that the slowest coordinate manifold rotates strictly less than one complete revolution across the entire 1,048,576-token sequence, preserving stable temporal geometry.
### 3. Salient Micro-Window Paging
During generation, ISOM-R2 performs two-tiered context activation:
1. **Macro Level:** Filters the repository down to the most relevant code modules.
2. **Micro Level:** Automatically isolates anchor-centered micro-windows (default 1,024 tokens) containing critical function implementations, enabling fast generation without attending over full megatoken buffers.
---
## ๐ Quickstart: Drop-In Usage
`Aazhi-Coder-1.5B` integrates natively with Hugging Face `transformers` using `trust_remote_code=True`.
### Installation
```bash
pip install -q transformers>=4.49.0 accelerate torch
```
### Loading the Model
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "Prannesshkva/Aazhi-Coder-1.5B"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="cuda",
)
model.eval()
# Configure bounded retrieval policy
model.config.isom_r2_num_retrieved_chunks = 2
model.config.isom_r2_micro_window_size = 1024
print(f"Model loaded on: {next(model.parameters()).device}")
print(f"Base VRAM: {torch.cuda.memory_allocated(0)/(1024**3):.2f} GB")
```
### Ingesting & Querying a Massive Codebase
```python
# prompt containing 100k to 1,000,000+ tokens of repository code
inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
output = model.generate(
inputs["input_ids"],
max_new_tokens=512,
num_retrieved_chunks=2,
micro_window_size=1024,
do_sample=False,
tokenizer=tokenizer,
)
response = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)
```
---
## ๐ Repository Structure
```text
โโโ config.json # Model hyperparameters & ISOM-R2 defaults
โโโ modeling_isom_qwen25_coder.py # Core ISOM-R2 engine with Aazhi architectural classes
โโโ configuration_isom_qwen25_coder.py # Configuration wrapper
โโโ isom_r2_module.py # Virtual SVD & Lie-algebra manifold operators
โโโ ISOM_R2.ipynb # Audited 1,055,402-token Colab benchmark execution proof
โโโ benchmarks/
โ โโโ ISOM_R2_1M_Benchmark.ipynb # Verified Colab benchmark notebook
โ โโโ audited_systems_benchmark_qwen25_coder.json
โ โโโ isom_r2_1m_real_repo_benchmark_results.json
โโโ model.safetensors # Model weights (Native FP16, ~3.09 GB)
```
---
## ๐ Citation & Attribution
If you use `Aazhi-Coder-1.5B` or the ISOM-R2 architecture in your research or applications, please cite:
```bibtex
@software{aazhi_coder_2026,
author = {Prannesh K. V. A.},
title = {Aazhi-Coder-1.5B: 1,048,576-Token Bounded-Memory Recurrent Code Intelligence},
year = {2026},
publisher = {Hugging Face},
doi = {10.5281/zenodo.14925828},
url = {https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B}
}
```
### Acknowledgements
`Aazhi-Coder-1.5B` is built upon the `Qwen2.5-Coder-1.5B-Instruct` model developed by the Qwen team at Alibaba Cloud, released under the Apache 2.0 license.
|