- 🌊 Aazhi-Coder-1.5B (ஆழி): 1,048,576-Token Bounded-Memory Code Intelligence
🌊 Aazhi-Coder-1.5B (ஆழி): 1,048,576-Token Bounded-Memory Code Intelligence
1M Codebase Streaming • 3.24 GB Peak VRAM • 100% Retrieval Precision • Powered by ISOM-R2
Executive Summary
Aazhi-Coder-1.5B (named after Aazhi [ஆழி] — representing the boundless ocean) is a production-grade 1,048,576-token (1 Million) long-context coding model built on Alibaba's exceptional Qwen2.5-Coder-1.5B-Instruct foundation and powered by the ISOM-R2 (Isometric State Space / Virtual SVD) memory architecture.
Standard Transformer models suffer catastrophic memory bottlenecks at million-token scales. For a 1.05M-token sequence, a conventional FP16 Key-Value cache demands 28.2 GiB of VRAM before allocating weights, instantly crashing single GPUs.
Aazhi-Coder-1.5B completely eliminates the 28.2 GiB KV cache explosion:
- 3.24 GB Peak VRAM: Streams and processes 1,055,402 continuous tokens within a strictly bounded memory pool ($< 3.5\text{ GB}$).
- 123.78s Ingestion Time: Reaches an effective throughput of ~8,526 tokens/second, digesting 181 production repository files in approximately 2 minutes.
- 100% Macro-Retrieval Precision: Correctly isolates target files across 516 candidate chunks with zero distractor noise.
- Reproducible Execution Proof: Verifiable directly via the audited benchmark notebook
ISOM_R2.ipynb.
⚡ The Memory Bottleneck: Conventional Attention vs. ISOM-R2
In standard Grouped-Query Attention (GQA) architectures (28 layers, 2 KV heads, head dimension 128), KV cache memory scales linearly:
| Parameter | Standard Transformer GQA | Aazhi-Coder-1.5B (ISOM-R2) | Advantage |
|---|---|---|---|
| KV Cache Footprint (1.05M tokens) | 28.18 GiB (Linear $O(N)$) | 4,160 Active Tokens (~0.11 GiB) | 99.6% Reduction |
| Prefill VRAM Behavior | Explodes to OOM | Flat 2.97 GB Invariant | Bounded State Space |
| Peak Execution VRAM | $> 33.5\text{ GB}$ (A100 required) | 3.24 GB Peak | 8.7× Total Compression |
| Target Hardware | 40GB / 80GB Data Center GPUs | Consumer 4GB / 6GB / 8GB GPUs | On-premise / Laptop deployment |
| Throughput (1.05M tokens) | Quadratic deceleration | 123.78 seconds (~8,526 tok/s) | Constant-time streaming |
📊 Audited Production Benchmark: 1,055,402 Real Tokens
Benchmark executed on real Python source files cloned from the official Hugging Face transformers repository (zero synthetic tokens, zero artificial padding):
- Corpus Scope: 181 production source files (covering
models/,pipelines/,generation/,trainer/) - Prompt Assembly: 1,055,402 total sequence tokens across 516 discrete 2,048-token chunks
- Target Multi-Hop Dependency: Cross-file inquiry spanning
configuration_llama.py(Chunk 218) andmodeling_llama.py(Chunk 217)
[ISOM-R2 Engine] Streaming 1,055,192 codebase context tokens across 516 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 210,944 / 1,055,192 tokens (20.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 421,888 / 1,055,192 tokens (40.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 632,832 / 1,055,192 tokens (60.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 843,776 / 1,055,192 tokens (80.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 1,054,720 / 1,055,192 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2 Engine] Micro-window sliced: Chunk 218, anchor=1030, range=[518:1542] (1024 tokens)
[ISOM-R2 Engine] Micro-window sliced: Chunk 217, anchor=1010, range=[498:1522] (1024 tokens)
[ISOM-R2] Retrieved salient context pages: [218, 217] | Active KV: 4160 tokens
============================================================
1,000,000-TOKEN MULTI-HOP GENERATION RESULTS:
============================================================
Input Sequence Length : 1,055,402 tokens
Total Time Taken : 123.78s
Peak GPU VRAM : 3.24 GB
============================================================
Retrieval & Memory Provenance
- Macro-Retrieval Precision: 100%. Out of 516 candidate chunks across 181 files, only Chunk 217 and Chunk 218 were paged. Zero distractor chunks were retrieved.
- Micro-Window Anchoring: Dynamic anchor-centered slicing restricted active attention to 1,024 tokens per salient chunk, pinning active KV cache memory to 4,160 tokens.
🏗️ Architecture: ISOM-R2 Under the Hood
1,048,576 Token Stream (Codebase)
│
├──► [ 1. Streaming Bounded State Ingestion ] ──────► 2,048-token chunk prefill @ flat 2.97 GB VRAM
│
├──► [ 2. Sub-Harmonic Lie Frequency Calibration ] ─► Zero phase aliasing across 1M tokens (ω_min = 5.99e-6)
│
├──► [ 3. Hybrid Salient Context Retrieval ] ──────► Filters 516 chunks down to top-k relevant files
│
├──► [ 4. Anchor-Centered Micro-Window Slicing ] ──► Paged attention restricted to 1,024 tokens/chunk
│
└──► [ 5. Focused Multi-Hop Decoding ] ────────────► High-fidelity generation with < 3.24 GB peak VRAM
1. Bounded State Space Ingestion
Rather than storing every intermediate key and value tensor in video memory, ISOM-R2 processes input sequences through bounded state projections. Token activations are compressed into isometric manifold representations, keeping GPU VRAM strictly flat at 2.97 GB throughout prefill.
2. Sub-Harmonic Lie Calibration ($\omega_{\min}$)
Standard rotational embeddings experience severe phase wrap and aliasing past 32k tokens. ISOM-R2 establishes a calibrated sub-harmonic frequency floor:
This guarantees that the slowest coordinate manifold rotates strictly less than one complete revolution across the entire 1,048,576-token sequence, preserving stable temporal geometry.
3. Salient Micro-Window Paging
During generation, ISOM-R2 performs two-tiered context activation:
- Macro Level: Filters the repository down to the most relevant code modules.
- Micro Level: Automatically isolates anchor-centered micro-windows (default 1,024 tokens) containing critical function implementations, enabling fast generation without attending over full megatoken buffers.
🚀 Quickstart: Drop-In Usage
Aazhi-Coder-1.5B integrates natively with Hugging Face transformers using trust_remote_code=True.
Installation
pip install -q transformers>=4.49.0 accelerate torch
Loading the Model
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL_ID = "Prannesshkva/Aazhi-Coder-1.5B"
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
MODEL_ID,
trust_remote_code=True,
torch_dtype=torch.float16,
device_map="cuda",
)
model.eval()
# Configure bounded retrieval policy
model.config.isom_r2_num_retrieved_chunks = 2
model.config.isom_r2_micro_window_size = 1024
print(f"Model loaded on: {next(model.parameters()).device}")
print(f"Base VRAM: {torch.cuda.memory_allocated(0)/(1024**3):.2f} GB")
Ingesting & Querying a Massive Codebase
# prompt containing 100k to 1,000,000+ tokens of repository code
inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to("cuda")
with torch.no_grad():
output = model.generate(
inputs["input_ids"],
max_new_tokens=512,
num_retrieved_chunks=2,
micro_window_size=1024,
do_sample=False,
tokenizer=tokenizer,
)
response = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)
📁 Repository Structure
├── config.json # Model hyperparameters & ISOM-R2 defaults
├── modeling_isom_qwen25_coder.py # Core ISOM-R2 engine with Aazhi architectural classes
├── configuration_isom_qwen25_coder.py # Configuration wrapper
├── isom_r2_module.py # Virtual SVD & Lie-algebra manifold operators
├── ISOM_R2.ipynb # Audited 1,055,402-token Colab benchmark execution proof
├── benchmarks/
│ ├── ISOM_R2_1M_Benchmark.ipynb # Verified Colab benchmark notebook
│ ├── audited_systems_benchmark_qwen25_coder.json
│ └── isom_r2_1m_real_repo_benchmark_results.json
└── model.safetensors # Model weights (Native FP16, ~3.09 GB)
📜 Citation & Attribution
If you use Aazhi-Coder-1.5B or the ISOM-R2 architecture in your research or applications, please cite:
@software{aazhi_coder_2026,
author = {Prannesh K. V. A.},
title = {Aazhi-Coder-1.5B: 1,048,576-Token Bounded-Memory Recurrent Code Intelligence},
year = {2026},
publisher = {Hugging Face},
doi = {10.5281/zenodo.14925828},
url = {https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B}
}
Acknowledgements
Aazhi-Coder-1.5B is built upon the Qwen2.5-Coder-1.5B-Instruct model developed by the Qwen team at Alibaba Cloud, released under the Apache 2.0 license.
- Downloads last month
- 373
Model tree for Prannesshkva/Aazhi-Coder-1.5B
Base model
Qwen/Qwen2.5-1.5B