🌊 Aazhi-Coder-1.5B (ஆழி): 1,048,576-Token Bounded-Memory Code Intelligence

1M Codebase Streaming • 3.24 GB Peak VRAM • 100% Retrieval Precision • Powered by ISOM-R2

DOI LinkedIn Aazhi Context Peak VRAM Ingestion Speed Audited Run


Executive Summary

Aazhi-Coder-1.5B (named after Aazhi [ஆழி] — representing the boundless ocean) is a production-grade 1,048,576-token (1 Million) long-context coding model built on Alibaba's exceptional Qwen2.5-Coder-1.5B-Instruct foundation and powered by the ISOM-R2 (Isometric State Space / Virtual SVD) memory architecture.

Standard Transformer models suffer catastrophic memory bottlenecks at million-token scales. For a 1.05M-token sequence, a conventional FP16 Key-Value cache demands 28.2 GiB of VRAM before allocating weights, instantly crashing single GPUs.

Aazhi-Coder-1.5B completely eliminates the 28.2 GiB KV cache explosion:

  • 3.24 GB Peak VRAM: Streams and processes 1,055,402 continuous tokens within a strictly bounded memory pool ($< 3.5\text{ GB}$).
  • 123.78s Ingestion Time: Reaches an effective throughput of ~8,526 tokens/second, digesting 181 production repository files in approximately 2 minutes.
  • 100% Macro-Retrieval Precision: Correctly isolates target files across 516 candidate chunks with zero distractor noise.
  • Reproducible Execution Proof: Verifiable directly via the audited benchmark notebook ISOM_R2.ipynb.

⚡ The Memory Bottleneck: Conventional Attention vs. ISOM-R2

In standard Grouped-Query Attention (GQA) architectures (28 layers, 2 KV heads, head dimension 128), KV cache memory scales linearly:

KV Cache Size=2×L×HKV×Dhead×Ntokens×2 bytes\text{KV Cache Size} = 2 \times L \times H_{KV} \times D_{\text{head}} \times N_{\text{tokens}} \times 2 \text{ bytes}

At 1,055,402 tokens: 2×28×2×128×1,055,402×2≈28.18 GiB\text{At } 1,055,402 \text{ tokens: } 2 \times 28 \times 2 \times 128 \times 1,055,402 \times 2 \approx \mathbf{28.18\text{ GiB}}

Parameter Standard Transformer GQA Aazhi-Coder-1.5B (ISOM-R2) Advantage
KV Cache Footprint (1.05M tokens) 28.18 GiB (Linear $O(N)$) 4,160 Active Tokens (~0.11 GiB) 99.6% Reduction
Prefill VRAM Behavior Explodes to OOM Flat 2.97 GB Invariant Bounded State Space
Peak Execution VRAM $> 33.5\text{ GB}$ (A100 required) 3.24 GB Peak 8.7× Total Compression
Target Hardware 40GB / 80GB Data Center GPUs Consumer 4GB / 6GB / 8GB GPUs On-premise / Laptop deployment
Throughput (1.05M tokens) Quadratic deceleration 123.78 seconds (~8,526 tok/s) Constant-time streaming

📊 Audited Production Benchmark: 1,055,402 Real Tokens

Benchmark executed on real Python source files cloned from the official Hugging Face transformers repository (zero synthetic tokens, zero artificial padding):

  • Corpus Scope: 181 production source files (covering models/, pipelines/, generation/, trainer/)
  • Prompt Assembly: 1,055,402 total sequence tokens across 516 discrete 2,048-token chunks
  • Target Multi-Hop Dependency: Cross-file inquiry spanning configuration_llama.py (Chunk 218) and modeling_llama.py (Chunk 217)
[ISOM-R2 Engine] Streaming 1,055,192 codebase context tokens across 516 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill  210,944 / 1,055,192 tokens (20.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  421,888 / 1,055,192 tokens (40.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  632,832 / 1,055,192 tokens (60.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill  843,776 / 1,055,192 tokens (80.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2] Prefill 1,054,720 / 1,055,192 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 2.97 GB
[ISOM-R2 Engine] Micro-window sliced: Chunk 218, anchor=1030, range=[518:1542] (1024 tokens)
[ISOM-R2 Engine] Micro-window sliced: Chunk 217, anchor=1010, range=[498:1522] (1024 tokens)
[ISOM-R2] Retrieved salient context pages: [218, 217] | Active KV: 4160 tokens

============================================================
1,000,000-TOKEN MULTI-HOP GENERATION RESULTS:
============================================================
Input Sequence Length : 1,055,402 tokens
Total Time Taken      : 123.78s
Peak GPU VRAM         : 3.24 GB
============================================================

Retrieval & Memory Provenance

  • Macro-Retrieval Precision: 100%. Out of 516 candidate chunks across 181 files, only Chunk 217 and Chunk 218 were paged. Zero distractor chunks were retrieved.
  • Micro-Window Anchoring: Dynamic anchor-centered slicing restricted active attention to 1,024 tokens per salient chunk, pinning active KV cache memory to 4,160 tokens.

🏗️ Architecture: ISOM-R2 Under the Hood

1,048,576 Token Stream (Codebase)
       │
       ├──► [ 1. Streaming Bounded State Ingestion ] ──────► 2,048-token chunk prefill @ flat 2.97 GB VRAM
       │
       ├──► [ 2. Sub-Harmonic Lie Frequency Calibration ] ─► Zero phase aliasing across 1M tokens (ω_min = 5.99e-6)
       │
       ├──► [ 3. Hybrid Salient Context Retrieval ] ──────► Filters 516 chunks down to top-k relevant files
       │
       ├──► [ 4. Anchor-Centered Micro-Window Slicing ] ──► Paged attention restricted to 1,024 tokens/chunk
       │
       └──► [ 5. Focused Multi-Hop Decoding ] ────────────► High-fidelity generation with < 3.24 GB peak VRAM

1. Bounded State Space Ingestion

Rather than storing every intermediate key and value tensor in video memory, ISOM-R2 processes input sequences through bounded state projections. Token activations are compressed into isometric manifold representations, keeping GPU VRAM strictly flat at 2.97 GB throughout prefill.

2. Sub-Harmonic Lie Calibration ($\omega_{\min}$)

Standard rotational embeddings experience severe phase wrap and aliasing past 32k tokens. ISOM-R2 establishes a calibrated sub-harmonic frequency floor:

ωmin⁡<2π1,048,576≈5.9921×10−6 rad/token\omega_{\min} < \frac{2\pi}{1,048,576} \approx 5.9921 \times 10^{-6} \text{ rad/token}

This guarantees that the slowest coordinate manifold rotates strictly less than one complete revolution across the entire 1,048,576-token sequence, preserving stable temporal geometry.

3. Salient Micro-Window Paging

During generation, ISOM-R2 performs two-tiered context activation:

  1. Macro Level: Filters the repository down to the most relevant code modules.
  2. Micro Level: Automatically isolates anchor-centered micro-windows (default 1,024 tokens) containing critical function implementations, enabling fast generation without attending over full megatoken buffers.

🚀 Quickstart: Drop-In Usage

Aazhi-Coder-1.5B integrates natively with Hugging Face transformers using trust_remote_code=True.

Installation

pip install -q transformers>=4.49.0 accelerate torch

Loading the Model

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "Prannesshkva/Aazhi-Coder-1.5B"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    trust_remote_code=True,
    torch_dtype=torch.float16,
    device_map="cuda",
)
model.eval()

# Configure bounded retrieval policy
model.config.isom_r2_num_retrieved_chunks = 2
model.config.isom_r2_micro_window_size = 1024

print(f"Model loaded on: {next(model.parameters()).device}")
print(f"Base VRAM: {torch.cuda.memory_allocated(0)/(1024**3):.2f} GB")

Ingesting & Querying a Massive Codebase

# prompt containing 100k to 1,000,000+ tokens of repository code
inputs = tokenizer(huge_codebase_prompt, return_tensors="pt").to("cuda")

with torch.no_grad():
    output = model.generate(
        inputs["input_ids"],
        max_new_tokens=512,
        num_retrieved_chunks=2,
        micro_window_size=1024,
        do_sample=False,
        tokenizer=tokenizer,
    )

response = tokenizer.decode(output[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
print(response)

📁 Repository Structure

├── config.json                             # Model hyperparameters & ISOM-R2 defaults
├── modeling_isom_qwen25_coder.py           # Core ISOM-R2 engine with Aazhi architectural classes
├── configuration_isom_qwen25_coder.py      # Configuration wrapper
├── isom_r2_module.py                       # Virtual SVD & Lie-algebra manifold operators
├── ISOM_R2.ipynb                           # Audited 1,055,402-token Colab benchmark execution proof
├── benchmarks/
│   ├── ISOM_R2_1M_Benchmark.ipynb          # Verified Colab benchmark notebook
│   ├── audited_systems_benchmark_qwen25_coder.json
│   └── isom_r2_1m_real_repo_benchmark_results.json
└── model.safetensors                       # Model weights (Native FP16, ~3.09 GB)

📜 Citation & Attribution

If you use Aazhi-Coder-1.5B or the ISOM-R2 architecture in your research or applications, please cite:

@software{aazhi_coder_2026,
  author       = {Prannesh K. V. A.},
  title        = {Aazhi-Coder-1.5B: 1,048,576-Token Bounded-Memory Recurrent Code Intelligence},
  year         = {2026},
  publisher    = {Hugging Face},
  doi          = {10.5281/zenodo.14925828},
  url          = {https://huggingface.co/Prannesshkva/Aazhi-Coder-1.5B}
}

Acknowledgements

Aazhi-Coder-1.5B is built upon the Qwen2.5-Coder-1.5B-Instruct model developed by the Qwen team at Alibaba Cloud, released under the Apache 2.0 license.

Downloads last month
373
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Prannesshkva/Aazhi-Coder-1.5B

Finetuned
(214)
this model