- Ael-Coder-16B-MoE
- Overview
- Ael Model Family — Audited Multi-Million-Token Benchmarks
- Audited Benchmark 1: 2,774,027-Token Multi-File MoE Code Synthesis
- Audited Benchmark 2: 1,089,849-Token Single-Hop Repository Retrieval
- Architectural Specifications
- Quickstart Usage
- Citation, Dual-Layer Licensing & Upstream Attribution
Ael-Coder-16B-MoE
15.71B Total / 2.36B Active Mixture-of-Experts Code Intelligence • ISOM-R2 Bounded-Memory Architecture
Overview
Ael-Coder-16B-MoE is a 15.71-billion-parameter Mixture-of-Experts code intelligence model (64 routed experts, 2 shared experts, 2.36B active parameters per token) equipped with Multi-Head Latent Attention (MLA, $d_k = 192$) and powered by the ISOM-R2 (Isometric State Operator Manifold) multi-million-token architecture.
Designed for repository-scale code synthesis and cross-module dependency resolution, Ael-Coder-16B-MoE streams 2,774,027 real codebase tokens (1,355 chunks across 82 source files from huggingface/transformers) in 12.18 seconds (227,686 tokens/sec) on a single NVIDIA A100-40GB GPU. Throughout the entire 2.77M-token prefill, active KV cache length remains locked at 2,112 tokens with only +0.06 GB streaming VRAM overhead (29.28 GB model weights $\rightarrow$ 29.34 GB streaming VRAM, 30.51 GB peak VRAM during synthesis).
Core Architectural Capabilities
- Hierarchical Paged Virtual SVD Cache (
ISOMR2VirtualSVDCache)
Streams multi-million-token repositories in 2,048-token chunks across MLA decoder layers while strictly bounding active GPU KV memory to 2,112 tokens (64 attention sinks + 2,048 rolling window). - Entity-Balanced Multi-Hop Retrieval (
ISOMR2Engine)
Automatically intercepts long-context inputs insidemodel.generate(), balancing per-entity lexical and spectral recall to isolate exact target definitions separated by over 1.39 million tokens ([339, 1020], zero distractors). - $SO(192)$ Multi-Head Latent Attention Lie-Manifold Transport
Enforces skew-symmetric Cayley orthogonal transport ($R = (I - \frac{1}{2}A)^{-1}(I + \frac{1}{2}A) \in SO(192)$) across the 192-dimensional MLA head space, achieving an audited orthogonality error of $4.47 \times 10^{-5}$ and an $O(1)$ multi-hop state rollback error of $3.06 \times 10^{-6}$.
Ael Model Family — Audited Multi-Million-Token Benchmarks
All four models in the Ael benchmark suite are evaluated on unpadded, real-world open-source repositories (huggingface/transformers, sympy/sympy, and django/django) on NVIDIA A100-SXM4-40GB hardware. Interactive comparisons and telemetry logs are hosted on the Official Ael Benchmark Space.
| Model | Architecture & Parameters | Benchmark Domain & Corpus | Audited Context | Inter-Hop Distance | Total Time (Throughput) | Weights / Peak VRAM | Multi-Hop Recall & Verification |
|---|---|---|---|---|---|---|---|
| Ael-Coder-1.5B | AelCoder15BForCausalLM1.54B Dense (28L GQA) |
Multi-File PyTorch Synthesistransformers (82 files) |
2,147,447 (1,049 chunks) |
1,077,049 tokens (526 chunks) |
7.42 s (289,500 tok/s) |
2.98 GB / 3.14 GB (+0.16 GB overhead) |
100% ([388, 791, 265])LoggedGELU (0.00e+00 err) |
| Ael-Reasoning-1.5B-Instruct | AelReasoning15BForCausalLM1.54B Dense (28L + $SO(64)$) |
Symbolic Math & Combinatoricssympy (42 modules) |
2,311,513 (1,129 chunks) |
932,802 tokens (455 chunks) |
9.97 s (231,748 tok/s) |
2.88 GB / 3.03 GB (+0.15 GB overhead) |
100% ([501, 956])DerangedFibonacci (Exact) |
| Ael-Coder-16B-MoE | AelCoder16BMoEForCausalLM15.71B Total / 2.36B Active MoE |
Repository-Scale MoE Synthesistransformers (82 files) |
2,774,027 (1,355 chunks) |
1,394,629 tokens (681 chunks) |
12.18 s (227,686 tok/s) |
29.28 GB / 30.51 GB (+1.23 GB overhead) |
100% ([339, 1020])LoggedGELU (0.00e+00 err) |
| Ael-Pro-40B | AelPro40BForCausalLM40.0B Dense (60L, 4-bit NF4) |
Enterprise Security & Crypto Auditdjango (122 modules) |
2,399,330 (1,172 chunks) |
1,171,652 tokens (572 chunks) |
13.94 s (172,072 tok/s) |
21.62 GB / 24.73 GB (+3.11 GB overhead) |
100% ([304, 876])SignedPBKDF2Hasher (Verified) |
Audited Benchmark 1: 2,774,027-Token Multi-File MoE Code Synthesis
Executed Notebook: Ael_Coder_16B.ipynb
Task Methodology
Evaluated on 82 real Python source files (9,779,186 characters, 2,774,027 real tokens across 1,355 chunks) from huggingface/transformers on an NVIDIA A100-SXM4-40GB GPU. The model must retrieve two target classes separated by 1,394,629 real tokens (681 chunks apart):
- Hop 1 (
Chunk 339, Token#695,058):class CaptureStdout(CaptureStd)intesting_utils.py - Hop 2 (
Chunk 1020, Token#2,089,687):class GELUActivation(nn.Module)inactivations.py
Using both retrieved definitions, the model synthesizes a LoggedGELU(GELUActivation) class that captures standard output during the forward pass and is verified live on GPU against PyTorch's native GELUActivation.
Hardware Execution Log (NVIDIA A100-SXM4-40GB)
Loaded AelCoder16BMoEForCausalLM (15.71B total / 2.36B active MoE params) | Weights VRAM: 29.28 GB
Tokenizing 2.25M+ real-world multi-file corpus (82 files | 9,779,186 chars)...
Hop 1 Target : class CaptureStdout(CaptureStd) @ Token #695,058 -> Chunk 339
Hop 2 Target : class GELUActivation(nn.Module) @ Token #2,089,687 -> Chunk 1020
Inter-File Gap : 1,394,629 real tokens (681 chunks apart)
[ISOM-R2 Engine] Streaming 2,773,953 codebase context tokens across 1355 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 552,960 / 2,773,953 tokens (19.9%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Prefill 1,105,920 / 2,773,953 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Prefill 1,658,880 / 2,773,953 tokens (59.8%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Prefill 2,211,840 / 2,773,953 tokens (79.7%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Prefill 2,764,800 / 2,773,953 tokens (99.7%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Prefill 2,773,953 / 2,773,953 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 29.34 GB
[ISOM-R2] Retrieved salient context pages: [339, 1020] | Active KV: 2112 tokens
============================================================================================
2.77M-TOKEN MULTI-FILE MOE SYNTHESIS BENCHMARK: Prannesshkva/Ael-Coder-16B-MoE
============================================================================================
Total Real Context Tokens : 2,774,027 (82 files | 1355 chunks)
Inter-File Hop Distance : 1,394,629 tokens (Chunk 339 <-> Chunk 1020)
Total Time (Prefill+Gen) : 12.18 s (227,686 tok/s)
Model Weights VRAM : 29.28 GB
Peak Total GPU VRAM : 30.51 GB (KV + Activation Overhead: +1.23 GB)
Active KV Cache Length : 2112 tokens
Retrieved Chunk Indices : [339, 1020]
Multi-Hop Recall : Hop 1 Hit: True | Hop 2 Hit: True
SO(192) MLA Orthogonality : 4.47e-05 ||R^T R - I||_F | O(1) Rollback Error: 3.06e-06
Live Execution Verification: PASSED (Output shape: [1, 4], Captured Stdout: 'torch.Size([1, 4])', max_err = 0.00e+00)
============================================================================================
Audited Benchmark 2: 1,089,849-Token Single-Hop Repository Retrieval
Evaluated on 93 real Python source files (3,867,627 characters, 1,089,849 tokens across 533 chunks) from huggingface/transformers and pytorch/pytorch:
Loaded AelCoder16BMoEForCausalLM | Model VRAM: 31.49 GB
Corpus: 93 real files | 3,867,627 chars | 1,089,849 tokens | Ground-Truth: Chunk 257
[ISOM-R2 Engine] Streaming 1,089,795 codebase context tokens across 533 chunks (active GPU buffer < 400 MB)...
[ISOM-R2] Prefill 1,089,795 / 1,089,795 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 31.49 GB
[ISOM-R2] Retrieved salient context pages: [258, 257] | Active KV: 2112 tokens
================================================================================
BENCHMARK RESULTS: Prannesshkva/Ael-Coder-16B-MoE
================================================================================
Total Real Context Tokens : 1,089,849
Total Time (Prefill+Gen) : 15.86 s (68,717 tok/s)
Model Weights VRAM : 31.49 GB
Peak Total GPU VRAM : 32.06 GB ( Overhead: +0.57 GB )
Active KV Cache Length : 2112 tokens
Retrieved Chunk Indices : [258, 257] (Exact Token Ground-Truth: Chunk 257)
Ground-Truth Base Class : CaptureStd
Model Generated Output : CaptureStd
================================================================================
Architectural Specifications
| Parameter | Specification |
|---|---|
| Model Architecture Class | AelCoder16BMoEForCausalLM (AelCoder16BMoEConfig) |
| Total / Active Parameters | 15.71 Billion Total / 2.36 Billion Active per Token |
| MoE & Attention Topology | 27 Layers, 64 Routed Experts, 2 Shared Experts, Multi-Head Latent Attention ($d_k = 192$) |
| Long-Context Engine | ISOM-R2 Paged Virtual SVD Cache (ISOMR2VirtualSVDCache) + $SO(192)$ Cayley Transport |
| Audited Context Length | 2,774,027 real tokens (1,355 chunks of 2,048 tokens) |
| Active GPU KV Cache | 2,112 tokens (64 attention sinks + 2,048 active sliding window) |
| VRAM Footprint (BF16) | 29.28 GB Weights • 29.34 GB Streaming (+0.06 GB) • 30.51 GB Peak Synthesis |
Quickstart Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "Prannesshkva/Ael-Coder-16B-MoE"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
).eval()
prompt = """<|im_start|>user
Write a high-performance concurrent queue in Python using lock-free atomic operations.<|im_end|>
<|im_start|>assistant
"""
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
outputs = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
tokenizer=tokenizer,
)
print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))
Citation, Dual-Layer Licensing & Upstream Attribution
@software{ael_coder_16b_moe_2026,
author = {Prannessh K. V. A.},
title = {Ael-Coder-16B-MoE: Multi-Million-Token Repository-Scale Bounded-Memory MoE Powered by ISOM-R2},
year = {2026},
publisher = {CERN Zenodo},
doi = {10.5281/zenodo.22649142},
url = {https://huggingface.co/Prannesshkva/Ael-Coder-16B-MoE}
}
Dual-Layer License Structure (DeepSeek License Section 3 & MIT Compliance)
Pursuant to Section 3 of the DeepSeek License Agreement and the MIT License, this repository separates licensing between the unmodified upstream pretrained foundation weights and the author's original architectural modifications:
| Component | Copyright Holder | Applicable License |
|---|---|---|
| ISOM-R2 Execution Engine & Architectural Modifications ( isom_r2_engine.py, ISOMR2VirtualSVDCache, $SO(192)$ MLA Cayley Lie-Group Transport, and custom AelCoder16BMoE* classes in modeling_isom_deepseek_coder_v2.py) |
Copyright © 2026 Prannessh K. V. A. | CC BY-NC-ND 4.0 (Non-Commercial Research) / Commercial Enterprise License via Author (LICENSE) |
| Base Pretrained Neural Weights & Unmodified Base Code (Initialized from deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct) |
Copyright © 2024 DeepSeek | DeepSeek License Agreement (Model Weights, including Attachment A Use-Based Restrictions) & MIT License (Base Code) (LICENSE & NOTICE) |
Mandatory Upstream Notice & Trademark Disclaimer (DeepSeek License Sections 3 & 4)
"DeepSeek-Coder-V2 is licensed under the MODEL LICENSE. Copyright (c) 2024 DeepSeek. All Rights Reserved."
- Prominent Notice of Modification (Section 3(b)): Modified by Prannessh K. V. A. to integrate the ISOM-R2 (Isometric State Operator Manifold) Paged Virtual SVD Cache (
ISOMR2VirtualSVDCache), entity-balanced multi-hop codebase retrieval (ISOMR2Engine), $SO(192)$ Multi-Head Latent Attention Lie-group transport, andAelCoder16BMoEForCausalLMexecution bindings. Full modification logs and Attachment A Use-Based Restrictions are included inLICENSEandNOTICE. - Distinct Naming & Non-Endorsement (Section 4 — Trademarks): In compliance with Section 4 of the DeepSeek License Agreement, this derivative architecture is published under the distinct Ael name (
Ael-Coder-16B-MoE) so as not to imply endorsement by or affiliation with the original licensor. DeepSeek is a trademark of DeepSeek AI. This independent research work is not affiliated with, sponsored by, or endorsed by DeepSeek AI.
- Downloads last month
- 2,177
Model tree for Prannesshkva/Ael-Coder-16B-MoE
Base model
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct