Ael-Pro-40B

40.0B Dense Flagship Foundation Model โ€ข ISOM-R2 Bounded-Memory Architecture

DOI Benchmark Suite Author


Overview

Ael-Pro-40B is the 40-billion-parameter dense flagship foundation model in the Ael model family, engineered for repository-scale security auditing, cryptographic compliance verification, and multi-module systems reasoning across multi-million-token contexts. Built on the 60-layer Falcon-40B architecture (128 query heads, 8 key-value heads per layer) and integrated with the ISOM-R2 (Isometric State Operator Manifold) execution engine, the model streams and synthesizes across 2,399,330 real tokens on a single NVIDIA A100-40GB GPU.

In a conventional Transformer, storing a 2.40M-token key-value cache across 60 layers requires 137.3 GiB of VRAM for the KV cache alone. Ael-Pro-40B bounds the active GPU attention buffer strictly to 2,112 tokens (+0.09 GB streaming KV overhead above the 21.62 GB 4-bit NF4 weights), completing a 2.40M-token prefill and multi-hop synthesis pass in 13.94 seconds (172,072 tokens/sec).

Core Architectural Capabilities

  1. Hierarchical Paged Virtual SVD Cache (ISOMR2VirtualSVDCache)
    Streams arbitrary-length context sequences in 2,048-token chunks while maintaining a constant 2,112-token active GPU KV buffer (64 attention sink tokens + 2,048-token rolling window).
  2. Entity-Balanced Multi-Hop Retrieval (ISOMR2Engine)
    Automatically intercepts long-context inputs inside model.generate(), indexing every chunk via low-rank spectral signatures and entity-balanced lexical scoring to retrieve exact distant modules across million-token gaps with zero distractor pages.
  3. $SO(64)$ Lie-Manifold Orthogonal Transport
    Preserves state norm and trajectory invertibility across deep 60-layer recurrent state transitions via closed-form Cayley transformations ($U = (I - \frac{1}{2}A)^{-1}(I + \frac{1}{2}A) \in SO(64)$), achieving an audited orthogonality error of $3.98 \times 10^{-6}$ and $O(1)$ multi-hop rollback error of $5.18 \times 10^{-7}$.

Ael Model Family โ€” Audited Multi-Million-Token Benchmarks

All four models in the Ael benchmark suite are evaluated on unpadded, real-world open-source repositories (huggingface/transformers, sympy/sympy, and django/django) on NVIDIA A100-SXM4-40GB hardware. Interactive comparisons and telemetry logs are hosted on the Official Ael Benchmark Space.

Model Architecture & Parameters Benchmark Domain & Corpus Audited Context Inter-Hop Distance Total Time (Throughput) Weights / Peak VRAM Multi-Hop Recall & Verification
Ael-Coder-1.5B AelCoder15BForCausalLM
1.54B Dense (28L GQA)
Multi-File PyTorch Synthesis
transformers (82 files)
2,147,447
(1,049 chunks)
1,077,049 tokens
(526 chunks)
7.42 s
(289,500 tok/s)
2.98 GB / 3.14 GB
(+0.16 GB overhead)
100% ([388, 791, 265])
LoggedGELU (0.00e+00 err)
Ael-Reasoning-1.5B-Instruct AelReasoning15BForCausalLM
1.54B Dense (28L + $SO(64)$)
Symbolic Math & Combinatorics
sympy (42 modules)
2,311,513
(1,129 chunks)
932,802 tokens
(455 chunks)
9.97 s
(231,748 tok/s)
2.88 GB / 3.03 GB
(+0.15 GB overhead)
100% ([501, 956])
DerangedFibonacci (Exact)
Ael-Coder-16B-MoE AelCoder16BMoEForCausalLM
15.71B Total / 2.36B Active MoE
Repository-Scale MoE Synthesis
transformers (82 files)
2,774,027
(1,355 chunks)
1,394,629 tokens
(681 chunks)
12.18 s
(227,686 tok/s)
29.28 GB / 30.51 GB
(+1.23 GB overhead)
100% ([339, 1020])
LoggedGELU (0.00e+00 err)
Ael-Pro-40B AelPro40BForCausalLM
40.0B Dense (60L, 4-bit NF4)
Enterprise Security & Crypto Audit
django (122 modules)
2,399,330
(1,172 chunks)
1,171,652 tokens
(572 chunks)
13.94 s
(172,072 tok/s)
21.62 GB / 24.73 GB
(+3.11 GB overhead)
100% ([304, 876])
SignedPBKDF2Hasher (Verified)

Audited Benchmark 1: 2,399,330-Token Enterprise Security & Cryptographic Audit

Task Methodology

To evaluate Ael-Pro-40B on large-scale enterprise security, cryptographic credential management, and cross-module compliance auditing, the model streams 122 real Python modules (8,973,788 characters, 2,399,330 real tokens across 1,172 chunks) from the django/django core framework (django/core/signing, django/contrib/auth, django/db/models, django/middleware, and django/dispatch).

The model is tasked with locating two distinct security primitives separated by 1,171,652 real tokens (572 chunks apart):

  • Hop 1 (Chunk 304, Token #623,966): class Signer in django/core/signing.py (salted HMAC-SHA256 cryptographic signing and BadSignature tamper detection).
  • Hop 2 (Chunk 876, Token #1,795,618): class PBKDF2PasswordHasher(BasePasswordHasher) in django/contrib/auth/hashers.py (1,800,000-iteration PBKDF2-SHA256 password derivation and constant-time verification).

From these retrieved definitions, the model must synthesize a compliant SignedPBKDF2Hasher(PBKDF2PasswordHasher) audit class that encodes credentials, signs the resulting hash using Signer(key=key), and verifies round-trip integrity while rejecting single-byte tampering.

Memory Scaling: Standard 40B MQA vs. Ael-Pro-40B (ISOM-R2)

Metric (at 2,399,330 Tokens) Standard 40B Transformer (60L, 8 KV Heads) Ael-Pro-40B (ISOM-R2) Measured Improvement
KV Cache Memory Footprint 137.29 GiB (Linear $O(N)$) +0.09 GB Streaming Overhead 99.9% KV Memory Reduction
Active Attention Context Length 2,399,330 tokens (CUDA OOM) 2,112 tokens (64 sinks + 2,048 window) Strict $O(1)$ Working State
Total GPU VRAM (4-bit NF4 Weights + KV) $> 158.9\text{ GB}$ (Requires $2\times$ A100-80GB) 21.71 GB Streaming / 24.73 GB Peak Runs on $1\times$ A100-40GB
End-to-End Execution Time Infeasible ($O(N^2)$ attention wall) 13.94 seconds (172,072 tok/s) Real-Time Repository Audit

Hardware Execution Log (NVIDIA A100-SXM4-40GB)

Loaded AelPro40BForCausalLM (40.0B Dense | 60 Layers | 128 Q / 8 KV Heads) | Weights VRAM: 21.62 GB
Tokenizing 2.3M+ Enterprise Security & ORM corpus (122 files | 8,973,788 chars)...
Hop 1 Target   : class Signer (HMAC-SHA256 Signing)         @ Token #623,966 -> Chunk 304
Hop 2 Target   : class PBKDF2PasswordHasher(BasePasswordHasher) @ Token #1,795,618 -> Chunk 876
Inter-Hop Gap  : 1,171,652 real tokens (572 chunks apart)

  [ISOM-R2 Engine] Streaming 2,399,185 codebase context tokens across 1172 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 479,232 / 2,399,185 tokens (20.0%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Prefill 958,464 / 2,399,185 tokens (39.9%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Prefill 1,437,696 / 2,399,185 tokens (59.9%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Prefill 1,916,928 / 2,399,185 tokens (79.9%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Prefill 2,396,160 / 2,399,185 tokens (99.9%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Prefill 2,399,185 / 2,399,185 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Retrieved salient context pages: [304, 876] | Active KV: 2112 tokens

============================================================================================
2.3M+ TOKEN ENTERPRISE SECURITY & AUDIT BENCHMARK: Prannesshkva/Ael-Pro-40B
============================================================================================
Model Architecture Class  : AelPro40BForCausalLM (40.0B Dense | 4-Bit NF4 | 60 Layers)
Enterprise Corpus         : 122 real Django modules (8,973,788 chars)
Total Real Context Tokens : 2,399,330
Inter-Module Hop Distance : 1,171,652 tokens (572 chunks apart)
Total Time (Prefill+Gen)  : 13.94 s (172,072 tok/s)
Model Weights VRAM        : 21.62 GB
Peak Total GPU VRAM       : 24.73 GB ( Overhead: +3.11 GB )
Active KV Cache Length    : 2112 tokens
Ground-Truth Chunks       : Hop 1 = Chunk 304 | Hop 2 = Chunk 876
Retrieved Chunk Indices   : [304, 876] (Hop 1 Hit: True | Hop 2 Hit: True)
SO(64) Lie Orthogonality  : ||A^T A - I||_F = 3.98e-06 | O(1) Rollback Err = 5.18e-07
--------------------------------------------------------------------------------------------
MODEL SYNTHESIZED ENTERPRISE CRYPTOGRAPHIC AUDIT CLASS:
--------------------------------------------------------------------------------------------
class SignedPBKDF2Hasher(PBKDF2PasswordHasher):
    @staticmethod
    def issue_audit_token(password, salt, key):
        encoded = PBKDF2PasswordHasher().encode(password, salt)
        signed_token = Signer(key=key).sign(encoded)
        verified = PBKDF2PasswordHasher().verify(password, Signer(key=key).unsign(signed_token))
        return (signed_token, verified)
--------------------------------------------------------------------------------------------
LIVE ENTERPRISE SECURITY & COMPLIANCE VERIFICATION:
  โ€ข Inheritance Check         : issubclass(SignedPBKDF2Hasher, PBKDF2PasswordHasher) = True
  โ€ข Signed PBKDF2 Audit Token : pbkdf2_sha256$1800000$ael_salt_99$l4kux5SXVPIiDs...vnWtZFJzaZm19ZgFflq6inYo
  โ€ข Round-Trip HMAC + PBKDF2  : verified=True (PASSED)
  โ€ข 1-Byte Tamper Rejection   : BadSignature raised=True (PASSED)
============================================================================================

Audited Benchmark 2: 1,056,780-Token Single-Hop Repository Retrieval

Evaluated on 93 real Python source files (3,867,627 characters, 1,056,780 tokens across 516 chunks) from huggingface/transformers and pytorch/pytorch:

Loaded AelPro40BForCausalLM | Model VRAM: 21.71 GB
Corpus: 93 real files | 3,867,627 chars | 1,056,780 tokens | Ground-Truth: Chunk 252
  [ISOM-R2 Engine] Streaming 1,056,738 codebase context tokens across 516 chunks (active GPU buffer < 400 MB)...
  [ISOM-R2] Prefill 1,056,738 / 1,056,738 tokens (100.0%) | Active buffer: 2112 tokens | VRAM: 21.71 GB
  [ISOM-R2] Retrieved salient context pages: [253, 252] | Active KV: 2112 tokens

================================================================================
BENCHMARK RESULTS: Prannesshkva/Ael-Pro-40B
================================================================================
Total Real Context Tokens : 1,056,780
Total Time (Prefill+Gen)  : 15.39 s (68,670 tok/s)
Model Weights VRAM        : 21.71 GB
Peak Total GPU VRAM       : 22.84 GB ( Overhead: +1.13 GB )
Active KV Cache Length    : 2112 tokens
Retrieved Chunk Indices   : [253, 252] (Exact Token Ground-Truth: Chunk 252)
Ground-Truth Base Class   : CaptureStd
Model Generated Output    : CaptureStd)
================================================================================

Architectural Specifications

Parameter Specification
Model Architecture Class AelPro40BForCausalLM (AelPro40BConfig)
Total Parameters 40.0 Billion Dense
Decoder Topology 60 Layers, $d_{\text{model}} = 8192$, 128 Query Heads, 8 KV Heads (Multi-Query / Grouped-Query)
Long-Context Engine ISOM-R2 Paged Virtual SVD Cache (ISOMR2VirtualSVDCache) + $SO(64)$ Cayley Transport
Audited Context Length 2,399,330 real tokens (1,172 chunks of 2,048 tokens)
Active GPU KV Cache 2,112 tokens (64 attention sinks + 2,048 active sliding window)
VRAM Footprint (4-bit NF4) 21.62 GB Weights โ€ข 21.71 GB Streaming (+0.09 GB) โ€ข 24.73 GB Peak Synthesis

Quickstart Usage

Ael-Pro-40B integrates directly with Hugging Face transformers using trust_remote_code=True. Multi-million-token contexts are automatically streamed and indexed by the built-in ISOM-R2 engine inside model.generate().

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig

model_id = "Prannesshkva/Ael-Pro-40B"

tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_use_double_quant=True,
    bnb_4bit_compute_dtype=torch.bfloat16,
)

model = AutoModelForCausalLM.from_pretrained(
    model_id,
    quantization_config=bnb_config,
    device_map="auto",
    trust_remote_code=True,
).eval()

prompt = "Analyze the cryptographic signing and password hashing compliance across this enterprise codebase."
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False,
        tokenizer=tokenizer,
    )

print(tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

Citation, Dual-Layer Licensing & Upstream Attribution

@article{prannessh2026ael_pro_40b,
  title   = {Ael-Pro-40B: Multi-Million-Token Bounded-Memory Flagship Foundation Model Powered by ISOM-R2},
  author  = {Prannessh K. V. A.},
  journal = {CERN Zenodo},
  year    = {2026},
  doi     = {10.5281/zenodo.22649142},
  url     = {https://huggingface.co/Prannesshkva/Ael-Pro-40B}
}

Dual-Layer License Structure (Apache 2.0 Section 4 Compliance)

Pursuant to Section 4 of the Apache License, Version 2.0, this repository separates licensing between the unmodified upstream pretrained foundation weights and the author's original architectural modifications:

Component Copyright Holder Applicable License
ISOM-R2 Execution Engine & Architectural Modifications
(isom_r2_engine.py, ISOMR2VirtualSVDCache, $SO(64)$ Cayley Lie-Manifold Transport, and custom AelPro40B* classes in modeling_falcon.py & configuration_falcon.py)
Copyright ยฉ 2026 Prannessh K. V. A. CC BY-NC-ND 4.0 (Non-Commercial Research) / BSL 1.1 / Commercial Enterprise License via Author (LICENSE)
Base Pretrained Neural Weights & Unmodified Base Falcon Code
(Initialized from tiiuae/falcon-40b)
Copyright ยฉ 2023 Technology Innovation Institute (TII) Apache License, Version 2.0 (LICENSE & NOTICE)

Statement of Modifications & Trademark Notice (Apache 2.0 Sections 4 & 6)

  1. Prominent Notice of Modification (Section 4(b)): Modified by Prannessh K. V. A. to integrate the ISOM-R2 (Isometric State Operator Manifold) Paged Virtual SVD Cache (ISOMR2VirtualSVDCache), entity-balanced multi-hop codebase retrieval (ISOMR2Engine), $SO(64)$ Cayley Lie-manifold orthogonal transport, dynamic symmetric INT8 KV quantization, and AelPro40BForCausalLM execution bindings. Full modification logs are documented in NOTICE.
  2. Distinct Naming & Non-Endorsement (Section 6 โ€” Trademarks): In compliance with Section 6 of the Apache License 2.0, this derivative architecture is published under the distinct Ael name (Ael-Pro-40B) so as not to imply endorsement by or affiliation with the original licensor. Falcon and TII are trademarks of the Technology Innovation Institute. This independent research work is not affiliated with, sponsored by, or endorsed by the Technology Innovation Institute (TII).
Downloads last month
3,352
Safetensors
Model size
42B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Prannesshkva/Ael-Pro-40B

Finetuned
(5)
this model

Spaces using Prannesshkva/Ael-Pro-40B 2

Collection including Prannesshkva/Ael-Pro-40B