algenta-llm

LLM, attention, RoPE, embedding and KV-cache kernels. Compiled Mojo, loaded in-process.

52 modules · 357 functions · CPU · Cuda · Metal

Get started

pip install kernels torch
from kernels import get_kernel

kernel = get_kernel(
    "thyn-ai/algenta-llm",
    version=1,
    trust_remote_code=["thyn-ai/algenta-llm"],
)
kernel.embeddings.cosine_similarity([1.0, 0.0], [1.0, 1.0])  # -> 0.7071

On macOS 26 the loader picks the Metal build. On macOS 15, add backend="cpu". On a Linux host with an NVIDIA GPU and a CUDA build of PyTorch the loader picks the CUDA build. trust_remote_code names the repositories you allow; Hugging Face's trusted publishers load without it.

Plain Python in, plain Python out. Lists, tuples, buffers and tensors are accepted wherever the contract expects a list; structured results are dictionaries. Some functions take a list and the number of elements to use from it, which may not exceed the list's length; multi-dimensional data is passed flattened, row-major, with its dimensions. Functions that update an argument do so in place, as help() says. Every call is checked against the published contract before it reaches native code. An invalid call raises KernelError with a stable code, never a crash. Engine kernels report shape and finiteness problems as a status; the wrapper raises KernelError named after it.

Any function can also be called by name, with args as a list or a dict of parameter names:

kernel.execute("accel.contract", "accel_backend_name", [...])

Resident decode

The Metal and CUDA builds run greedy Qwen3-family decode on the GPU. One session is open at a time; each step blocks until its token is computed and returns the token id.

with kernel.DecodeSession("/path/to/converted-checkpoint", prompt_token_ids, max_new_tokens=32) as session:
    tokens = list(session)

prompt_token_ids come from the checkpoint's own tokenizer. The checkpoint directory is the converted form (manifest.json + weights.blob) written by prepare_model_checkpoint() in the algenta package. kernel.decode(...) runs a whole session in one call; session.step() returns the full record for one token. Decode needs a GPU build; on the CPU build DecodeSession raises KernelError("gpu_build_required").

The GPU builds stage decode for models with max(d_model, h_q*d_head, d_ff) <= 6144, for example Qwen3-0.6B and Qwen3-1.7B; a larger model is refused with KernelError("resident_dimension_too_large").

What's inside

Module Functions What it does
accel.contract 6 Status codes and backend selection for accelerator kernels
accel.device 3 Accelerator detection for this host
accel.gemm_metal 3 Threadgroup-tiled GEMM on the GPU
attnk.backward 2 Gradients of scaled dot-product attention
attnk.contract 10 Status codes and work budget for attention kernels
attnk.flash 3 Tiled attention with online softmax
attnk.mask 7 Attention masks as additive bias: causal, sliding-window, padding, packed
attnk.rope 9 Rotary position embedding: frequency tables and application
attnk.sdpa 3 Scaled dot-product attention, materialized reference
attnk.tiled 5 Blockwise online-softmax attention over flat buffers, with grouped-query attention
bpe_tokenizer 10 BPE and SentencePiece tokenization primitives
causal_inference.propensity 5 Propensity scores, overlap diagnostics, inverse-probability weighting, matching
embeddings 20 Vector embeddings: similarity, distance, centroids, nearest neighbours
flash_attention 10 Blockwise memory-efficient attention math
gemm.blocked 4 Cache-blocked SIMD GEMM and GEMV over flat row-major buffers
gemm.contract 5 Status codes, dimension check and work budget for GEMM
gemm.quantized 4 Matmul with dequantization fused into the k-loop
kv_cache 10 KV-cache memory math and incremental decoding
llm_sampling 11 Token sampling: temperature, top-k, top-p, min-p, penalties, beam scores
llmk.contract 5 Status codes and shape check for the LLM kernels
llmk.gemv_fast 1 Parallel matrix-vector multiply for the transposed decode shape
llmk.generate 2 Autoregressive decode primitives: greedy argmax, KV-cache append
llmk.residual 1 Elementwise add of two flat buffers
lossk.contract 8 Status codes and work budget for training objectives
lossk.cross_entropy 5 Fused projection and cross-entropy, chunked
metaldispatch.contract 2 Status codes for the Metal dispatch layer
metaldispatch.qwen3_decode 1 Int8 Metal GPU decode for Qwen3-family checkpoints
metaldispatch.qwen3_prefill_decode 10 Batched-prefill Metal GPU decode for Qwen3-family checkpoints
metaldispatch.session_prefill_policy 5 Resolved execution plan for a resident decode session
mlpk.activation 7 Gated feed-forward activations, forward and backward
mlpk.contract 11 Status codes, work budget and the GELU convention for feed-forward kernels
mlpk.embedding 3 Token embedding lookup and its scatter-add gradient
mlpk.gated_ffn 6 Fused gated feed-forward activations over flat buffers
normk.contract 5 Status codes and work budget for normalization kernels
normk.layernorm 4 LayerNorm, forward and backward, over flat row-major buffers
normk.rmsnorm 4 RMSNorm, forward and backward, over flat row-major buffers
paged_kv_cache 10 Paged KV cache: virtualization, compaction, prefix sharing
propensity 9 Propensity score matching: scores, calipers, balance checks, matched estimates
residentk.contract 12 Status codes, shape check and dispatch budget for resident sessions
residentk.probe 3 Whether this host can hold a resident GPU decode session
sparse_attention 10 Sparse, local and long-context attention variants
tensorx.buffer 7 Byte buffers tagged with a dtype and a shape
tensorx.contract 7 Status codes, finiteness predicate and shape check for tensors
tensorx.dtype 14 Dtype tags, limits, accumulation policy and round-trip
tensorx.elementwise 5 Add, subtract, multiply, divide and cast over dtype-tagged buffers, with broadcasting
tensorx.pack 12 Blockwise low-bit quantization: int4 and int8 grids, NF4 codebook, nibble packing
tensorx.reduce 7 Sum, mean, min, max, variance and standard deviation over dtype-tagged buffers
tensorx.shape 9 N-dimensional shapes, row-major strides and broadcasting
text.tokenizer 10 Text tokenization: words, sentences, n-grams, stop words, stemming
token_bucket 11 Rate limiting: token bucket, leaky bucket, sliding window
transformer_attention 11 Transformer attention calculations
transformer_blocks 10 Transformer block primitives: attention heads, feed-forward, residual, norms

kernel.CONTRACT holds every signature, including the length rules for list arguments; help(kernel.bpe_tokenizer) documents each function.

Requirements

  • Apple silicon: macOS 15 or later for the CPU build, macOS 26 or later for Metal.
  • Linux arm64 and x86-64, glibc 2.35 or later.
  • NVIDIA: driver 580 or later; the build carries kernels for T4 (sm_75), A100 (sm_80), L4 (sm_89) and H100/H200 (sm_90) and picks the one matching your GPU.
  • kernels 0.17 or later and PyTorch 2.5 to 2.14. PyTorch has to be installed: the loader picks the build for your PyTorch version. The kernel itself never imports it.

Windows is not supported.

Notes

Calls into one kernel instance run one at a time; use processes for parallelism. Runtime state does not survive fork(); start worker processes with spawn.

License

Algenta Community License 1.1 (LICENSE). Free for personal, research and open-source use, and for internal use at organizations with fewer than 50 employees and under $5M in annual revenue. Beyond that, a commercial license is required: https://algenta.ai/pricing.

Enforced in the compiled library, not just in this text: one concurrent native worker per device (ABI §9). A second process, family or thread waits its turn rather than running in parallel. That is the Community licence's worker floor made real; parallel execution comes with a commercial license.

Support

Generally Available on the platforms listed under Requirements. Within v1, functions are only added; removals or signature changes ship as v2. Platforms, accelerators and PyTorch releases not listed are not supported. Documentation: https://docs.algenta.ai (the kernels guide: https://docs.algenta.ai/guides/kernels-on-hugging-face). Community: https://discord.gg/w8NDsph9an or this repository's Community tab. Commercial licences and support: https://algenta.ai/pricing.

Downloads last month
98
algenta
mojo
cpu
cuda
metal
other
Supported hardwares new
CUDA
7.58.08.99.0
NVIDIA SXM
H200
141GB
NVIDIA SXM
H100
80GB
GPU
H800
80GB
GPU
H20
96GB
GPU
L40s
48GB
GPU
L40
48GB
GPU
L20
48GB
GPU
L4
24GB
GPU
RTX 6000 Ada
48GB
GPU
RTX 5880 Ada
48GB
RTX
RTX 5000 Ada
32GB
GPU
RTX 4500 Ada
24GB
RTX
RTX 4000 Ada
20GB
RTX
RTX 4000 SFF Ada
20GB
GPU
RTX 3500 Ada Mobile
12GB
GPU
RTX 2000 Ada
16GB
GPU
RTX A6000
48GB
GPU
RTX A5000
8GB
GPU
RTX A5000 Max-Q
16GB
GPU
RTX A5000 Mobile
16GB
GPU
RTX A4000
16GB
GPU
RTX A4000 Max-Q
8GB
GPU
RTX A4000 Mobile
8GB
GPU
RTX A3000 Mobile
6GB
GPU
RTX A2000
6GB
GPU
RTX A2000 Embedded
4GB
GPU
RTX A2000 Max-Q
4GB
GPU
RTX A2000 Mobile
4GB
GPU
A800
40GB
GPU
A100
80GB
GPU
A40
48GB
GPU
A30
24GB
GPU
A10
24GB
GPU
A2
16GB
RTX
RTX 4090
24GB
RTX
RTX 4090D
24GB
RTX
RTX 4090 Mobile
16GB
RTX
RTX 4080 SUPER
16GB
RTX
RTX 4080
16GB
RTX
RTX 4080 Mobile
12GB
RTX
RTX 4070
12GB
RTX
RTX 4070 Mobile
8GB
RTX
RTX 4070 Ti
12GB
RTX
RTX 4070 Super
12GB
RTX
RTX 4070 Ti Super
16GB
RTX
RTX 4060
8GB
RTX
RTX 4060 Ti
8GB
RTX
RTX 4090 Laptop
16GB
RTX
RTX 4080 Laptop
12GB
RTX
RTX 4070 Laptop
8GB
RTX
RTX 4060 Laptop
8GB
RTX
RTX 4050 Laptop
6GB
RTX
RTX 3090
24GB
RTX
RTX 3090 Ti
24GB
RTX
RTX 3080
12GB
RTX
RTX 3080 Ti
12GB
RTX
RTX 3080 Mobile
16GB
RTX
RTX 3070
8GB
RTX
RTX 3070 Ti
8GB
RTX
RTX 3070 Ti Mobile
8GB
RTX
RTX 3060 Ti
8GB
RTX
RTX 3060
12GB
RTX
RTX 3050
8GB
GPU
RTX 2080 Ti
11GB
GPU
RTX 2080
8GB
GPU
RTX 2080 SUPER
8GB
GPU
RTX 2070
8GB
GPU
RTX 2070 SUPER Mobile
8GB
GPU
RTX 2070 SUPER
8GB
RTX
RTX 3060 Mobile
6GB
RTX
RTX 3050 Mobile
4GB
GPU
RTX 2060 SUPER
8GB
GPU
RTX 2060
6GB
GPU
RTX 2060 12GB
12GB
GPU
RTX 2060 Mobile
6GB
GPU
RTX 2050 Mobile
4GB
GPU
RTX Titan
24GB
GPU
GTX 1660 Ti
6GB
GPU
GTX 1660 Ti Mobile
6GB
GPU
GTX 1660 SUPER
6GB
GPU
GTX 1660
6GB
GPU
GTX 1650 SUPER
4GB
GPU
GTX 1650
4GB
GPU
GTX 1650 Ti Mobile
4GB
GPU
GTX 1650 Mobile
4GB
GPU
GTX 1630
4GB
NVIDIA T4
T4
16GB
GPU
T40
24GB
GPU
T10
16GB
GPU
Quadro RTX 8000
48GB
GPU
Quadro RTX 6000
24GB
GPU
Quadro RTX 5000
16GB
Jetson
Jetson AGX Orin 64GB
64GB
Jetson
Jetson AGX Orin 32GB
32GB
Jetson
Jetson Orin NX 16GB
16GB
Jetson
Jetson Orin NX 8GB
8GB
Jetson
Jetson Orin Nano 8GB
8GB
Jetson
Jetson Orin Nano 4GB
4GB
Metal
Apple Silicon
Apple MacBook Neo
8GB
Apple Silicon
Apple M1
8GB
Apple Silicon Pro
Apple M1 Pro
16GB
Apple Silicon Max
Apple M1 Max
16GB
Apple Silicon Ultra
Apple M1 Ultra
16GB
Apple Silicon
Apple M2
8GB
Apple Silicon Pro
Apple M2 Pro
16GB
Apple Silicon Max
Apple M2 Max
32GB
Apple Silicon Ultra
Apple M2 Ultra
64GB
Apple Silicon
Apple M3
8GB
Apple Silicon Pro
Apple M3 Pro
18GB
Apple Silicon Max
Apple M3 Max
36GB
Apple Silicon Ultra
Apple M3 Ultra
96GB
Apple Silicon
Apple M4
16GB
Apple Silicon Pro
Apple M4 Pro
24GB
Apple Silicon Max
Apple M4 Max
36GB
Apple Silicon
Apple M5
16GB
Apple Silicon Pro
Apple M5 Pro
24GB
Apple Silicon Max
Apple M5 Max
36GB
Apple Silicon Ultra
Apple M5 Ultra
96GB
Torch
2.14
OS
macoslinux
Arch
x86_64aarch64