algenta-llm
LLM, attention, RoPE, embedding and KV-cache kernels. Compiled Mojo, loaded in-process.
52 modules · 357 functions · CPU · Cuda · Metal
Get started
pip install kernels torch
from kernels import get_kernel
kernel = get_kernel(
"thyn-ai/algenta-llm",
version=1,
trust_remote_code=["thyn-ai/algenta-llm"],
)
kernel.embeddings.cosine_similarity([1.0, 0.0], [1.0, 1.0]) # -> 0.7071
On macOS 26 the loader picks the Metal build. On macOS 15, add backend="cpu". On a Linux host with
an NVIDIA GPU and a CUDA build of PyTorch the loader picks the CUDA build. trust_remote_code names
the repositories you allow; Hugging Face's trusted publishers load without it.
Plain Python in, plain Python out. Lists, tuples, buffers and tensors are accepted wherever the
contract expects a list; structured results are dictionaries. Some functions take a list and the
number of elements to use from it, which may not exceed the list's length; multi-dimensional data is
passed flattened, row-major, with its dimensions. Functions that update an argument do so in place,
as help() says. Every call is checked against the published contract before it reaches native
code. An invalid call raises KernelError with a stable code, never a crash. Engine kernels
report shape and finiteness problems as a status; the wrapper raises KernelError named after it.
Any function can also be called by name, with args as a list or a dict of parameter names:
kernel.execute("accel.contract", "accel_backend_name", [...])
Resident decode
The Metal and CUDA builds run greedy Qwen3-family decode on the GPU. One session is open at a time; each step blocks until its token is computed and returns the token id.
with kernel.DecodeSession("/path/to/converted-checkpoint", prompt_token_ids, max_new_tokens=32) as session:
tokens = list(session)
prompt_token_ids come from the checkpoint's own tokenizer. The checkpoint directory is the
converted form (manifest.json + weights.blob) written by prepare_model_checkpoint() in the
algenta package. kernel.decode(...) runs a whole session in one call; session.step() returns
the full record for one token. Decode needs a GPU build; on the CPU build DecodeSession raises
KernelError("gpu_build_required").
The GPU builds stage decode for models with max(d_model, h_q*d_head, d_ff) <= 6144, for example
Qwen3-0.6B and Qwen3-1.7B; a larger model is refused with
KernelError("resident_dimension_too_large").
What's inside
| Module | Functions | What it does |
|---|---|---|
accel.contract |
6 | Status codes and backend selection for accelerator kernels |
accel.device |
3 | Accelerator detection for this host |
accel.gemm_metal |
3 | Threadgroup-tiled GEMM on the GPU |
attnk.backward |
2 | Gradients of scaled dot-product attention |
attnk.contract |
10 | Status codes and work budget for attention kernels |
attnk.flash |
3 | Tiled attention with online softmax |
attnk.mask |
7 | Attention masks as additive bias: causal, sliding-window, padding, packed |
attnk.rope |
9 | Rotary position embedding: frequency tables and application |
attnk.sdpa |
3 | Scaled dot-product attention, materialized reference |
attnk.tiled |
5 | Blockwise online-softmax attention over flat buffers, with grouped-query attention |
bpe_tokenizer |
10 | BPE and SentencePiece tokenization primitives |
causal_inference.propensity |
5 | Propensity scores, overlap diagnostics, inverse-probability weighting, matching |
embeddings |
20 | Vector embeddings: similarity, distance, centroids, nearest neighbours |
flash_attention |
10 | Blockwise memory-efficient attention math |
gemm.blocked |
4 | Cache-blocked SIMD GEMM and GEMV over flat row-major buffers |
gemm.contract |
5 | Status codes, dimension check and work budget for GEMM |
gemm.quantized |
4 | Matmul with dequantization fused into the k-loop |
kv_cache |
10 | KV-cache memory math and incremental decoding |
llm_sampling |
11 | Token sampling: temperature, top-k, top-p, min-p, penalties, beam scores |
llmk.contract |
5 | Status codes and shape check for the LLM kernels |
llmk.gemv_fast |
1 | Parallel matrix-vector multiply for the transposed decode shape |
llmk.generate |
2 | Autoregressive decode primitives: greedy argmax, KV-cache append |
llmk.residual |
1 | Elementwise add of two flat buffers |
lossk.contract |
8 | Status codes and work budget for training objectives |
lossk.cross_entropy |
5 | Fused projection and cross-entropy, chunked |
metaldispatch.contract |
2 | Status codes for the Metal dispatch layer |
metaldispatch.qwen3_decode |
1 | Int8 Metal GPU decode for Qwen3-family checkpoints |
metaldispatch.qwen3_prefill_decode |
10 | Batched-prefill Metal GPU decode for Qwen3-family checkpoints |
metaldispatch.session_prefill_policy |
5 | Resolved execution plan for a resident decode session |
mlpk.activation |
7 | Gated feed-forward activations, forward and backward |
mlpk.contract |
11 | Status codes, work budget and the GELU convention for feed-forward kernels |
mlpk.embedding |
3 | Token embedding lookup and its scatter-add gradient |
mlpk.gated_ffn |
6 | Fused gated feed-forward activations over flat buffers |
normk.contract |
5 | Status codes and work budget for normalization kernels |
normk.layernorm |
4 | LayerNorm, forward and backward, over flat row-major buffers |
normk.rmsnorm |
4 | RMSNorm, forward and backward, over flat row-major buffers |
paged_kv_cache |
10 | Paged KV cache: virtualization, compaction, prefix sharing |
propensity |
9 | Propensity score matching: scores, calipers, balance checks, matched estimates |
residentk.contract |
12 | Status codes, shape check and dispatch budget for resident sessions |
residentk.probe |
3 | Whether this host can hold a resident GPU decode session |
sparse_attention |
10 | Sparse, local and long-context attention variants |
tensorx.buffer |
7 | Byte buffers tagged with a dtype and a shape |
tensorx.contract |
7 | Status codes, finiteness predicate and shape check for tensors |
tensorx.dtype |
14 | Dtype tags, limits, accumulation policy and round-trip |
tensorx.elementwise |
5 | Add, subtract, multiply, divide and cast over dtype-tagged buffers, with broadcasting |
tensorx.pack |
12 | Blockwise low-bit quantization: int4 and int8 grids, NF4 codebook, nibble packing |
tensorx.reduce |
7 | Sum, mean, min, max, variance and standard deviation over dtype-tagged buffers |
tensorx.shape |
9 | N-dimensional shapes, row-major strides and broadcasting |
text.tokenizer |
10 | Text tokenization: words, sentences, n-grams, stop words, stemming |
token_bucket |
11 | Rate limiting: token bucket, leaky bucket, sliding window |
transformer_attention |
11 | Transformer attention calculations |
transformer_blocks |
10 | Transformer block primitives: attention heads, feed-forward, residual, norms |
kernel.CONTRACT holds every signature, including the length rules for list arguments;
help(kernel.bpe_tokenizer) documents each function.
Requirements
- Apple silicon: macOS 15 or later for the CPU build, macOS 26 or later for Metal.
- Linux arm64 and x86-64, glibc 2.35 or later.
- NVIDIA: driver 580 or later; the build carries kernels for T4 (sm_75), A100 (sm_80), L4 (sm_89) and H100/H200 (sm_90) and picks the one matching your GPU.
kernels0.17 or later and PyTorch 2.5 to 2.14. PyTorch has to be installed: the loader picks the build for your PyTorch version. The kernel itself never imports it.
Windows is not supported.
Notes
Calls into one kernel instance run one at a time; use processes for parallelism. Runtime state
does not survive fork(); start worker processes with spawn.
License
Algenta Community License 1.1 (LICENSE). Free for personal, research and open-source use, and
for internal use at organizations with fewer than 50 employees and under $5M in annual revenue.
Beyond that, a commercial license is required: https://algenta.ai/pricing.
Enforced in the compiled library, not just in this text: one concurrent native worker per device (ABI §9). A second process, family or thread waits its turn rather than running in parallel. That is the Community licence's worker floor made real; parallel execution comes with a commercial license.
Support
Generally Available on the platforms listed under Requirements. Within v1, functions are only added;
removals or signature changes ship as v2. Platforms, accelerators and PyTorch releases not listed are not
supported. Documentation: https://docs.algenta.ai (the kernels guide:
https://docs.algenta.ai/guides/kernels-on-hugging-face). Community: https://discord.gg/w8NDsph9an or this
repository's Community tab. Commercial licences and support: https://algenta.ai/pricing.
- Downloads last month
- 98
- Torch
- 2.14
- OS
- macoslinux
- Arch
- x86_64aarch64







