algenta-ml

Machine-learning kernels: activations, gradients, optimizers, losses, layers. Compiled Mojo, loaded in-process.

37 modules · 294 functions · CPU

Get started

pip install kernels torch
from kernels import get_kernel

kernel = get_kernel(
    "thyn-ai/algenta-ml",
    version=1,
    backend="cpu",
    trust_remote_code=["thyn-ai/algenta-ml"],
)
kernel.activations.gelu(1.0)  # -> 0.8412  learning rate at that step

backend="cpu" selects the CPU build. On a Mac the loader otherwise looks for a Metal build, which this family does not ship. trust_remote_code names the repositories you allow; Hugging Face's trusted publishers load without it.

Plain Python in, plain Python out. Lists, tuples, buffers and tensors are accepted wherever the contract expects a list; structured results are dictionaries. Some functions take a list and the number of elements to use from it, which may not exceed the list's length; multi-dimensional data is passed flattened, row-major, with its dimensions. Functions that update an argument do so in place, as help() says. Every call is checked against the published contract before it reaches native code. An invalid call raises KernelError with a stable code, never a crash. Engine kernels report shape and finiteness problems as a status; the wrapper raises KernelError named after it.

Any function can also be called by name, with args as a list or a dict of parameter names:

kernel.execute("causal_inference.base", "ci_all_finite", [...])

What's inside

Module Functions What it does
activations 10 Neural network activations: GELU, Swish/SiLU, SwiGLU, Mish, ELU, SELU, hard swish, ReLU6
bayesian_inference 10 Bayesian basics: posterior update, ELBO, mean-field steps, Laplace approximation, Gaussian KL
calibration 5 Probability calibration: isotonic regression (PAVA), Platt scaling, reliability curves
causal_inference.base 13 Shared checks for causal estimators: binary treatment, finite values, size limits, sort orders
causal_inference.dag 6 Causal graphs: acyclicity check, ancestors, descendants, d-separation, backdoor adjustment sets
causal_inference.estimators 5 Treatment effects: naive and stratified ATE, outcome regression, doubly robust AIPW, bootstrap
causal_inference.panel 2 Panel causal methods: difference-in-differences with pre-trend slope, synthetic control weights
causal_inference.sensitivity 7 Unmeasured confounding: E-values, robustness value, omitted variable bias, Rosenbaum bounds
causal_inference.uplift 2 Uplift modeling: treatment effect by score-ranked bins, Qini curve and Qini coefficient
config.toml_parser 5 TOML config parsing: sections, nested tables, strings, numbers, booleans, arrays, dotted paths
config.yaml_parser 5 YAML config parsing: nested maps, lists, scalars, and dotted-path lookups with defaults
constraint 7 Constraint solving: backtracking, N-Queens, graph coloring, Sudoku and magic square validation
decision_ml_fusion 10 Decision scoring: mean-variance utility, regret, UCB1, Chebyshev scalarization, reward shaping
gradient_descent 10 Optimizer steps: SGD, momentum, Adam, AdamW, RMSProp, AdaGrad, LARS, Lion, clipping, cosine LR
inference.assume 5 Tests with assumption verdicts: Welch t, two-proportion z, chi-squared, Levene, Shapiro-Wilk
inference.core 18 Student t and F tail probabilities, t quantiles, incomplete beta, weighted least squares solver
inference.glm 3 Generalized linear models: logistic and Poisson regression by IRLS, refusing separated data
inference.multiplicity 5 Multiple testing: Bonferroni, Holm step-down, Benjamini-Hochberg false discovery rate control
inference.power 5 Power and sample size for a two-sample mean comparison, normal and Student t, one- or two-sided
inference.resample 3 Seeded percentile bootstrap intervals for mean, median or SD, and a two-sample permutation test
inference.robust 3 Huber robust regression by iteratively reweighted least squares, with MAD scale and weights
inference_cost_latency 11 LLM serving estimates: token counts, KV-cache bytes, attention FLOPs, latency, cost, queueing
inference_engine 10 LLM serving metrics: batching priority, KV-cache hit rate, throughput, p99 latency, decode time
inference_security 10 LLM safety risk scores: prompt injection, jailbreaks, toxicity, PII exposure, data leakage
latency_optimizer 10 LLM latency heuristics: early exit, adaptive depth, layer and attention skipping, speculation
loss_functions 11 Losses: binary and categorical cross-entropy, focal, Huber, KL, contrastive, triplet, MSE, MAE
ml_metrics_extended 10 ML metrics: trapezoid AUC, average precision, NDCG, F-beta, MCC, Cohen's kappa, ECE, Brier
nn_layers 11 Neural net basics: dense layer, ReLU, sigmoid, tanh, softmax, dropout, batch norm, Xavier init
normalization_layers 10 Normalization: LayerNorm, RMSNorm, batch, group, instance and spectral norm, AdaLN, DeepNorm
optimization_constrained 8 Constrained optimization: exterior penalty, interior barrier and augmented Lagrangian methods
optimizer_full 10 Optimizers: L-BFGS scaling, Adadelta, Nadam, Sophia, Prodigy, Adan, schedule-free SGD, Muon
quantization 10 INT8 and INT4 quantization: symmetric and asymmetric quantize and dequantize, scale, zero point
quantization_stack 10 Quantization: absmax and percentile scales, SmoothQuant, GPTQ Hessian, AWQ, double quant, STE
seq2_trainer 10 Two-layer LSTM regressor training: mini-batch Adam, backprop through time, dropout, early stop
seq_training 4 Single-layer LSTM forecasting: sliding windows, full-batch backprop through time with Adam
state_space_models 10 State space model steps: zero-order-hold discretization, S4 kernel, Mamba scan, HiPPO, RetNet
training_stability 10 Training stability: dynamic loss scaling, warmup and cosine schedules, EMA weights, NaN guards

kernel.CONTRACT holds every signature, including the length rules for list arguments; help(kernel.activations) documents each function.

Not included

  • constraint.magic_square_check — reads its grid argument as square; a non-square nested list is not expressible as a length rule.

Requirements

  • Apple silicon: macOS 15 or later for the CPU build.
  • Linux arm64 and x86-64, glibc 2.35 or later.
  • kernels 0.17 or later and PyTorch 2.5 to 2.14. PyTorch has to be installed: the loader picks the build for your PyTorch version. The kernel itself never imports it.

Windows is not supported.

Notes

Calls into one kernel instance run one at a time; use processes for parallelism. Runtime state does not survive fork(); start worker processes with spawn.

License

Algenta Community License 1.1 (LICENSE). Free for personal, research and open-source use, and for internal use at organizations with fewer than 50 employees and under $5M in annual revenue. Beyond that, a commercial license is required: https://algenta.ai/pricing.

Enforced in the compiled library, not just in this text: one concurrent native worker per device (ABI §9). A second process, family or thread waits its turn rather than running in parallel. That is the Community licence's worker floor made real; parallel execution comes with a commercial license.

Support

Generally Available on the platforms listed under Requirements. Within v1, functions are only added; removals or signature changes ship as v2. Platforms, accelerators and PyTorch releases not listed are not supported. Documentation: https://docs.algenta.ai (the kernels guide: https://docs.algenta.ai/guides/kernels-on-hugging-face). Community: https://discord.gg/w8NDsph9an or this repository's Community tab. Commercial licences and support: https://algenta.ai/pricing.

Downloads last month
2
algenta
mojo
cpu
other
Torch
2.14
OS
macoslinux
Arch
x86_64aarch64