algenta-ml
Machine-learning kernels: activations, gradients, optimizers, losses, layers. Compiled Mojo, loaded in-process.
37 modules · 294 functions · CPU
Get started
pip install kernels torch
from kernels import get_kernel
kernel = get_kernel(
"thyn-ai/algenta-ml",
version=1,
backend="cpu",
trust_remote_code=["thyn-ai/algenta-ml"],
)
kernel.activations.gelu(1.0) # -> 0.8412 learning rate at that step
backend="cpu" selects the CPU build. On a Mac the loader otherwise looks for a Metal build, which
this family does not ship. trust_remote_code names the repositories you allow; Hugging Face's
trusted publishers load without it.
Plain Python in, plain Python out. Lists, tuples, buffers and tensors are accepted wherever the
contract expects a list; structured results are dictionaries. Some functions take a list and the
number of elements to use from it, which may not exceed the list's length; multi-dimensional data is
passed flattened, row-major, with its dimensions. Functions that update an argument do so in place,
as help() says. Every call is checked against the published contract before it reaches native
code. An invalid call raises KernelError with a stable code, never a crash. Engine kernels
report shape and finiteness problems as a status; the wrapper raises KernelError named after it.
Any function can also be called by name, with args as a list or a dict of parameter names:
kernel.execute("causal_inference.base", "ci_all_finite", [...])
What's inside
| Module | Functions | What it does |
|---|---|---|
activations |
10 | Neural network activations: GELU, Swish/SiLU, SwiGLU, Mish, ELU, SELU, hard swish, ReLU6 |
bayesian_inference |
10 | Bayesian basics: posterior update, ELBO, mean-field steps, Laplace approximation, Gaussian KL |
calibration |
5 | Probability calibration: isotonic regression (PAVA), Platt scaling, reliability curves |
causal_inference.base |
13 | Shared checks for causal estimators: binary treatment, finite values, size limits, sort orders |
causal_inference.dag |
6 | Causal graphs: acyclicity check, ancestors, descendants, d-separation, backdoor adjustment sets |
causal_inference.estimators |
5 | Treatment effects: naive and stratified ATE, outcome regression, doubly robust AIPW, bootstrap |
causal_inference.panel |
2 | Panel causal methods: difference-in-differences with pre-trend slope, synthetic control weights |
causal_inference.sensitivity |
7 | Unmeasured confounding: E-values, robustness value, omitted variable bias, Rosenbaum bounds |
causal_inference.uplift |
2 | Uplift modeling: treatment effect by score-ranked bins, Qini curve and Qini coefficient |
config.toml_parser |
5 | TOML config parsing: sections, nested tables, strings, numbers, booleans, arrays, dotted paths |
config.yaml_parser |
5 | YAML config parsing: nested maps, lists, scalars, and dotted-path lookups with defaults |
constraint |
7 | Constraint solving: backtracking, N-Queens, graph coloring, Sudoku and magic square validation |
decision_ml_fusion |
10 | Decision scoring: mean-variance utility, regret, UCB1, Chebyshev scalarization, reward shaping |
gradient_descent |
10 | Optimizer steps: SGD, momentum, Adam, AdamW, RMSProp, AdaGrad, LARS, Lion, clipping, cosine LR |
inference.assume |
5 | Tests with assumption verdicts: Welch t, two-proportion z, chi-squared, Levene, Shapiro-Wilk |
inference.core |
18 | Student t and F tail probabilities, t quantiles, incomplete beta, weighted least squares solver |
inference.glm |
3 | Generalized linear models: logistic and Poisson regression by IRLS, refusing separated data |
inference.multiplicity |
5 | Multiple testing: Bonferroni, Holm step-down, Benjamini-Hochberg false discovery rate control |
inference.power |
5 | Power and sample size for a two-sample mean comparison, normal and Student t, one- or two-sided |
inference.resample |
3 | Seeded percentile bootstrap intervals for mean, median or SD, and a two-sample permutation test |
inference.robust |
3 | Huber robust regression by iteratively reweighted least squares, with MAD scale and weights |
inference_cost_latency |
11 | LLM serving estimates: token counts, KV-cache bytes, attention FLOPs, latency, cost, queueing |
inference_engine |
10 | LLM serving metrics: batching priority, KV-cache hit rate, throughput, p99 latency, decode time |
inference_security |
10 | LLM safety risk scores: prompt injection, jailbreaks, toxicity, PII exposure, data leakage |
latency_optimizer |
10 | LLM latency heuristics: early exit, adaptive depth, layer and attention skipping, speculation |
loss_functions |
11 | Losses: binary and categorical cross-entropy, focal, Huber, KL, contrastive, triplet, MSE, MAE |
ml_metrics_extended |
10 | ML metrics: trapezoid AUC, average precision, NDCG, F-beta, MCC, Cohen's kappa, ECE, Brier |
nn_layers |
11 | Neural net basics: dense layer, ReLU, sigmoid, tanh, softmax, dropout, batch norm, Xavier init |
normalization_layers |
10 | Normalization: LayerNorm, RMSNorm, batch, group, instance and spectral norm, AdaLN, DeepNorm |
optimization_constrained |
8 | Constrained optimization: exterior penalty, interior barrier and augmented Lagrangian methods |
optimizer_full |
10 | Optimizers: L-BFGS scaling, Adadelta, Nadam, Sophia, Prodigy, Adan, schedule-free SGD, Muon |
quantization |
10 | INT8 and INT4 quantization: symmetric and asymmetric quantize and dequantize, scale, zero point |
quantization_stack |
10 | Quantization: absmax and percentile scales, SmoothQuant, GPTQ Hessian, AWQ, double quant, STE |
seq2_trainer |
10 | Two-layer LSTM regressor training: mini-batch Adam, backprop through time, dropout, early stop |
seq_training |
4 | Single-layer LSTM forecasting: sliding windows, full-batch backprop through time with Adam |
state_space_models |
10 | State space model steps: zero-order-hold discretization, S4 kernel, Mamba scan, HiPPO, RetNet |
training_stability |
10 | Training stability: dynamic loss scaling, warmup and cosine schedules, EMA weights, NaN guards |
kernel.CONTRACT holds every signature, including the length rules for list arguments;
help(kernel.activations) documents each function.
Not included
constraint.magic_square_check— reads its grid argument as square; a non-square nested list is not expressible as a length rule.
Requirements
- Apple silicon: macOS 15 or later for the CPU build.
- Linux arm64 and x86-64, glibc 2.35 or later.
kernels0.17 or later and PyTorch 2.5 to 2.14. PyTorch has to be installed: the loader picks the build for your PyTorch version. The kernel itself never imports it.
Windows is not supported.
Notes
Calls into one kernel instance run one at a time; use processes for parallelism. Runtime state
does not survive fork(); start worker processes with spawn.
License
Algenta Community License 1.1 (LICENSE). Free for personal, research and open-source use, and
for internal use at organizations with fewer than 50 employees and under $5M in annual revenue.
Beyond that, a commercial license is required: https://algenta.ai/pricing.
Enforced in the compiled library, not just in this text: one concurrent native worker per device (ABI §9). A second process, family or thread waits its turn rather than running in parallel. That is the Community licence's worker floor made real; parallel execution comes with a commercial license.
Support
Generally Available on the platforms listed under Requirements. Within v1, functions are only added;
removals or signature changes ship as v2. Platforms, accelerators and PyTorch releases not listed are not
supported. Documentation: https://docs.algenta.ai (the kernels guide:
https://docs.algenta.ai/guides/kernels-on-hugging-face). Community: https://discord.gg/w8NDsph9an or this
repository's Community tab. Commercial licences and support: https://algenta.ai/pricing.
- Downloads last month
- 2
- Torch
- 2.14
- OS
- macoslinux
- Arch
- x86_64aarch64