Papers
arxiv:2609.32814

Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers

Published on Sep 26
· Submitted by
ilya koziev
on Sep 29
Authors:
,

Abstract

Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.

Community

Paper submitter

We study whether the matrix multiplication used in Transformer projections can be replaced by a different, cheaper associative product while keeping the same weight bank and parameter count.

We construct an associative algebra product with a lower bilinear rank than standard matrix multiplication. For the q=2 case, the product has rank 6, compared with rank 7 for Strassen's 2×2 algorithm. This allows us to replace linear projections without introducing additional nonlinearities, while reducing their arithmetic cost.

We evaluate the approach at three levels: GPU projection kernels across several Transformer shapes, including Qwen and DeepSeek; and two approximately 110M-parameter decoder-only LMs trained for 12.3B tokens, differing only in the MLP multiplication law. The algebraic model achieves a 6.2–7.8% increase in end-to-end generation throughput across four prompt domains, while obtaining lower scores on GSM8K, MBPP, and IFEval.

The results provide a small-scale feasibility and trainability study of changing the multiplication law of Transformer projections while retaining the full parameter bank.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.32814
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.32814 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.32814 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.32814 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.