Upload 3 files
Browse files- stacklm-tiny/README.md +66 -0
- stacklm-tiny/config.json +20 -0
- stacklm-tiny/pytorch_model.bin +3 -0
stacklm-tiny/README.md
ADDED
|
@@ -0,0 +1,66 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
library_name: stacklm
|
| 3 |
+
license: apache-2.0
|
| 4 |
+
tags:
|
| 5 |
+
- stacklm
|
| 6 |
+
- multi-task
|
| 7 |
+
- lora-composition
|
| 8 |
+
- query-time-fit
|
| 9 |
+
- unlearning
|
| 10 |
+
---
|
| 11 |
+
|
| 12 |
+
# stacklm-tiny
|
| 13 |
+
|
| 14 |
+
A tiny transformer (~11K parameters) demonstrating **additive stack
|
| 15 |
+
composition** for multi-task language modeling.
|
| 16 |
+
|
| 17 |
+
## Architecture
|
| 18 |
+
|
| 19 |
+
Frozen base transformer + N additive residual stacks on output logits.
|
| 20 |
+
No router parameters. Alpha (stack mixing weights) is fit at query time
|
| 21 |
+
on a small labeled example set.
|
| 22 |
+
|
| 23 |
+
f(x) = base(x) + sum_i alpha_i * stack_i(x)
|
| 24 |
+
|
| 25 |
+
## Validated claims
|
| 26 |
+
|
| 27 |
+
Tested locally, synthetic tasks, seed 0:
|
| 28 |
+
|
| 29 |
+
| Claim | Result | Baseline |
|
| 30 |
+
|---|---|---|
|
| 31 |
+
| Query-fit alpha ~ oracle | ratio 1.003 | 20 labeled examples |
|
| 32 |
+
| Query-fit vs softmax router | ratio 0.926 | Trained router |
|
| 33 |
+
| Anti-stack cancellation | 97% exact | log-space ratio 0.030 |
|
| 34 |
+
| Composition linearity | 98% exact | [1,1] vs [2,0] |
|
| 35 |
+
| Per-sample alpha beats joint | ratio 0.892 | 11% improvement |
|
| 36 |
+
|
| 37 |
+
## Usage
|
| 38 |
+
|
| 39 |
+
from stacklm_tiny import StackLM, StackLMConfig, TrainConfig
|
| 40 |
+
|
| 41 |
+
model = StackLM.from_pretrained('./stacklm-tiny')
|
| 42 |
+
alpha = model.fit_alpha_joint(X_adapt, Y_adapt, model.n_active, tcfg)
|
| 43 |
+
logits = model(X_test, alpha=alpha)
|
| 44 |
+
|
| 45 |
+
## Revocation
|
| 46 |
+
|
| 47 |
+
anti_idx = model.train_anti_stack(task, target_idx=0, tcfg=tcfg)
|
| 48 |
+
alpha = torch.tensor([1., 1.])
|
| 49 |
+
out = model(X, alpha=alpha, n=2)
|
| 50 |
+
|
| 51 |
+
## What this is / is not
|
| 52 |
+
|
| 53 |
+
**Is:** a proof-of-concept demonstrating that (a) multi-task can be
|
| 54 |
+
additive rather than routed, (b) mixing weights are optimally fitted
|
| 55 |
+
at query time, (c) adapters can be partially revoked by adding a
|
| 56 |
+
cancellation stack.
|
| 57 |
+
|
| 58 |
+
**Is not:** a useful language model. It is a demonstration model.
|
| 59 |
+
For real use cases, the same architecture applies to LoRA stacks on
|
| 60 |
+
a real base.
|
| 61 |
+
|
| 62 |
+
## Files
|
| 63 |
+
|
| 64 |
+
- pytorch_model.bin -- base + stacks weights
|
| 65 |
+
- config.json -- architecture config
|
| 66 |
+
- stacklm_tiny.py -- model code
|
stacklm-tiny/config.json
ADDED
|
@@ -0,0 +1,20 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "stacklm",
|
| 3 |
+
"architectures": [
|
| 4 |
+
"StackLM"
|
| 5 |
+
],
|
| 6 |
+
"config": {
|
| 7 |
+
"vocab_size": 16,
|
| 8 |
+
"n_tasks": 5,
|
| 9 |
+
"seq_len": 20,
|
| 10 |
+
"d_model": 32,
|
| 11 |
+
"n_heads": 4,
|
| 12 |
+
"n_layers": 1,
|
| 13 |
+
"d_ff": 64,
|
| 14 |
+
"stack_rank": 8,
|
| 15 |
+
"max_stacks": 16,
|
| 16 |
+
"dropout": 0.1,
|
| 17 |
+
"model_type": "stacklm"
|
| 18 |
+
},
|
| 19 |
+
"n_active": 4
|
| 20 |
+
}
|
stacklm-tiny/pytorch_model.bin
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:df365f8ef324f318542f106be307dc5fb44f2b3a124d64c42d8334666491752b
|
| 3 |
+
size 78091
|