|
Download README.md from zeechimp/stacklm-tiny: direct link, hf CLI and curl.
- Browser
- Download file 7.79 kB
-
https://huggingface.co/zeechimp/stacklm-tiny/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/stacklm-tiny/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/stacklm-tiny/resolve/main/README.md
7.79 kB
| language: | |
| - en | |
| license: apache-2.0 | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| tags: | |
| - stacklm | |
| - multi-task | |
| - lora-composition | |
| - query-time-fit | |
| - unlearning | |
| - custom-architecture | |
| - tiny | |
| - research | |
| model-index: | |
| - name: stacklm-tiny | |
| results: | |
| - task: | |
| type: text-generation | |
| name: Multi-task composition | |
| dataset: | |
| name: Synthetic Markov chains | |
| type: synthetic-markov-chains | |
| metrics: | |
| - name: Query-fit vs oracle (ratio) | |
| type: query_oracle_ratio | |
| value: 1.005 | |
| - name: Query-fit vs softmax (ratio) | |
| type: query_softmax_ratio | |
| value: 0.915 | |
| - name: Anti-stack cancellation | |
| type: cancellation_ratio | |
| value: 0.091 | |
| - name: Composition linearity | |
| type: linearity_log_diff | |
| value: 0.023 | |
| - name: Per-sample vs joint alpha (ratio) | |
| type: per_sample_joint_ratio | |
| value: 0.928 | |
| # stacklm-tiny | |
| A tiny transformer (~15K parameters) demonstrating **additive stack composition** for multi-task language modeling. | |
| ## Architecture | |
| Frozen base transformer + N additive residual stacks on output logits. No router parameters. Alpha (stack mixing weights) is fit at query time on a small labeled example set. | |
| The base is a 1-layer causal transformer with `d_model=32`, 4 heads, and a 16-token vocabulary. Each stack is a rank-8 low-rank projection from the base's hidden state to the output logit space. Stacks are trained sequentially: stack `i` fits the residual left by `base + stacks[0..i-1]` on task `i`. | |
| At inference, no router is used. The mixing weights `α` are fit directly on a query batch by gradient descent on a cross-entropy objective. | |
| ## Validated claims | |
| All numbers below are **mean ± std across 3 seeds** (seed 0, 1, 2). The model card template recommends reporting evaluation results in a structured format . Run `python stacklm_tiny.py --seed {0,1,2}` to reproduce. | |
| | Claim | Mean ± std | Baseline | Interpretation | | |
| |---|---|---|---| | |
| | Query-fit α ≈ oracle | **1.005 ± 0.002** | 20 labeled examples | Query-fit matches oracle within 0.5% | | |
| | Query-fit vs softmax | **0.915 ± 0.009** | Trained router | Query-fit beats softmax by 8.5% | | |
| | Anti-stack cancellation | **0.091 ± 0.029** | log-space ratio | ~91% cancellation | | |
| | Composition linearity | **0.023 ± 0.006** | [1,1] vs [2,0] | 2.3% deviation from exact | | |
| | Per-sample vs joint α | **0.928 ± 0.003** | 7.2% improvement | Per-sample α is consistently better | | |
| ### What each claim means | |
| **1. Query-fit ≈ oracle.** Fitting α on 20 labeled examples produces perplexity within 0.5% of fitting α on the full test set. The mixing weights do not need a trained router; they can be solved at query time. | |
| **2. Query-fit vs softmax.** A softmax router (a task classifier trained on 6,000 examples) is 8.5% worse than query-fit. The softmax router learns to predict a task ID from input tokens; query-fit learns the optimal mixing weights directly from labeled examples. The latter is more robust because it doesn't require the input to carry a task-identifying signal. | |
| **3. Anti-stack cancellation.** Training a stack to fit `−stack_0` reduces the composed output's divergence from the base by ~91%. This is partial cancellation, not exact erasure. For "unlearning" in the regulatory sense, this is not sufficient. For soft revocation or A/B testing, it is. | |
| **4. Composition linearity.** `[1,1]` weights on (stack, copy-stack) approximates `[2,0]` weights on (stack, zero) to within 2.3% in log-space. The raw stack logits are exactly linear; the deviation comes from the softmax, which is nonlinear. Composition is approximately linear, not exactly. | |
| **5. Per-sample α.** Fitting a separate α vector for each input sample beats fitting a single α vector for the whole batch by 7.2%, reproducibly across all 3 seeds. This is the strongest single result: the optimal mixing weights genuinely vary per input, not just per task. | |
| ## Usage | |
| ```python | |
| from stacklm_tiny import StackLM, StackLMConfig, TrainConfig | |
| import torch | |
| # Load the model | |
| model = StackLM.from_pretrained("./stacklm-tiny") | |
| tcfg = TrainConfig() | |
| # Fit alpha on 20 labeled examples | |
| X_adapt, Y_adapt = get_adapt_examples() # shape (20, seq_len-1) | |
| alpha = model.fit_alpha_joint(X_adapt, Y_adapt, model.n_active, tcfg) | |
| # Inference | |
| logits = model(X_test, alpha=alpha) | |
| # Per-sample refinement (better quality, same 20 examples) | |
| alpha_ps = model.fit_alpha_per_sample(X_test, Y_test, model.n_active, tcfg) | |
| logits = model(X_test, alpha=alpha_ps) | |
| ``` | |
| ### Revocation | |
| ```python | |
| # Train an anti-stack to cancel stack 0 | |
| anti_idx = model.train_anti_stack(task, target_idx=0, tcfg=tcfg) | |
| # Apply both: base + stack0 + anti ≈ base (91% cancellation) | |
| alpha = torch.tensor([1., 1.]) | |
| out = model(X, alpha=alpha, n=2) | |
| ``` | |
| ## Training data | |
| Synthetic Markov chains over a 16-token vocabulary. Five chains: one for the base model (task 0) and four for the stacks (tasks 1–4). Chains share 70% of their transition structure and have 30% task-specific structure. Each task has a distinct initial-token bias to give the router a weak input signal. | |
| This is a **demonstration dataset**, not a language modeling benchmark. It is designed to make the composition mechanics observable, not to test language quality. | |
| ## Training procedure | |
| - **Base**: 400 steps, AdamW, lr 1e-3, weight decay 0.05, early stopping on validation loss | |
| - **Stacks**: 300 steps each, Adam, lr 3e-3, fit on the residual left by prior stacks | |
| - **α fit**: 80 Adam steps on a length-N parameter, lr 5e-2 | |
| - **Per-sample α**: 20 Adam steps on a (B, N) parameter | |
| Hardware: CPU only. Total training time: ~90–106 seconds per seed. | |
| ## Evaluation | |
| Evaluated on 300 held-out sequences per task. The primary metric is perplexity (exponentiated cross-entropy on the task's test split). All claims are measured with the same code that produces the numbers. No cherry-picking. | |
| ## Limitations | |
| - **Tiny scale.** 15K parameters, 16-token vocabulary, 20-token sequences. Nothing about this model generalizes to real LLMs without re-testing. | |
| - **Synthetic tasks.** Markov chains, not natural language. Composition mechanics may behave differently on real text. | |
| - **No causal masking bug check.** The base uses a standard causal mask; the composition is applied to output logits post-attention. | |
| - **Single architecture.** Only one base shape tested. Different depths or attention patterns may produce different composition behavior. | |
| - **Cancellation is partial.** ~91% is not 100%. Do not rely on this for data erasure. | |
| - **Per-sample α is stochastic.** The 7.2% improvement is consistent across seeds but the mechanism is not understood. It may be an artifact of the specific synthetic setup. | |
| ## What this is / is not | |
| **Is:** a proof-of-concept demonstrating that (a) multi-task can be additive rather than routed, (b) mixing weights are optimally fitted at query time, (c) adapters can be partially revoked by adding a cancellation stack, (d) per-sample mixing weights beat batch-level weights. | |
| **Is not:** a useful language model, a benchmark result, or evidence that these claims hold at scale. For real use cases, the same architecture would apply to LoRA stacks on a real base model, and all claims would need re-testing. | |
| ## Files | |
| - `pytorch_model.bin` — base + stack weights | |
| - `config.json` — architecture config and `n_active` (number of trained stacks) | |
| - `stacklm_tiny.py` — model code (self-contained) | |
| - `README.md` — this file | |
| ## Citation | |
| ```bibtex | |
| @misc{stacklm-tiny, | |
| title={stacklm-tiny: Additive Stack Composition for Multi-Task Language Modeling}, | |
| author={zeechimp}, | |
| year={2026}, | |
| howpublished={\url{https://huggingface.co/zeechimp/stacklm-tiny}} | |
| } | |
| ``` | |
| ## Contact | |
| For questions or to report issues, open a discussion on the model repository. |