|
Download della.md from Wiself/Voice: direct link, hf CLI and curl.
- Browser
- Download file 12.2 kB
-
https://huggingface.co/Wiself/Voice/resolve/main/della.md
- Command line
-
hf download hf://Wiself/Voice/della.md
-
curl -L -o della.md https://huggingface.co/Wiself/Voice/resolve/main/della.md
12.2 kB
| # DELLA for voices β implementation plan | |
| Goal: a native `vlib/della.py` implementing DELLA magnitude pruning + task-vector combine over single voice head tensors. No mergekit dependency, no vendored code. Clean-room port from the paper spec (arXiv:2406.11617), cross-checked against two read-only references: upstream `mergekit/sparsify.py` + `generalized_task_arithmetic.py` (LGPL-3.0) and the authors' `declare-lab/della` snapshot (no LICENSE file β read-only, never copy). Do not paste either source into this tree. | |
| Non-goals (separate steps): base-head fetching (reuse `_resolve_base` / `voice delta --base` path when that lands), real 3-voice run. CLI is `voice chorus` / `voice graft` β this doc is the spec for both the tensor core and the UX that drives it. | |
| ## Settled decisions (do not re-litigate) | |
| These were debated in thread and are now closed. Implementation must honor them; polish wording, don't reopen. | |
| 1. **Vocabulary β pulse / support / lead** (Apple-plain, chorus-frame native): `pulse` is the base (weakest presence β survives only where donors stay silent β strongest structure; constant, unshowy, everything syncs to it). Chosen over `main`/`first` (clashed with lead for the crown), `anchor` (broadcast frame leak), `rhythm` (read as strong), `drone`/`canvas`/`stock`/`choir`/`bass`/`key` (frame, pitch-expectation, or ambiguity failures). `support` = harmony grafted first at 0.8; `lead` = melody grafted last at 0.9. Prompts read exactly those three words. Output default `chorus-<pulse>-<lead>`. | |
| 2. **Two modes, one code path:** `voice chorus` (plain) seals house defaults and sings straight through. `voice chorus --config` surfaces knobs staged per graft. Same tensor path, one extra screen per graft when flagged. No third mode. | |
| 3. **No recipe card in plain mode.** Three prompts β echo line (`Scarlett pulse + StyleTune support + Boulesis lead β chorus-scarlett-boulesis`) β progress β `voice info` + audition nudge. No confirmation screen. `--yes` unnecessary because there is nothing to confirm; flags silently override defaults when present. | |
| 4. **Support optional, lead required.** `Choose support (Enter to skip):` β inline, discoverable, no `--2` flag. Skipping support β single graft pulseβlead at 0.9. Trio β 0.8 + 0.9 chain. A chorus with support but no lead is nonsense (no melody), so lead is never skippable. No flag beats a flag users must discover. | |
| 5. **Staged `--config` with auditionable duet.** Pick pulse + support β graft-1 knobs surface (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name `chorus-<pulse>-<support>`) β Enter accepts each, typing overrides β β sings **duet**, saved as a real registry voice. Audition checkpoint: cast the duet and hear it before the lead lands (bad duet = retune graft 1 now). Pick lead β graft-2 knobs surface (lead density 0.9, rest prefilled) β sings final `chorus-<pulse>-<lead>`, provenance chaining through the duet. If support was skipped, only one knob screen appears (single 0.9 graft). Intermediate duet stays in registry either way β delete after if final satisfies, keep if duet is the better voice (happens). | |
| 6. **Densities are intentional:** 0.8 preserves the duet's blend; 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs. Plain `merge()` default 0.5 when called directly (not via chorus). | |
| 7. **Output = registry.** Every chorus/graft result is a named registry voice, castable immediately via `voice cast`. File export rides the existing `voice get` paths later if wanted. Every run appends `operations.log`. | |
| 8. **Validation before singing:** distinct voices, identical head geometry or fail loud with shape table. No lineage check β can't verify ancestry, so say so in one line rather than pretend. | |
| 9. **Non-interactive twin β undecided, do not bake.** Need exists (scripts / ops log) but flag spelling (`--yes` vs `--no-interaction` vs positionals) stays open. Plain `voice chorus` with flags silently overriding defaults is the direction; final flag names land with the CLI implementation. Spec records the need, leaves spelling out of v1. | |
| 10. **Module stays ~100 lines native numpy, no torch, no mergekit.** Streaming row-blocks keep peak flat; blocking is exactly correct (sign election is per-position). | |
| ## Module shape: `vlib/della.py` | |
| ```python | |
| def magprune(delta, density=0.5, epsilon=0.1, rescale="l1", rng=None): ... | |
| def combine(base, deltas, weights, lambda_=1.0, normalize=True): ... | |
| def merge(base, voices, weights, density=0.5, epsilon=0.1, lambda_=1.0, | |
| rescale="l1", normalize=True, seed=None, block_rows=4096): ... | |
| ``` | |
| - `magprune`: row-wise rank β keep-probs `(densityβeps)+rank_normΒ·2Β·eps` β Bernoulli β rescale. Works on F32 numpy; repo convention is numpy (`vlib/tensors.decode_to_f32`), so port with `np.argsort` + `np.random.Generator.binomial` (seeded determinism), not torch. | |
| - `combine`: weight β TIES sign election (`sum`, always on) β divide by weight sum β times lambda β add base. Della only; no linear variant. | |
| - `merge`: streaming driver β processes independent row-blocks (default 4096 rows) so peak memory stays flat regardless of head size. Full 262144Γ2816 F32 stack (~8.8 GB for 3 deltas) is never materialized. Row-locality makes blocking exactly correct (sign election votes per-position across voices, no neighbors). | |
| ## Parameters β house defaults (chorus recipe) | |
| | Param | Scope | Default | Constraint | | |
| |---|---|---|---| | |
| | `weight` | per voice | 1.0 | β | | |
| | `density` | per voice | **0.8 first graft, 0.9 lead graft** (plain `merge()` default 0.5 when called directly) | (0,1) | | |
| | `epsilon` | per voice | **0.1 first graft, 0.05 lead graft** (upstream 0.15; authors ran 0.14) | `density Β± epsilon β (0,1)`, else `ValueError` β 0.9+0.1 would touch 1.0, so the lead runs 0.05 | | |
| | `lambda_` | global | 1.0 | β | | |
| | `normalize` | global | True (della default) | zero-sum weights guarded | | |
| | `rescale` | global | `"l1"` (stable) / `"inv_p"` (paper-faithful flag) | β | | |
| | `seed` | global | **42** for chorus; `None` for raw `merge()` (nondeterministic unless set) | int β `np.random.default_rng(seed)` | | |
| Rescale semantics (the one real spec fork β record, don't collapse): | |
| - `"l1"` (default): scale each row so `sum|x|` matches pre-prune. Bounded, stable. | |
| - `"inv_p"` (paper): divide each survivor by its own keep-prob. Per-element unbiased, heavy-tailed at aggressive sparsity. Authors ran eps 0.14 (fn default 0.05; CLI default 0 is a footgun β collapses to uniform DARE). | |
| Chorus densities are intentional: 0.8 preserves the duet's blend, 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs. | |
| ## Chorus UX β `voice chorus` (settled) | |
| **Two modes, one code path:** | |
| 1. `voice chorus` β three prompts, house defaults sealed, then it sings: | |
| ``` | |
| Choose pulse: [numbered registry list] | |
| Choose support (Enter to skip): [same list, pulse excluded] | |
| Choose lead: [remaining] | |
| β echo: Scarlett pulse + StyleTune support + Boulesis lead β chorus-scarlett-boulesis | |
| β progress β voice info of result + audition nudge | |
| ``` | |
| No recipe card, no confirmation screen. Output goes straight to **registry** (named, castable immediately via `voice cast`); file export rides `voice get` paths later. Validation before singing: distinct voices, identical head geometry or fail loud with shape table. No lineage check (can't verify ancestry β say so in one line rather than pretend). Every run appends `operations.log`. | |
| 2. `voice chorus --config` β staged, same three picks, knobs surface **per graft**: | |
| - Pick pulse + support β graft-1 knobs (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name `chorus-<pulse>-<support>`) β Enter accepts each, typing overrides β sings **duet**, saved as a real registry voice. | |
| - Audition checkpoint: cast the duet and hear it before the lead lands. Bad duet = retune graft 1 now, not after. Duet stays in registry either way β delete after if final satisfies, keep if duet is the better voice. | |
| - Pick lead β graft-2 knobs (lead density 0.9, rest prefilled) β sings final `chorus-<pulse>-<lead>`, provenance chaining through the duet (pulse β support@0.8 β lead@0.9). | |
| If support was skipped (duet mode), only one knob screen appears for the single pulseβlead graft at 0.9. | |
| **Two-voice mode:** support Enter-to-skip β single graft pulseβlead at 0.9. Lead is never skippable (pulse+support with no lead is not a chorus). No `--2` flag β inline `(Enter to skip)` beats a flag users must discover. Densities follow automatically: trio 0.8+0.9, duet 0.9. | |
| **Non-interactive twin β undecided (do not bake):** earlier draft proposed `voice chorus <pulse> <support> <lead> --lead-density 0.9 --density 0.8 --out <name> --seed 42 --yes`. Flag shape (`--yes` vs `--no-interaction` vs positional) is still open β spec records the need (scripts / ops log) but leaves the exact spelling out of v1. Plain `voice chorus` with flags silently overriding defaults is the direction; final flag names land with the CLI implementation. | |
| **`voice graft`:** single-graft primitive (pulseβdonor) exposed separately for auditioning donors solo and for scripting; chorus reuses its code path. | |
| ## Edge cases (all observed in upstream β mirror each) | |
| 1. `density >= 1` β return delta untouched (no-op short-circuit). | |
| 2. `density <= 0` β return zeros. | |
| 3. `density Β± epsilon` outside (0,1) β `ValueError` (fail loud, never print-warn; authors' snapshot only printed β do not repeat that). | |
| 4. 1-D input β `unsqueeze(0)` semantics, reshape back (latent for heads, free). | |
| 5. F16 working precision β compute in F32 (CPU Bernoulli has no fp16 path anywhere; StyleTune-V2 voice is F16 and trips this without upcast). | |
| 6. All-pruned row β L1 path would divide 0/0: guard `before/after < 1e-7` returns masked-unscaled. `inv_p` path: assert no NaN/Inf post-rescale (upstream-crash behavior is correct here β silent NaN in a voice is worse). | |
| 7. Zero divisors in combine: `divisor == 0 β 1` on the elected mask. Covers positions all voices pruned / weights cancel. | |
| 8. Empty voice list β return base. | |
| 9. Dtype discipline: promote mixed F16/BF16 inputs to F32 working precision (matching BF16 may stay BF16); cast output back to base dtype; reuse `encode_from_f32` for the final store. | |
| 10. Determinism: unseeded Bernoulli is nondeterministic β `seed` param mandatory for reproducible voices; record seed in `voice.json` provenance. | |
| 11. Chorus validation: pulse/support/lead must be distinct; head shapes must match; fail loud with shape table before any tensor work. | |
| ## Verification | |
| - Unit tests in existing `tests/_helpers` fake style (tiny random tensors, no downloads): kept-fraction β density per row (Β± tolerance); seeded run bit-identical; `density Β± epsilon` violations raise; empty list returns base; F16 input β F32 compute β base dtype out; all-pruned row finite. | |
| - `VOICE_NO_VENV=1 python3 -m unittest tests.test_units` green (106 tests) before/after. | |
| - Equivalence spot-check (throwaway, needs `pip install mergekit` in scratch env only, never in tree): same seed/tensors through `merge_tensors(..., "della")` vs native `merge()` β expect close (not bit-exact: op order may differ in last-ulp; document tolerance). | |
| - Real run + `voice info` + `audition/` only after the above. Paper-faithful vs upstream rescale decided by audition, not theory. | |
| ## Provenance (`voice.json` on merged output) | |
| Record: source voice ids + base id, method `della`, per-voice `weight`/`density`/`epsilon`, `lambda_`, `rescale`, `seed`, `voice chorus` config hash, chain (duet β chorus when staged). Intermediate duet gets its own `voice.json` so the final can chain through it. | |
| ## Open questions (do not block module work) | |
| - Base head source for the real run: `google/gemma-4-26B-A4B-it` head shard via the existing delta base path (~1.5 GB for [262144,2816], check cache first). | |
| - Non-interactive flag spelling for `voice chorus` (see undecided note above). | |
| - `voice merge merge.json` JSON shape β superseded by `voice chorus`/`voice graft`; keep JSON only if a file-driven batch mode is still wanted. | |