Download della.md from Wiself/Voice: direct link, hf CLI and curl.
- Browser
- Download file 12.2 kB
-
https://huggingface.co/Wiself/Voice/resolve/main/della.md
- Command line
-
hf download hf://Wiself/Voice/della.md
-
curl -L -o della.md https://huggingface.co/Wiself/Voice/resolve/main/della.md
DELLA for voices β implementation plan
Goal: a native vlib/della.py implementing DELLA magnitude pruning + task-vector combine over single voice head tensors. No mergekit dependency, no vendored code. Clean-room port from the paper spec (arXiv:2406.11617), cross-checked against two read-only references: upstream mergekit/sparsify.py + generalized_task_arithmetic.py (LGPL-3.0) and the authors' declare-lab/della snapshot (no LICENSE file β read-only, never copy). Do not paste either source into this tree.
Non-goals (separate steps): base-head fetching (reuse _resolve_base / voice delta --base path when that lands), real 3-voice run. CLI is voice chorus / voice graft β this doc is the spec for both the tensor core and the UX that drives it.
Settled decisions (do not re-litigate)
These were debated in thread and are now closed. Implementation must honor them; polish wording, don't reopen.
Vocabulary β pulse / support / lead (Apple-plain, chorus-frame native):
pulseis the base (weakest presence β survives only where donors stay silent β strongest structure; constant, unshowy, everything syncs to it). Chosen overmain/first(clashed with lead for the crown),anchor(broadcast frame leak),rhythm(read as strong),drone/canvas/stock/choir/bass/key(frame, pitch-expectation, or ambiguity failures).support= harmony grafted first at 0.8;lead= melody grafted last at 0.9. Prompts read exactly those three words. Output defaultchorus-<pulse>-<lead>.Two modes, one code path:
voice chorus(plain) seals house defaults and sings straight through.voice chorus --configsurfaces knobs staged per graft. Same tensor path, one extra screen per graft when flagged. No third mode.No recipe card in plain mode. Three prompts β echo line (
Scarlett pulse + StyleTune support + Boulesis lead β chorus-scarlett-boulesis) β progress βvoice info+ audition nudge. No confirmation screen.--yesunnecessary because there is nothing to confirm; flags silently override defaults when present.Support optional, lead required.
Choose support (Enter to skip):β inline, discoverable, no--2flag. Skipping support β single graft pulseβlead at 0.9. Trio β 0.8 + 0.9 chain. A chorus with support but no lead is nonsense (no melody), so lead is never skippable. No flag beats a flag users must discover.Staged
--configwith auditionable duet. Pick pulse + support β graft-1 knobs surface (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output namechorus-<pulse>-<support>) β Enter accepts each, typing overrides β β sings duet, saved as a real registry voice. Audition checkpoint: cast the duet and hear it before the lead lands (bad duet = retune graft 1 now). Pick lead β graft-2 knobs surface (lead density 0.9, rest prefilled) β sings finalchorus-<pulse>-<lead>, provenance chaining through the duet. If support was skipped, only one knob screen appears (single 0.9 graft). Intermediate duet stays in registry either way β delete after if final satisfies, keep if duet is the better voice (happens).Densities are intentional: 0.8 preserves the duet's blend; 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs. Plain
merge()default 0.5 when called directly (not via chorus).Output = registry. Every chorus/graft result is a named registry voice, castable immediately via
voice cast. File export rides the existingvoice getpaths later if wanted. Every run appendsoperations.log.Validation before singing: distinct voices, identical head geometry or fail loud with shape table. No lineage check β can't verify ancestry, so say so in one line rather than pretend.
Non-interactive twin β undecided, do not bake. Need exists (scripts / ops log) but flag spelling (
--yesvs--no-interactionvs positionals) stays open. Plainvoice choruswith flags silently overriding defaults is the direction; final flag names land with the CLI implementation. Spec records the need, leaves spelling out of v1.Module stays ~100 lines native numpy, no torch, no mergekit. Streaming row-blocks keep peak flat; blocking is exactly correct (sign election is per-position).
Module shape: vlib/della.py
def magprune(delta, density=0.5, epsilon=0.1, rescale="l1", rng=None): ...
def combine(base, deltas, weights, lambda_=1.0, normalize=True): ...
def merge(base, voices, weights, density=0.5, epsilon=0.1, lambda_=1.0,
rescale="l1", normalize=True, seed=None, block_rows=4096): ...
magprune: row-wise rank β keep-probs(densityβeps)+rank_normΒ·2Β·epsβ Bernoulli β rescale. Works on F32 numpy; repo convention is numpy (vlib/tensors.decode_to_f32), so port withnp.argsort+np.random.Generator.binomial(seeded determinism), not torch.combine: weight β TIES sign election (sum, always on) β divide by weight sum β times lambda β add base. Della only; no linear variant.merge: streaming driver β processes independent row-blocks (default 4096 rows) so peak memory stays flat regardless of head size. Full 262144Γ2816 F32 stack (~8.8 GB for 3 deltas) is never materialized. Row-locality makes blocking exactly correct (sign election votes per-position across voices, no neighbors).
Parameters β house defaults (chorus recipe)
| Param | Scope | Default | Constraint |
|---|---|---|---|
weight |
per voice | 1.0 | β |
density |
per voice | 0.8 first graft, 0.9 lead graft (plain merge() default 0.5 when called directly) |
(0,1) |
epsilon |
per voice | 0.1 first graft, 0.05 lead graft (upstream 0.15; authors ran 0.14) | density Β± epsilon β (0,1), else ValueError β 0.9+0.1 would touch 1.0, so the lead runs 0.05 |
lambda_ |
global | 1.0 | β |
normalize |
global | True (della default) | zero-sum weights guarded |
rescale |
global | "l1" (stable) / "inv_p" (paper-faithful flag) |
β |
seed |
global | 42 for chorus; None for raw merge() (nondeterministic unless set) |
int β np.random.default_rng(seed) |
Rescale semantics (the one real spec fork β record, don't collapse):
"l1"(default): scale each row sosum|x|matches pre-prune. Bounded, stable."inv_p"(paper): divide each survivor by its own keep-prob. Per-element unbiased, heavy-tailed at aggressive sparsity. Authors ran eps 0.14 (fn default 0.05; CLI default 0 is a footgun β collapses to uniform DARE).
Chorus densities are intentional: 0.8 preserves the duet's blend, 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs.
Chorus UX β voice chorus (settled)
Two modes, one code path:
voice chorusβ three prompts, house defaults sealed, then it sings:Choose pulse: [numbered registry list] Choose support (Enter to skip): [same list, pulse excluded] Choose lead: [remaining] β echo: Scarlett pulse + StyleTune support + Boulesis lead β chorus-scarlett-boulesis β progress β voice info of result + audition nudgeNo recipe card, no confirmation screen. Output goes straight to registry (named, castable immediately via
voice cast); file export ridesvoice getpaths later. Validation before singing: distinct voices, identical head geometry or fail loud with shape table. No lineage check (can't verify ancestry β say so in one line rather than pretend). Every run appendsoperations.log.voice chorus --configβ staged, same three picks, knobs surface per graft:- Pick pulse + support β graft-1 knobs (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name
chorus-<pulse>-<support>) β Enter accepts each, typing overrides β sings duet, saved as a real registry voice. - Audition checkpoint: cast the duet and hear it before the lead lands. Bad duet = retune graft 1 now, not after. Duet stays in registry either way β delete after if final satisfies, keep if duet is the better voice.
- Pick lead β graft-2 knobs (lead density 0.9, rest prefilled) β sings final
chorus-<pulse>-<lead>, provenance chaining through the duet (pulse β support@0.8 β lead@0.9).
If support was skipped (duet mode), only one knob screen appears for the single pulseβlead graft at 0.9.
- Pick pulse + support β graft-1 knobs (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name
Two-voice mode: support Enter-to-skip β single graft pulseβlead at 0.9. Lead is never skippable (pulse+support with no lead is not a chorus). No --2 flag β inline (Enter to skip) beats a flag users must discover. Densities follow automatically: trio 0.8+0.9, duet 0.9.
Non-interactive twin β undecided (do not bake): earlier draft proposed voice chorus <pulse> <support> <lead> --lead-density 0.9 --density 0.8 --out <name> --seed 42 --yes. Flag shape (--yes vs --no-interaction vs positional) is still open β spec records the need (scripts / ops log) but leaves the exact spelling out of v1. Plain voice chorus with flags silently overriding defaults is the direction; final flag names land with the CLI implementation.
voice graft: single-graft primitive (pulseβdonor) exposed separately for auditioning donors solo and for scripting; chorus reuses its code path.
Edge cases (all observed in upstream β mirror each)
density >= 1β return delta untouched (no-op short-circuit).density <= 0β return zeros.density Β± epsilonoutside (0,1) βValueError(fail loud, never print-warn; authors' snapshot only printed β do not repeat that).- 1-D input β
unsqueeze(0)semantics, reshape back (latent for heads, free). - F16 working precision β compute in F32 (CPU Bernoulli has no fp16 path anywhere; StyleTune-V2 voice is F16 and trips this without upcast).
- All-pruned row β L1 path would divide 0/0: guard
before/after < 1e-7returns masked-unscaled.inv_ppath: assert no NaN/Inf post-rescale (upstream-crash behavior is correct here β silent NaN in a voice is worse). - Zero divisors in combine:
divisor == 0 β 1on the elected mask. Covers positions all voices pruned / weights cancel. - Empty voice list β return base.
- Dtype discipline: promote mixed F16/BF16 inputs to F32 working precision (matching BF16 may stay BF16); cast output back to base dtype; reuse
encode_from_f32for the final store. - Determinism: unseeded Bernoulli is nondeterministic β
seedparam mandatory for reproducible voices; record seed invoice.jsonprovenance. - Chorus validation: pulse/support/lead must be distinct; head shapes must match; fail loud with shape table before any tensor work.
Verification
- Unit tests in existing
tests/_helpersfake style (tiny random tensors, no downloads): kept-fraction β density per row (Β± tolerance); seeded run bit-identical;density Β± epsilonviolations raise; empty list returns base; F16 input β F32 compute β base dtype out; all-pruned row finite. VOICE_NO_VENV=1 python3 -m unittest tests.test_unitsgreen (106 tests) before/after.- Equivalence spot-check (throwaway, needs
pip install mergekitin scratch env only, never in tree): same seed/tensors throughmerge_tensors(..., "della")vs nativemerge()β expect close (not bit-exact: op order may differ in last-ulp; document tolerance). - Real run +
voice info+audition/only after the above. Paper-faithful vs upstream rescale decided by audition, not theory.
Provenance (voice.json on merged output)
Record: source voice ids + base id, method della, per-voice weight/density/epsilon, lambda_, rescale, seed, voice chorus config hash, chain (duet β chorus when staged). Intermediate duet gets its own voice.json so the final can chain through it.
Open questions (do not block module work)
- Base head source for the real run:
google/gemma-4-26B-A4B-ithead shard via the existing delta base path (~1.5 GB for [262144,2816], check cache first). - Non-interactive flag spelling for
voice chorus(see undecided note above). voice merge merge.jsonJSON shape β superseded byvoice chorus/voice graft; keep JSON only if a file-driven batch mode is still wanted.