Voice / della.md
Wiself's picture
Upload della.md with huggingface_hub
c45358d verified
|
Raw History Blame Contribute Delete
12.2 kB

DELLA for voices β€” implementation plan

Goal: a native vlib/della.py implementing DELLA magnitude pruning + task-vector combine over single voice head tensors. No mergekit dependency, no vendored code. Clean-room port from the paper spec (arXiv:2406.11617), cross-checked against two read-only references: upstream mergekit/sparsify.py + generalized_task_arithmetic.py (LGPL-3.0) and the authors' declare-lab/della snapshot (no LICENSE file β€” read-only, never copy). Do not paste either source into this tree.

Non-goals (separate steps): base-head fetching (reuse _resolve_base / voice delta --base path when that lands), real 3-voice run. CLI is voice chorus / voice graft β€” this doc is the spec for both the tensor core and the UX that drives it.

Settled decisions (do not re-litigate)

These were debated in thread and are now closed. Implementation must honor them; polish wording, don't reopen.

  1. Vocabulary β€” pulse / support / lead (Apple-plain, chorus-frame native): pulse is the base (weakest presence β€” survives only where donors stay silent β€” strongest structure; constant, unshowy, everything syncs to it). Chosen over main/first (clashed with lead for the crown), anchor (broadcast frame leak), rhythm (read as strong), drone/canvas/stock/choir/bass/key (frame, pitch-expectation, or ambiguity failures). support = harmony grafted first at 0.8; lead = melody grafted last at 0.9. Prompts read exactly those three words. Output default chorus-<pulse>-<lead>.

  2. Two modes, one code path: voice chorus (plain) seals house defaults and sings straight through. voice chorus --config surfaces knobs staged per graft. Same tensor path, one extra screen per graft when flagged. No third mode.

  3. No recipe card in plain mode. Three prompts β†’ echo line (Scarlett pulse + StyleTune support + Boulesis lead β†’ chorus-scarlett-boulesis) β†’ progress β†’ voice info + audition nudge. No confirmation screen. --yes unnecessary because there is nothing to confirm; flags silently override defaults when present.

  4. Support optional, lead required. Choose support (Enter to skip): — inline, discoverable, no --2 flag. Skipping support → single graft pulse→lead at 0.9. Trio → 0.8 + 0.9 chain. A chorus with support but no lead is nonsense (no melody), so lead is never skippable. No flag beats a flag users must discover.

  5. Staged --config with auditionable duet. Pick pulse + support β†’ graft-1 knobs surface (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name chorus-<pulse>-<support>) β€” Enter accepts each, typing overrides β€” β†’ sings duet, saved as a real registry voice. Audition checkpoint: cast the duet and hear it before the lead lands (bad duet = retune graft 1 now). Pick lead β†’ graft-2 knobs surface (lead density 0.9, rest prefilled) β†’ sings final chorus-<pulse>-<lead>, provenance chaining through the duet. If support was skipped, only one knob screen appears (single 0.9 graft). Intermediate duet stays in registry either way β€” delete after if final satisfies, keep if duet is the better voice (happens).

  6. Densities are intentional: 0.8 preserves the duet's blend; 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs. Plain merge() default 0.5 when called directly (not via chorus).

  7. Output = registry. Every chorus/graft result is a named registry voice, castable immediately via voice cast. File export rides the existing voice get paths later if wanted. Every run appends operations.log.

  8. Validation before singing: distinct voices, identical head geometry or fail loud with shape table. No lineage check β€” can't verify ancestry, so say so in one line rather than pretend.

  9. Non-interactive twin β€” undecided, do not bake. Need exists (scripts / ops log) but flag spelling (--yes vs --no-interaction vs positionals) stays open. Plain voice chorus with flags silently overriding defaults is the direction; final flag names land with the CLI implementation. Spec records the need, leaves spelling out of v1.

  10. Module stays ~100 lines native numpy, no torch, no mergekit. Streaming row-blocks keep peak flat; blocking is exactly correct (sign election is per-position).

Module shape: vlib/della.py

def magprune(delta, density=0.5, epsilon=0.1, rescale="l1", rng=None): ...
def combine(base, deltas, weights, lambda_=1.0, normalize=True): ...
def merge(base, voices, weights, density=0.5, epsilon=0.1, lambda_=1.0,
          rescale="l1", normalize=True, seed=None, block_rows=4096): ...
  • magprune: row-wise rank β†’ keep-probs (densityβˆ’eps)+rank_normΒ·2Β·eps β†’ Bernoulli β†’ rescale. Works on F32 numpy; repo convention is numpy (vlib/tensors.decode_to_f32), so port with np.argsort + np.random.Generator.binomial (seeded determinism), not torch.
  • combine: weight β†’ TIES sign election (sum, always on) β†’ divide by weight sum β†’ times lambda β†’ add base. Della only; no linear variant.
  • merge: streaming driver β€” processes independent row-blocks (default 4096 rows) so peak memory stays flat regardless of head size. Full 262144Γ—2816 F32 stack (~8.8 GB for 3 deltas) is never materialized. Row-locality makes blocking exactly correct (sign election votes per-position across voices, no neighbors).

Parameters β€” house defaults (chorus recipe)

Param Scope Default Constraint
weight per voice 1.0 β€”
density per voice 0.8 first graft, 0.9 lead graft (plain merge() default 0.5 when called directly) (0,1)
epsilon per voice 0.1 first graft, 0.05 lead graft (upstream 0.15; authors ran 0.14) density Β± epsilon ∈ (0,1), else ValueError β€” 0.9+0.1 would touch 1.0, so the lead runs 0.05
lambda_ global 1.0 β€”
normalize global True (della default) zero-sum weights guarded
rescale global "l1" (stable) / "inv_p" (paper-faithful flag) β€”
seed global 42 for chorus; None for raw merge() (nondeterministic unless set) int β†’ np.random.default_rng(seed)

Rescale semantics (the one real spec fork β€” record, don't collapse):

  • "l1" (default): scale each row so sum|x| matches pre-prune. Bounded, stable.
  • "inv_p" (paper): divide each survivor by its own keep-prob. Per-element unbiased, heavy-tailed at aggressive sparsity. Authors ran eps 0.14 (fn default 0.05; CLI default 0 is a footgun β€” collapses to uniform DARE).

Chorus densities are intentional: 0.8 preserves the duet's blend, 0.9 lets the lead dominate contested coordinates. Last graft wins 90% of where it differs.

Chorus UX β€” voice chorus (settled)

Two modes, one code path:

  1. voice chorus β€” three prompts, house defaults sealed, then it sings:

    Choose pulse:   [numbered registry list]
    Choose support (Enter to skip):   [same list, pulse excluded]
    Choose lead:    [remaining]
    β†’ echo:  Scarlett pulse + StyleTune support + Boulesis lead β†’ chorus-scarlett-boulesis
    β†’ progress β†’ voice info of result + audition nudge
    

    No recipe card, no confirmation screen. Output goes straight to registry (named, castable immediately via voice cast); file export rides voice get paths later. Validation before singing: distinct voices, identical head geometry or fail loud with shape table. No lineage check (can't verify ancestry β€” say so in one line rather than pretend). Every run appends operations.log.

  2. voice chorus --config β€” staged, same three picks, knobs surface per graft:

    • Pick pulse + support β†’ graft-1 knobs (density 0.8, epsilon 0.1, lambda 1.0, seed 42, output name chorus-<pulse>-<support>) β€” Enter accepts each, typing overrides β†’ sings duet, saved as a real registry voice.
    • Audition checkpoint: cast the duet and hear it before the lead lands. Bad duet = retune graft 1 now, not after. Duet stays in registry either way β€” delete after if final satisfies, keep if duet is the better voice.
    • Pick lead β†’ graft-2 knobs (lead density 0.9, rest prefilled) β†’ sings final chorus-<pulse>-<lead>, provenance chaining through the duet (pulse β†’ support@0.8 β†’ lead@0.9).

    If support was skipped (duet mode), only one knob screen appears for the single pulse→lead graft at 0.9.

Two-voice mode: support Enter-to-skip → single graft pulse→lead at 0.9. Lead is never skippable (pulse+support with no lead is not a chorus). No --2 flag — inline (Enter to skip) beats a flag users must discover. Densities follow automatically: trio 0.8+0.9, duet 0.9.

Non-interactive twin β€” undecided (do not bake): earlier draft proposed voice chorus <pulse> <support> <lead> --lead-density 0.9 --density 0.8 --out <name> --seed 42 --yes. Flag shape (--yes vs --no-interaction vs positional) is still open β€” spec records the need (scripts / ops log) but leaves the exact spelling out of v1. Plain voice chorus with flags silently overriding defaults is the direction; final flag names land with the CLI implementation.

voice graft: single-graft primitive (pulse→donor) exposed separately for auditioning donors solo and for scripting; chorus reuses its code path.

Edge cases (all observed in upstream β€” mirror each)

  1. density >= 1 β†’ return delta untouched (no-op short-circuit).
  2. density <= 0 β†’ return zeros.
  3. density Β± epsilon outside (0,1) β†’ ValueError (fail loud, never print-warn; authors' snapshot only printed β€” do not repeat that).
  4. 1-D input β†’ unsqueeze(0) semantics, reshape back (latent for heads, free).
  5. F16 working precision β†’ compute in F32 (CPU Bernoulli has no fp16 path anywhere; StyleTune-V2 voice is F16 and trips this without upcast).
  6. All-pruned row β†’ L1 path would divide 0/0: guard before/after < 1e-7 returns masked-unscaled. inv_p path: assert no NaN/Inf post-rescale (upstream-crash behavior is correct here β€” silent NaN in a voice is worse).
  7. Zero divisors in combine: divisor == 0 β†’ 1 on the elected mask. Covers positions all voices pruned / weights cancel.
  8. Empty voice list β†’ return base.
  9. Dtype discipline: promote mixed F16/BF16 inputs to F32 working precision (matching BF16 may stay BF16); cast output back to base dtype; reuse encode_from_f32 for the final store.
  10. Determinism: unseeded Bernoulli is nondeterministic β€” seed param mandatory for reproducible voices; record seed in voice.json provenance.
  11. Chorus validation: pulse/support/lead must be distinct; head shapes must match; fail loud with shape table before any tensor work.

Verification

  • Unit tests in existing tests/_helpers fake style (tiny random tensors, no downloads): kept-fraction β‰ˆ density per row (Β± tolerance); seeded run bit-identical; density Β± epsilon violations raise; empty list returns base; F16 input β†’ F32 compute β†’ base dtype out; all-pruned row finite.
  • VOICE_NO_VENV=1 python3 -m unittest tests.test_units green (106 tests) before/after.
  • Equivalence spot-check (throwaway, needs pip install mergekit in scratch env only, never in tree): same seed/tensors through merge_tensors(..., "della") vs native merge() β€” expect close (not bit-exact: op order may differ in last-ulp; document tolerance).
  • Real run + voice info + audition/ only after the above. Paper-faithful vs upstream rescale decided by audition, not theory.

Provenance (voice.json on merged output)

Record: source voice ids + base id, method della, per-voice weight/density/epsilon, lambda_, rescale, seed, voice chorus config hash, chain (duet β†’ chorus when staged). Intermediate duet gets its own voice.json so the final can chain through it.

Open questions (do not block module work)

  • Base head source for the real run: google/gemma-4-26B-A4B-it head shard via the existing delta base path (~1.5 GB for [262144,2816], check cache first).
  • Non-interactive flag spelling for voice chorus (see undecided note above).
  • voice merge merge.json JSON shape β€” superseded by voice chorus/voice graft; keep JSON only if a file-driven batch mode is still wanted.