goldfish-en-pol_latn__average__aligned

Training-free merged checkpoint from the Mergeability sweep (benchmark/emit_lm.py --real), produced by weight-space merging of two independently trained parents. No gradient steps were taken.

field value
pair_id goldfish-en-pol_latn
parent_a goldfish-models/eng_latn_1000mb
parent_b goldfish-models/pol_latn_1000mb
ceiling catherinearnett/B-GPT_en_pl_simultaneous
operator average
alignment aligned
align_method permutation
regime different_corpora
eval_langs eng_latn+pol_latn
nll_merge 8.3791
nll_floor 4.5982
param_coverage 1.0
MS -5.8104

How it was made

Parents were loaded, activations extracted on a shared calibration corpus, and the merge applied either naive (parents combined in their own coordinates) or aligned (parent B carried into parent A's residual-stream basis via common.alignment.residual_basis_map before merging โ€” permutation for same-width pairs, orthogonal/rectangular for cross-width).

MS is the recovery score from common.eval.mergeability_score (merged vs. floor vs. ceiling), the same normalisation used by Zhou et al., so it is comparable across rows of the sweep.

Caveats

Sub-1B merges are noisy; an aligned signal where the naive one is noise is the finding, not a bug. Rows without a joint ceiling are floor-relative and must not be read as absolute recovery.

Generated automatically โ€” see the mergeability repo.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support