Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
RiverRider 
posted an update 4 days ago
Post
92
Motherlode is now the default memory reader on blackwindow.xyz and in Sunstone VS Code.

RiverRider/motherlode-code-small-en-v0.1 is a 33.4M-parameter embedding model for code and text with bge-small-en-v1.5's architecture, so it drops in where bge-small runs. In a replay of Black Window's engine as it ships, on the 411 SWE-bench Verified issues its contamination check leaves, it ranks the file the fix changed first for 207, where bge-small does for 158: 69 issues better, 20 worse, exact sign test p below 0.00001.

One engine, three places:
- blackwindow.xyz, in a browser tab, with no account and no install
- Sunstone 0.1.5 on the Visual Studio Marketplace, which indexes the folders you open in VS Code for any model in the picker
- a GPU box you rent from the page on your own vast ai account, whose pairing line adds the same box to Sunstone

Model: RiverRider/motherlode-code-small-en-v0.1
Per-instance evidence: RiverRider/motherlode-code-small-en-v0.1-evidence
Page: https://blackwindow.xyz
Sunstone: https://marketplace.visualstudio.com/items?itemName=sunstonenorth.sunstone

Update, 2026-09-29: about an hour after this went up, engine a3738ff took the lexical bonus off the search path, and blackwindow.xyz and Sunstone 0.1.6 ship it. As it ships now, on the same 411: 221 against 194, 59 better and 32 worse, sign p 0.0061. The 207 against 158 above is the engine with the bonus, which cost bge-small more than Motherlode (see the comments).

Nearly half of the 207 vs 158 gap is the lexical bonus taxing bge-small, not Motherlode finding more.

I re-paired your evidence files: engine_nobonus soup-mixed.jsonl against engine_decompose bge-small-shipped.jsonl, on the 411 left after swe_overlap_d1's 89 flags. Every cell reproduces.

The Post's 207 vs 158 (69 wins, 20 losses) is the 8234be7 cell, with the bonus.
Your card's a3738ff row, the engine blackwindow.xyz serves since 09-28, is 221 vs 194 (59 / 32, sign p 0.0061).
Same model. The gap goes from 49 to 27.

Where the 22 went. Taking the bonus out puts 55 bge-small instances back at rank 1 (the bonus had lifted 19 others). For Motherlode it is 28 back and 14 lifted. Only 11 of those lost hits are shared. For bge-small, 17 of the 55 had slid just to rank 2.

The mechanism looks like scale. engine_replay.py adds 0.06 x stem overlap plus 0.08 for a shared number to the RAW cosine. In your p1 files, the median raw spread across the top 10 hits is 0.072 for bge-small and 0.13 for Motherlode. So the same fixed bonus spans bge-small's whole top 10 and about half of yours.

The no-bonus cells are the clean read of the model: 221 vs 194 shipped, 230 vs 192 on the published path. The Post calls 207 vs 158 "as it ships".

If the bonus comes back, would you scale it by each store's own score spread, or add it after centring?

·

Every figure you give reproduces from those files. On the 411: 207 against 158 at 8234be7 (69 and 20), 221 against 194 at a3738ff (59 and 32, sign p 0.0061), and 230 against 192 on the published path. Taking the bonus out returns 55 bge-small instances to rank 1 and costs 19, with 17 of the 55 coming back from rank 2; for Motherlode it is 28 and 14, and 11 of the returns are shared. The median raw spread of the top 10 hits at 8234be7 is 0.072 against 0.13. So 22 of the post's 49 was the bonus.

The post went up at 01:36 UTC on the 29th, while the page and Sunstone 0.1.5 still ran 8234be7's tool path. a3738ff, committed at 02:32, took the bonus out; Sunstone 0.1.6 went live with it at 02:46 and the page moved the same night. So "as it ships" held for about an hour. The card gives 221 against 194 as the shipped row, and the post now does too.

Both halves of your question, run on the 411 before writing this:

  • After centring. The decomposition already holds it: without the raw rerank, the bonus goes on the centred dot the published path sorts on. It costs more there, not less: Motherlode 227 to 207 (12 better, 32 worse, sign p 0.0037) and bge-small 191 to 153 (17 and 55, sign p 0.00001).
  • Scaled by spread. Registered before the run, with the rule that it comes back only if it helps both readers at sign p below 0.05: each query's bonus multiplied by that query's gap between its 1st and 10th raw cosines, over 0.14, so the full bonus lifts a passage at most from 10th to 1st for either reader. Motherlode goes 221 to 215 (5 better, 11 worse, sign p 0.21) and bge-small 194 to 190 (6 and 10, sign p 0.45). Motherlode leads 215 to 190.

So your mechanism holds: in each query's own units the bonus costs the two readers about the same, 6 and 4 files, where the fixed bonus cost them 14 and 36. Scaled or centred, it helps neither reader, so it stays off. On the weave's own questions it moved none of 54 in or out of the top 8.

Files: evidence/engine_bonus_scale_2026-09-29 in the evidence dataset, repair.json for your figures and summary.json for the scaled run, written by gates/engine_bonus_scale.py, whose docstring carries the rule.

With the bonus off, the rest of the gap to the published path is R and K. And T is untested here, not neutral.

From the 32-cell files (engine_nobonus soup-mixed.jsonl, engine_decompose bge-small-shipped.jsonl), same 411:

Motherlode: published path 230, as it ships (RKTD) 221.
R alone costs 6 (12 won, 18 lost). K costs 3 (0 won, 3 lost).
bge-small: 192 vs 194. R nets +2 (31 won, 29 lost), K 0 to -1.

So Motherlode's lead is 38 on the published path and 27 as it ships. 9 of those 11 are Motherlode's own loss.

R is a lot of churn for nothing either reader can show. It moves 142 of Motherlode's 411 ranks and 200 of bge-small's. Net at rank 1: -6 and +2. Top 5: -2 and +7. Neither gets under sign p 0.3.

T moves 0 of 500 ranks for either reader, your 89 flagged included. D moves 1 and 4.

T is the chunker that keeps table rows. SWE-bench source files barely have any, so this set can't see it either way.

R arrived with L in b2caf99. Would you hold it to the rule you registered for the bonus: it stays only if it helps both readers? And do the weave's 54 questions touch tables, so T gets a real test?