maxffarrell commited on
Commit
546d6ba
·
verified ·
1 Parent(s): 3b96e72

Upload current experimental gpu-search model with runtime and provenance

Browse files
LICENSE ADDED
@@ -0,0 +1,21 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Max Farrell
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
README.md ADDED
@@ -0,0 +1,61 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ language:
3
+ - en
4
+ license: mit
5
+ tags:
6
+ - gpu-search
7
+ - experimental
8
+ - embeddings
9
+ - int8
10
+ - typescript
11
+ - navigation
12
+ ---
13
+
14
+ # gpu-search
15
+
16
+ A small experimental encoder for short English software-navigation queries and candidate menus. This repository contains the **current demo checkpoint**, `navigation-align0p5-seed29-word085`, with its original **32,768-byte int8 weights**, manifest, feature specification, and TypeScript CPU runtime.
17
+
18
+ [Source code](https://github.com/maxffarrell/gpu-search) · [Live demo](https://gpu-search.vercel.app) · [Detailed model history](docs/model-card.md) · [Training-source attribution](docs/data-sources.md)
19
+
20
+ ## Model and intended use
21
+
22
+ The encoder pools hashed word and character n-gram embeddings into a normalized 16-dimensional vector. Each table has 1,024 rows. Candidate labels and aliases are averaged, with optional context contribution. Ranking uses cosine similarity. The manifest supplies the exact int8 dequantization scales; weights alone are insufficient to reproduce the selected model.
23
+
24
+ This checkpoint reduces the previous navigation model's word-table scale by 15%. It is intended for research and explicitly labeled model suggestions in command palettes, settings, and navigation menus. The application's deterministic lexical matching takes precedence over model suggestions.
25
+
26
+ **Experimental:** no learned model has passed the project's production semantic release gate, and `validatedSemanticCutoff` is null. Scores are not confidence probabilities. Raw inference still ranks **Invoices** above **Profile** for `profle`; the demo handles that query with its separate lexical tier. No WebGPU runtime or GPU acceleration is claimed. This custom binary format is loaded by the included runtime, rather than Transformers `AutoModel`.
27
+
28
+ ## Load the model
29
+
30
+ Download this repository, preserve its directory structure, and use a TypeScript-capable environment with Web Crypto. For example, save this as `example.ts` at the repository root and run `npx tsx example.ts` with Node.js 22.12+:
31
+
32
+ ```ts
33
+ import { readFile } from 'node:fs/promises';
34
+ import { loadModel } from './packages/model/runtime.ts';
35
+
36
+ const manifest = JSON.parse(await readFile('./manifest.json', 'utf8'));
37
+ const bytes = await readFile('./weights.bin');
38
+ const payload = bytes.buffer.slice(bytes.byteOffset, bytes.byteOffset + bytes.byteLength);
39
+ const model = await loadModel(manifest, payload);
40
+ const index = model.prepare([
41
+ { id: 'profile', label: 'Profile' },
42
+ { id: 'members', label: 'Members' },
43
+ { id: 'invoices', label: 'Invoices' },
44
+ ]);
45
+ console.log(index.score('coworkers'));
46
+ index.dispose();
47
+ ```
48
+
49
+ Browser applications can fetch `manifest.json` and `weights.bin`, call `loadModel(manifest, await response.arrayBuffer())`, and reuse a prepared index. The loader verifies the payload hash and manifest format. See [prepared-index documentation](docs/prepared-model-index.md) for lifecycle and input limits.
50
+
51
+ ## Evidence and limitations
52
+
53
+ The selected scale adjustment improved top-one accuracy on 600 reserved **synthetic** typo cases from 64.33% to 67.33%, while retaining the 32 KiB payload. These historical results are documented in [typo-weight experiments](docs/typo-weight-experiments.md), with original evaluation receipts under `eval/`. They do not establish broad semantic generalization, human-reviewed relevance, or reliable abstention. The spelling holdout is now consumed.
54
+
55
+ Hash collisions, bag-of-features pooling, short unseen labels, ambiguous menus, and action opposites remain limitations. Preserve lexical precedence and treat learned results as suggestions. Documentation includes historical checkpoints; this repository ships only the current demo artifact identified above.
56
+
57
+ ## Provenance and license
58
+
59
+ Copied from source commit `0277121c3e4127c96604b25c73a9af627663da80`, artifact path `packages/model/experiments/navigation-align0p5-seed29-word085`. `provenance.json` records SHA-256 hashes of copied files. Weight SHA-256: `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`.
60
+
61
+ Project code and original model artifacts are distributed under the source project's MIT license. Training and research sources retain their own terms: CLINC150 (CC BY 3.0), BANKING77 (CC BY 4.0), VS Code (MIT), and navigation metadata from GNOME, KDE, and Xfce. See [data attribution](docs/data-sources.md), [navigation provenance](docs/navigation-data.md), and [credits](docs/credits.md). Raw third-party datasets and other experimental checkpoints are not included.
docs/credits.md ADDED
@@ -0,0 +1,33 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Reference provenance
2
+
3
+ Inspected 2026-09-12 via primary GitHub repository metadata and READMEs. All three repository license identifiers were MIT. No source code, model weights, datasets, branding assets, or benchmark figures were copied. These are architectural inspirations, not dependencies or claims of equivalent performance.
4
+
5
+ | Project | Pinned revision | Contribution to the direction |
6
+ | --- | --- | --- |
7
+ | [gpu-lexer — Shu Ding / Vercel Labs](https://github.com/vercel-labs/gpu-lexer) · [demo](https://gpu-lexer.vercel.app) | `3bf10853186b20070096ed40d5986e0bf48742c1` | Small specialized learned browser workloads and custom WebGPU execution |
8
+ | [gpu-time — Arik Chakma](https://github.com/arikchakma/gpu-time) · [demo](https://gpu-time.arikko.dev) | `aba27e54aabe7310cba5c160fa2079045096ffb1` | Local natural-language functionality with explicit CPU/WebGPU backends |
9
+ | [gpu-cron — Manu Schiller](https://github.com/manuschillerdev/gpu-cron) · [demo](https://gpu-cron.vercel.app) | `0ddaa91d88c33838e1f76c11c0d830810f1d7914` | Compact learned natural-language parsing as a focused browser experiment |
10
+
11
+ Additional primary references: [fastText subword representations](https://fasttext.cc/docs/en/unsupervised-tutorial.html), [StarSpace shared-feature retrieval](https://arxiv.org/abs/1709.03856), [Apple MLX](https://ml-explore.github.io/mlx/build/html/index.html), and [W3C WGSL](https://www.w3.org/TR/WGSL/). Pooled encoders here are inspired baselines, not exact reproductions. Vite, TypeScript, pnpm, esbuild, Playwright, NumPy, and uv provide development and validation tooling; their licenses remain with their respective projects. Mind2Web and WorkArena were suggested research sources in the spec but their data was not ingested.
12
+
13
+
14
+ ## Desktop navigation metadata
15
+
16
+ Author-provided English search keywords, destination names, and descriptions were extracted from these official repositories without executing their code:
17
+
18
+ | Provider | Source | Pinned revision |
19
+ | --- | --- | --- |
20
+ | GNOME | [GNOME Settings / gnome-control-center](https://github.com/GNOME/gnome-control-center) | `cd79a897190989ca395c6f00962f0929734103c7` |
21
+ | KDE | [Plasma Desktop](https://github.com/KDE/plasma-desktop) | `19dae94ec1d563d095b08602dd5eed60d485a5e0` |
22
+ | KDE | [Plasma Workspace](https://github.com/KDE/plasma-workspace) | `df3bba49a2c0657d4165f6511b1771138ccc75ea` |
23
+ | Xfce | [Xfce Settings](https://github.com/xfce-mirror/xfce4-settings) | `488c919d23c979c9b28abd192c05b9b2310fa732` |
24
+
25
+ Credit the GNOME, KDE, and Xfce contributors for this search metadata. [KDE's KCM documentation](https://develop.kde.org/docs/features/configuration/kcm/) explains the authored search-keyword field. [Navigation data provenance](navigation-data.md) records extraction, immutable provider/product lineage, original source paths, hashes, license evidence, and shared-keyword positive destinations. An absent keyword is not a negative relevance judgment.
26
+
27
+ Original metadata and notices remain in [data/navigation/raw](../data/navigation/raw/). GNOME/Xfce repository GPL license texts and KDE's mixed per-file GPL/LGPL and other license evidence are retained; missing file-level SPDX assignments are explicitly recorded. These materials are not relicensed under this project's MIT code license.
28
+
29
+ ## Open English WordNet
30
+
31
+ Offline positive-pair experiments credit the **Open English WordNet team** and **Princeton University WordNet**. The source is [Open English WordNet, 2025 edition](https://github.com/globalwordnet/english-wordnet/releases/tag/2025-edition); the recorded repository revision is `dc343f2683279ecbb13fab4e2fd778d7b162d287`. The downloaded XML archive has SHA-256 `9ca6d1dcb75f822fdd66617f7d9da48142ace38dd544d6ad5e2feca1674ad3fe`.
32
+
33
+ Open English WordNet's additions use **CC BY 4.0**, while the underlying Princeton material retains its WordNet license and attribution. Both are preserved in [LICENSE.md](../data/lexicon/LICENSE.md) and [WNDB_License.txt](../data/lexicon/WNDB_License.txt). The [extraction manifest](../data/lexicon/manifest.json) records the source, hashes, train-only seed selection, filters, and 504 derived positive pairs. No negative judgments are inferred from missing word relationships. This is offline experimental supervision; no WordNet dictionary or lookup is shipped to the browser.
docs/data-sources.md ADDED
@@ -0,0 +1,45 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Expanded training data provenance
2
+
3
+ This dataset adapts licensed intent-classification examples and software-setting documentation into local retrieval menus. The labels and menu relevance are **mechanically derived, not human-reviewed retrieval judgments**. No source establishes reliable generalization to arbitrary software interfaces.
4
+
5
+ ## Sources and attribution
6
+
7
+ - **CLINC150**, Stefan Larson and collaborators, [An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction](https://aclanthology.org/D19-1131/) (EMNLP-IJCNLP 2019). [Primary repository](https://github.com/clinc/oos-eval/tree/828f8093932c8fe6ca7936c3d2e52903b1c523de), revision `828f8093932c8fe6ca7936c3d2e52903b1c523de`. Licensed **CC BY 3.0**, preserved in `data/external/clinc150/LICENSE`; original README and attribution preserved alongside. Changes: selected interface-adjacent intent families, humanized class IDs, generated distractor menus, and removed-positive negatives.
8
+ - **BANKING77**, Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, Ivan Vulić / PolyAI, [Efficient Intent Detection with Dual Sentence Encoders](https://aclanthology.org/2020.nlp4convai-1.5/) (2020). [Primary repository](https://github.com/PolyAI-LDN/task-specific-datasets/tree/57ec275d8078af65b7731c2a98be812d844a6d6b), revision `57ec275d8078af65b7731c2a98be812d844a6d6b`. Licensed **CC BY 4.0**, preserved in `data/external/banking77/LICENSE`; original README preserved alongside. Changes: capped source-train records, grouped intents, humanized class IDs, and generated menus/negatives.
9
+ - **Visual Studio Code**, Microsoft Corporation. [Primary repository](https://github.com/microsoft/vscode/tree/8e35945bae3f2b0b3d0276963281180f1ce10cb0), revision `8e35945bae3f2b0b3d0276963281180f1ce10cb0`. **MIT**, preserved in `data/ui-settings/LICENSE-vscode.txt`. Root extraction records 460 description/setting-ID pairs and source locations in `data/ui-settings/pairs.json` and `source.json`. Documentation descriptions are weak query supervision, not naturally collected search queries. Changes: cleaned original descriptions and mechanically humanized setting IDs, whole-namespace splits, deduplication, menus, negatives. Each emitted row includes the pinned source URL and line.
10
+ - Original `data/train.jsonl` synthetic examples only. They retain their `synthetic-unreviewed` status and source identity. Original `data/dev.jsonl` is separately reserved for regression evaluation; preparation never opens original dev or test.
11
+
12
+ Third-party dataset material retains its source license and attribution, independently of this repository's code license. Neither source authors nor Microsoft endorse this derived dataset or model. Source README files retain the full recommended citations.
13
+
14
+ ## Reproduction
15
+
16
+ ```sh
17
+ python -m training.prepare_external --download
18
+ python -m training.prepare_external
19
+ ```
20
+
21
+ The first command fetches only pinned public source URLs and saves complete license texts. CLINC distributes train/val/test in one transport file: the JSON container is decoded, but `test` and `oos_test` members are never accessed, exported, sampled, or scored. BANKING77 `test.csv` is not downloaded. Extracted CLINC train, val, and OOS train/val members are preserved; OOS examples are currently unused. Full upstream transport hashes are recorded by `--download`. Source asset hashes, output counts and SHA-256 values appear in `data/expanded/manifest.json` and `data/external/download-manifest.json`.
22
+
23
+ The expanded adapter consumes the checked-in VSCode pairs; `--download` refreshes CLINC/BANKING assets only. To independently re-extract the VSCode pairs after installing the pinned JavaScript dependencies:
24
+
25
+ ```sh
26
+ mkdir -p /tmp/gpu-search-vscode-source
27
+ curl --fail --location https://codeload.github.com/microsoft/vscode/tar.gz/8e35945bae3f2b0b3d0276963281180f1ce10cb0 --output /tmp/gpu-search-vscode-source/source.tar.gz
28
+ tar -xzf /tmp/gpu-search-vscode-source/source.tar.gz -C /tmp/gpu-search-vscode-source
29
+ node --import tsx training/prepare_settings.ts /tmp/gpu-search-vscode-source/vscode-8e35945bae3f2b0b3d0276963281180f1ce10cb0
30
+ python -m training.prepare_external
31
+ ```
32
+
33
+ This inspects TypeScript ASTs without executing upstream application code. Source locations, the upstream MIT license, and the extraction script hash are preserved alongside the pairs. Retain the dataset attribution and license files with redistributed training material and checkpoints; the repository's MIT code license does not replace the CLINC/BANKING source licenses.
34
+
35
+ ## Splits and derivation
36
+
37
+ Intent-family groups are declared in `training/prepare_external.py` before partition assignment. CLINC music/location and BANKING identity families are unseen development targets; CLINC employment and BANKING virtual-card families are unseen calibration targets. None of those intent labels appears in any training menu. Held CLINC families use source validation only, reserving their source-train wording unused. Known CLINC families keep source train exclusively in training and divide official validation wording between dev/calibration. BANKING77 has no official validation partition: source train is deterministically split and capped at 40/10/10 examples per known intent for train/dev/calibration, with 20 source-train examples per held intent. Its official test remains reserved.
38
+
39
+ VSCode setting namespaces are assigned wholly to train/dev/calibration using a fixed SHA-256 namespace rule. No setting ID or namespace occurs in two splits, and candidate menus only draw from the same split's settings. Exact normalized duplicate queries are globally removed with training taking precedence, then dev, then calibration. VSCode cross-split near duplicates use a frozen token-set Jaccard threshold of 0.85 (minimum four tokens), checked for target labels and queries; queries are also checked against other-source examples. This conservative lexical audit does not prove absence of paraphrase leakage.
40
+
41
+ Menus contain ten label-only candidates for new source rows. Distractors favor shared intent/label tokens and fill deterministically. Source class inequality is treated as a weak negative, though neighboring intents can be jointly relevant. Approximately one tenth of training and one quarter of evaluation examples remove the source-positive candidate entirely and are marked `removed-positive-uncertain`. They model unavailable destinations but are not reviewed proof that every remaining option is irrelevant. Existing original synthetic menus retain their six-item form.
42
+
43
+ No alias or context text is added. CLINC food/travel/chitchat and unrelated intents are excluded, BANKING training is capped, and real VSCode descriptions add a software-domain transfer slice. Banking and voice-assistant queries still dominate the count: macro source/slice reporting and the separately reserved original UI regression set are required. Do not report aggregate auxiliary improvements as a passed software-search release gate.
44
+
45
+ `data/expanded/split-audit.json` records duplicate drops, zero normalized query overlap, and the held-family candidate exclusion check. `manifest.json` records exact final counts, source counts, no-match counts, slice counts, intent-family definitions and hashes. Development and calibration have distinct queries and separate unseen-intent families, but known-intent labels are intentionally shared. These are development resources; no final test has been opened.
docs/decisions.md ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Implementation decisions
2
+
3
+ - 2026-09-12: The attached v1.1 specification is the design input. The user separately requested public GitHub publication and a Vercel Hobby demo. No npm publication, paid teacher calls, or paid hosting features are authorized or used.
4
+ - Preserve the normative API names. Until a semantic model meets release gates, the main entry returns `backend: 'lexical', degraded: true` when semantics are requested. The lexical entry/default demo intentionally disables semantics and reports no degradation. Context is validated but has no lexical effect.
5
+ - Keep normalized label equality above alias equality. Lexical subclass precedence is independent of score. Ordered token prefix scores are capped at one: token separators can otherwise make the numerator larger than the original field length (for example `foo-bar` and `fooBar`).
6
+ - Freeze feature serialization before synthetic generation; see `packages/model/feature-spec.json`. ASCII case rules and explicit Unicode whitespace prevent Python casefold divergence.
7
+ - Synthetic data is an experiment, not reviewed relevance ground truth. A synthetic positive result cannot satisfy the required unseen-product quality gate. The sealed synthetic set is not used to select a model; independent reviewed final evaluation remains unavailable.
8
+ - Do not build or advertise a GPU inference backend before semantic feasibility. `auto` remains lexical in this baseline release. The project name describes the research direction, not a claim of shipped acceleration.
9
+ - Ship a static Vite demo without telemetry, remote fonts, server functions, databases, or inference endpoints. All query processing remains in the browser. Run local browser validation before deployment.
10
+
11
+ ## Interactive model demo revision
12
+
13
+ The user explicitly requested a real model-backed demo despite the failed release gate. The demo now evaluates the existing experimental pooled16 int8 weights on CPU and displays raw cosine rankings alongside the unchanged lexical library. It does not impose a made-up validated cutoff or present model scores as confidence. Initial candidates have no aliases or contexts. Arbitrary edited candidates use the same encoder and composition, and model load failure is shown as an error rather than replaced with baseline results. The stable library still defaults to explicit lexical degradation when semantics are requested. No WebGPU claim is made.
14
+
15
+ ## 2026-09-13: expanded-data experiment
16
+
17
+ Freeze source revisions and split hashes before training. Add CLINC150/BANKING77 intents and VS Code setting descriptions with explicit source licenses and weak-label provenance. Keep original dev out of training/checkpoint selection. Compare three seeds against repeated old-data controls with equal optimizer updates at epoch1/10. Calibrate abstention on separate calibration rows. Among development-selected checkpoints, choose seed17 (epoch10) because it is the only one meeting the pre-existing 5% no-match gate; reject higher-nDCG checkpoints that exceed that budget. Preserve pilot artifacts, deploy the improved experimental int8 artifact, and retain raw-score/limited-generalization disclosure. Quality regression thresholds freeze measured development floors; they are not independent test evidence.
18
+ # Optimization follow-up: retain the live model
19
+
20
+ After committing baseline `c5a0a11`, 39 new training runs and 90 fixed-model calibration variants were evaluated within the existing size target. Ordered word-pair hashing produced a 32 KiB research candidate with development semantic nDCG 0.5225 versus 0.4644. The combined no-match objective offered no clear additional benefit; capacity, distillation and learned calibration alternatives were rejected.
21
+
22
+ Selection was frozen before opening the reserved synthetic test. Both candidate and live baseline had zero calibrated semantic coverage there; raw overall nDCG fell from 0.5452 to 0.5345. Keep the current live weights. Preserve the candidate, complete results, isolated runtime and regressions for research, without calling it an improved release. The consumed synthetic test is no longer fresh evaluation data. Details: [optimization experiments](optimization-experiments.md).
docs/deployment.md ADDED
@@ -0,0 +1,11 @@
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Deployment
2
+
3
+ Public demo: https://gpu-search.vercel.app
4
+
5
+ GitHub: https://github.com/maxffarrell/gpu-search
6
+
7
+ Published 2026-09-12 to Vercel project `gpu-search` in `maxffarrells-projects`. The Vercel team API reported billing plan `hobby` before project creation. No paid features or plan changes were enabled. The demo is static; it requires no secrets, server functions, database, analytics, or inference service.
8
+
9
+ The initial CLI production deployment `dpl_BYWZ9EjyZYaZL8aZKLegvPk3G8K5` reached READY and was aliased to the public URL. Vercel connected the GitHub repository for subsequent builds. `vercel.json` sets the frozen pnpm install, Vite build, output directory and restrictive security headers. `.vercelignore` excludes research data and float checkpoints. The demo revision includes the experimental manifest and 32KiB int8 weights as explicit assets for real local CPU inference.
10
+
11
+ Run `vercel deploy --prod --scope maxffarrells-projects` to republish when explicitly intended. Runtime behavior was validated on the local production build; deployment readiness and the alias were verified from Vercel's build result.
docs/expanded-evaluation.md ADDED
@@ -0,0 +1,57 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Expanded training evaluation
2
+
3
+ The new 16-dimensional pooled model improves the measured weakly labeled development task. Its quantized artifact is `packages/model/candidate`; the original deployed pilot artifact remains unchanged at `packages/model/experimental`. This is an experimental model, not an independently validated release.
4
+
5
+ The expanded corpus has 7,961 training rows, 1,397 calibration rows, and 1,472 development rows. Sources are licensed CLINC150 and BANKING77 intent datasets, VS Code settings descriptions, and the original 120 synthetic training records. Source classes and descriptions were converted into retrieval judgments; none of these retrieval judgments were human reviewed. See [data sources](data-sources.md) for licenses, transformations, split audits, and limitations. No sealed test was used.
6
+
7
+ The calibration partition alone selects each system's global cosine cutoff. Development selects checkpoints and compares systems. Of the three expanded runs, seed 17 at epoch 10 meets the preexisting 5% development no-match cap; the higher-nDCG seed 29 and seed 43 candidates exceed it. Selection therefore uses the previously registered constraint rather than silently relaxing it. Original pilot development data is evaluated only after selection as a transfer regression check.
8
+
9
+ An independent NumPy checkpoint audit also compares training-data controls at the same saved epoch. At epoch 10, semantic nDCG@5 for expanded versus original-data-only training is 0.4644 versus 0.0292 (seed 17), 0.5154 versus 0.0447 (seed 29), and 0.4948 versus 0.0324 (seed 43). Seed 17 improves from 0.2741 at epoch 1 to 0.4644 at epoch 10. This supports a benefit from the expanded supervision on this evaluation distribution; it does not make all expanded runs pass abstention. All selected, epoch-1, and epoch-10 checkpoints are recorded in `eval/expanded-checkpoint-audit.json`.
10
+
11
+ ## Actual int8 artifact comparison
12
+
13
+ The table reports nDCG@5 on the 1,118 development queries having a relevant candidate but no relevant eligible lexical match. Coverage is the fraction returning any result. No-match rate is the fraction of all 352 no-match development menus returning a semantic item. Cosines are not confidence probabilities.
14
+
15
+ | System | Semantic nDCG@5 | No-match semantic rate |
16
+ | --- | ---: | ---: |
17
+ | Lexical | 0.0000 | 0.00% |
18
+ | Lexical plus all train-derived aliases | 0.0304 | 0.00% |
19
+ | Train-fitted word/character TF-IDF | 0.1232 | 3.98% |
20
+ | Nearest train paraphrase TF-IDF | 0.3852 | 7.10% |
21
+ | Original int8, calibrated | 0.0277 | 3.12% |
22
+ | New int8, calibrated | **0.4644** | **4.83%** |
23
+
24
+ The new artifact improves over the strongest baseline by **0.0792** nDCG@5, with paired query-bootstrap 95% interval **[0.0495, 0.1089]**. Improvement over the original artifact is **0.4366**, interval **[0.4073, 0.4655]**. These are descriptive development intervals after checkpoint selection, not independent generalization evidence; source-related queries are correlated. The strongest paraphrase baseline's calibration cutoff transfers to a 7.10% development false-positive rate, which exceeds the target and is reported rather than retuned on development.
25
+
26
+ Every system sees identical host candidates, aliases, and context. The complete train alias baseline can use more than the runtime's eight-alias limit, intentionally making this a stronger offline comparator. The nearest-paraphrase baseline scores each candidate against its complete collection of positive training utterances, not only the candidate label. Word/character vocabulary and IDF are fit exclusively on training texts. No development wording is inserted into either baseline.
27
+
28
+ Without abstention, new raw semantic nDCG@5 is 0.8491 versus 0.4537 for the old artifact. Raw scoring returns a score for every nonzero embedding, including unrelated candidates, and consequently returns something on 100% of no-match menus. It is useful for inspection, not a calibrated search policy. Calibrated semantic-query coverage is 49.55%; abstention removes many relevant results as well as unrelated ones.
29
+
30
+ ## Seen intents and transfer
31
+
32
+ | Development slice | New calibrated nDCG@5 | New raw nDCG@5 | Nearest train paraphrase nDCG@5 |
33
+ | --- | ---: | ---: | ---: |
34
+ | Known intent, different wording | 0.5479 | 0.9153 | 0.4721 |
35
+ | Unseen intent family | 0.1673 | 0.5480 | 0.0698 |
36
+ | VS Code settings, unseen namespace | 0.1733 | 0.7915 | 0.1020 |
37
+
38
+ Most gains concern wording for known intents. Unseen-family calibrated coverage remains weak. In particular, direct train-fitted TF-IDF obtains 0.5174 on the unseen settings namespace slice, beating the new calibrated model's 0.1733 there. The global improvement must not be presented as superiority on every software-interface domain.
39
+
40
+ The original pilot development set was never added to training or checkpoint selection. Its combined calibrated overall nDCG@5 improves from 0.4000 to 0.4167; raw semantic-query nDCG improves from 0.5221 to 0.5256. Raw overall nDCG changes from 0.6548 to 0.6460, a 0.0088 regression within the preexisting 0.01 allowance. Both models have zero calibrated no-match false positives on that small transfer set. This small old synthetic set is a regression probe, not a generalization benchmark.
41
+
42
+ ## Gates and regression checks
43
+
44
+ The expanded source-derived development comparison passes the provisional numerical gates registered in `eval/expanded-gates.json`. The reviewed-final-test gate remains **not passed**. No final test, broad human judgment study, general embedding teacher baseline, or multilingual evaluation was executed.
45
+
46
+ Actual float and int8 evaluation shows zero nDCG loss with independently calibrated cutoffs. At the frozen float cutoff, quantization loses 0.00179 semantic nDCG@5, also within the unchanged 0.01 allowance. CPU/NumPy numerical parity is tested separately from retrieval quality; WebGPU is not implemented.
47
+
48
+ `training.test_quality` checks metric math with multiple relevant answers and grade-1 judgments, hard tier precedence against adversarial semantic scores, cutoff tie behavior, train-only baseline fit, invalid/zero embedding abstention, and actual selected artifact quality against relevance judgments. The frozen regression floor is semantic nDCG 0.45, semantic coverage 0.48, and no-match rate at most 0.05. It also enforces the original transfer regression allowance and proves zeroed weights fail the semantic floor. Dataset SHA-256s are fixed in `eval/expanded-quality-floors.json`; missing data or artifacts fail, rather than silently skipping. This suite uses selection-used development data as a repeatable regression corpus, not as new independent evidence.
49
+
50
+ Reproduce from the repository root:
51
+
52
+ ```sh
53
+ uv run python -m training.evaluate_expanded --old packages/model/experimental --new packages/model/candidate
54
+ uv run python -m unittest training.test_quality -v
55
+ ```
56
+
57
+ The evaluation command writes `eval/expanded-evaluation.json`, `eval/expanded-transfer.json`, and `eval/expanded-quantization.json`. Reports retain per-query top-five indices and aggregate metrics, not large score matrices. Floor thresholds are intentionally not regenerated by the evaluator. Original pilot reports are preserved.
docs/experiment.md ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Reproducing the synthetic feasibility experiment
2
+
3
+ This experiment deliberately reports a failed feasibility gate. It does not supply a production semantic model.
4
+
5
+ ```sh
6
+ uv sync --frozen
7
+ uv run python -m training.prepare --config configs/data.yaml
8
+ uv run python -m training.train --config configs/base.yaml
9
+ uv run python -m training.report
10
+ uv run python -m training.audit
11
+ uv run python -m training.evaluate --split dev --checkpoint runs/pooled-16-both-contrastive-seed17
12
+ uv run python -m unittest training.test_export
13
+ uv run python -m training.export --checkpoint runs/pooled-16-both-contrastive-seed17 --bits 8 --experimental
14
+ ```
15
+
16
+ `uv.lock` pins the Python dependencies. Training requires macOS Apple Silicon with Metal access; the original sandbox could not initialize Metal, and training succeeded with authorized direct GPU access. Feature preparation, NumPy evaluation and exporter tests work without MLX. `training.evaluate` intentionally has no final-test option; a reviewed sealed-test protocol is still required. `runs/selected` does not exist because no model qualified. Export without `--experimental` rejects instead of silently blessing a checkpoint.
17
+
18
+ The source fixtures are entirely original synthetic text under this repository's MIT license. `data/sources.json` contains provenance and the preparation-script hash. `data/leakage-audit.json` records menu-family splits, literal-query overlaps and encoder truncation counts. No paid API or external dataset was used.
19
+
20
+ ## Actual development results
21
+
22
+ The 80 development records contain 24 exact/typo records, 36 semantic queries and 20 no-match records. Overall nDCG below averages the 60 relevant-query records; no-match behavior is reported separately. There are no grade-1-only cases. Semantic-only means no judged relevant candidate has an eligible lexical match. Gain is `2^grade - 1`, MRR and recall use grade at least 2. Because each record has one positive, recall equals hit rate in this particular dataset.
23
+
24
+ | System | Semantic nDCG@5 after cutoff | Overall nDCG@5 | No-match any result |
25
+ | --- | ---: | ---: | ---: |
26
+ | Lexical | 0.0000 | 0.4000 | 0% |
27
+ | Lexical + training aliases | 0.0000 | 0.4000 | 0% |
28
+ | Character TF-IDF + same lexical tiers | 0.1462 | 0.4877 | 5% |
29
+ | Pooled 16d contrastive seed 17 | 0.0833 | 0.4500 | 5% |
30
+ | Pooled 16d contrastive seed 29 | 0.0000 | 0.4000 | 5% |
31
+ | Pooled 16d contrastive seed 43 | 0.0278 | 0.4167 | 5% |
32
+ | Projected 16d contrastive seed 17 | 0.0278 | 0.4167 | 5% |
33
+ | Projected 16d contrastive seed 29 | 0.0278 | 0.4167 | 5% |
34
+ | Projected 16d contrastive seed 43 | 0.0278 | 0.4167 | 5% |
35
+
36
+ All 18 individual ablations, cutoffs, coverage, recall, MRR, raw per-query ranks/scores and bootstrap intervals are in `eval/feasibility.json` and each `runs/*/dev-evaluation.json`. The pooled seed-17 semantic improvement interval is [0.0000, 0.1944], so its point estimate does not pass the interval gate. No seed passes. Metrics before abstention look much better while returning something for every no-match query; that is precisely why both cutoff behavior and coverage are reported.
37
+
38
+ Cutoffs were selected on development and are reused for these diagnostics. Bootstrap uses 10,000 paired query samples, seed 20260912, and is descriptive after development selection; it is not a final held-out confidence claim. The global cutoff allows at most 5% of development no-match queries to receive any semantic item. Results remain dominated by tiny sample size and invented judgments.
39
+
40
+ The alias map comes from training only. Host aliases/context are identically absent for every system. TF-IDF fits document frequencies on unique training labels, with smoothed unseen-gram document frequency zero, and uses the same fixed lexical tiers. It is a lexical n-gram similarity diagnostic, not a general embedding model. The default library uses neither TF-IDF nor an experimental encoder. The website explicitly runs the seed-17 experimental encoder without a cutoff so users can inspect its actual scores; this does not change the failed gate.
41
+
42
+ The reserved test records were never scored. The next meaningful step is human-reviewed product menus and disjoint relevant/no-match queries, with a stronger reviewed train/development alias baseline, followed by preregistered evaluation. More shader work cannot repair the current evidence gap.
docs/feature-experiments.md ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Feature experiments
2
+
3
+ The selected research candidate is `features-bigram25-seed29-e10`, epoch 9: two 1024×16 tables, 32,768 int8 weight bytes, trained with MLX 0.31.1 on Metal. It adds ordered adjacent token bigrams to the existing word table. No extra embedding table, vocabulary or inference framework is required. The byte figure covers weights, not the complete browser download.
4
+
5
+ The artifact is frozen at `packages/model/experiments/features-bigram25-seed29-e10/`, payload SHA-256 `166f105df6a3f4e0941828b1130d06671c2b5bbd759471811f607a2fa7ae30f8`. The standard experimental manifest deliberately retains `validatedSemanticCutoff: null`; its research calibration result is not a final release approval. `feature-contract.json` specifies serialization and pooling. Sixteen numerical fixtures cover multibyte text, emoji, normalization, punctuation, repeated occurrences, 32-token bounds and 64-scalar token bounds. The portable NumPy reference is `training/features_bigram.py`; browser implementations must explicitly recognize `gpu-search-features-bigram25-v1`.
6
+
7
+ ## Feature comparison
8
+
9
+ The ten-epoch screen uses the same pooled 16 architecture, training rows, optimizer settings and per-seed shuffled update sequence. It compares original features, 75% word/25% character weighting, and 25% bigram weighting inside the word family. Calibration alone sets the cutoff. Development selects epochs that satisfy the 5% no-match constraint, then maximizes semantic nDCG@5 and relevant coverage. Original pilot development is excluded from selection.
10
+
11
+ | Arm | Seed | Selected epoch | Int8 semantic nDCG@5 | Dev no-match |
12
+ |---|---:|---:|---:|---:|
13
+ | Original features |17|10|.46526|4.83%|
14
+ | Original features |29|1|.27704|4.55%|
15
+ | Original features |43|None feasible|—|Above5%|
16
+ |75% word weighting|17|None feasible|—|Above5%|
17
+ | Bigram25 |17|8|.50253|4.83%|
18
+ | Bigram25 |29|9|.52254|4.83%|
19
+ | Bigram25 |43|8|.52161|3.69%|
20
+
21
+ The selected-checkpoint comparison favors bigrams, but the seed 29 control only satisfies the no-match condition at epoch 1. That large difference combines feature quality with calibration/optimization behavior. At the equal ten-epoch budget, float semantic improvements are+.0625,+.0050,+.0206 for seeds 17/29/43; paired bootstrap intervals include zero for the latter two. Do not present this as a significant equal-update gain for every seed. Independent NumPy results and per-source/held-family slices are in `eval/feature-experiments-independent.json`.
22
+
23
+ The selected seed 29 has expanded-dev semantic nDCG .522537, relevant coverage and source slices recorded in the report, and 4.83% global no-match rate. Its VSCode slice nDCG is .290658. Aggregate safety can hide weaker unseen-domain slices; the final evaluator must report those separately.
24
+
25
+ ## Post-selection transfer
26
+
27
+ `eval/feature-experiments-selection.json` froze seed 29 before the one original-pilot transfer check. Raw pilot nDCG is .684893 versus .654799 for the original pilot and the supplied deployed reference .6460, passing the .01 regression guards. The paired interval versus the original pilot is [-.0731,.1346], so significant improvement is not established. Raw scores return results for no-match queries; the calibrated combined system instead has .43333 nDCG, .45 relevant coverage and .05 no-match rate on that pilot slice. No tuning followed the transfer result.
28
+
29
+ ## No-match objective interaction
30
+
31
+ A separate bounded experiment combines the bigram features with the independently tested no-match training objective. It uses all-zero TRAIN judgments only, seeds 17/29/43 and at most 30 epochs. Its int8 development semantic nDCGs are .531107, .447057, .526181, with global no-match rates 4.83%, 3.98%, 4.55%. The best combination gains only .00857 over the selected simpler model; the paired interval [-.01614,.03274] includes zero. It has a worse weakest seed and lower selected VSCode nDCG(.271versus .291). The combination is therefore rejected for this candidate; its original-pilot transfer was not opened for tuning.
32
+
33
+ Those outputs remain separate in `runs/features-bigram-no-match/`, `packages/model/experiments/features-bigram-no-match-*`, and `eval/feature-experiments-bigram-no-match.json`. The training wrapper injects a feature cache only inside its own process, leaving baseline modules untouched. Its fixtures use the bigram encoder, never the original feature encoder.
34
+
35
+ ## Verification and reproduction
36
+
37
+ The independent MLX-versus-NumPy pooling audit over 139 texts found maximum vector/cosine error below 4.18e-7, within the 1e-4/2e-4 tolerances. The same development set showed no semantic nDCG loss after int8 quantization. Browser parity and complete transfer-size results are separate integration evidence.
38
+
39
+ ```sh
40
+ uv run python -m training.experiment_features --epochs 10
41
+ uv run python -m training.experiment_features --arms control bigram25 --seeds 29 43 --epochs 10
42
+ uv run python eval/feature-experiments-reference .py
43
+ uv run python eval/feature-experiments-parity.py
44
+ uv run python -m training.experiment_features_objective
45
+ ```
46
+
47
+ Training refuses to overwrite existing run directories; use a fresh experiment checkout to reproduce it. Source dataset hashes are saved with runs. These results use weakly derived retrieval judgments and development-selected checkpoints. Neither the source datasets nor these experiments establish reviewed final-test generalization.
docs/model-card.md ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model card: expanded experimental model
2
+
3
+ Latest decision: the experimental demo now uses `navigation-align0p5-seed29-word085`, a 15% reduction to the word-table dequantization scale of the selected navigation checkpoint. The int8 payload remains **32,768 bytes** with SHA-256 `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`; the manifest SHA-256 is `019b113275824658b6ac18aa40afb7a28187ce5855620f5be393bdeda2e24388`. [Typo-weight experiments](typo-weight-experiments.md) document the 3-point reserved synthetic spelling gain, rejected retraining, and still-failing raw `profle` case. No production semantic release is claimed.
4
+
5
+ The previous demo artifact remains under `packages/model/candidate/`, the prior word-pair research artifact under `packages/model/candidate-v2/`, and the original pilot under `packages/model/experimental/`. Both the original synthetic test and new GNOME provider evaluation are now consumed; neither is fresh evidence for future selection. The sections below retain historical pilot and expanded-data evidence.
6
+
7
+ Training records: 7,961 total / 7,165 positive supervised, development 1,472, separate calibration 1,397. Sources are CLINC150, BANKING77, MIT VS Code descriptions and the original training fixtures. Source intent labels and documentation are adapted into weak menu relevance, not newly human-reviewed search judgments. Original dev is used only for post-selection transfer; upstream tests are unused.
8
+
9
+ On expanded development, int8 semantic nDCG@5 is 0.4644, versus old-model 0.0277 and nearest-training-paraphrase TF-IDF 0.3852. The +0.0792 baseline delta has paired bootstrap 95% interval [0.0495, 0.1089], descriptive after checkpoint selection. No-match false positives are 4.83%; semantic relevant-query coverage is 49.55%. Measured int8 nDCG loss is zero with separately calibrated cutoffs and 0.00179 at the frozen float cutoff, within the 0.01 allowance. Provisional numerical development gates pass; independent reviewed final evaluation remains absent.
10
+
11
+ The model is much stronger on new wording of seen intents than on unseen domains. Original pilot transfer semantic nDCG@5 remains only 0.0278 with the expanded calibration cutoff. Raw demo rankings deliberately apply no cutoff and can return irrelevant answers. The stable lexical engine remains responsible for guaranteed exact/typo ordering; there is no WebGPU implementation or generalization guarantee.
12
+
13
+ Only seed17 met the development no-match limit among the three selected new checkpoints. Other seeds with higher retrieval scores were rejected for excessive false positives. Matched-update old-data controls and fixed epoch comparisons separate extra examples from extra optimization. See [expanded evaluation](expanded-evaluation.md) and [source/license provenance](data-sources.md). More training did help this larger dataset through epoch10, while later selected checkpoints/other seeds did not establish better acceptable coverage.
14
+
15
+ ## Original pilot history (superseded findings, retained for reproducibility)
16
+
17
+ The library remains lexical with optional developer aliases. **No learned model passed the semantic release gate.** The first model demo explicitly ran the pilot MLX-trained pooled16 int8 weights through a custom TypeScript CPU encoder, exposing raw cosine rankings alongside the lexical engine. The default package does not load those weights. The model panel has no validated cutoff and may rank unrelated candidates.
18
+
19
+ ## Intended scope
20
+
21
+ Short English software-navigation queries and candidate menus. Character hashing borrows the subword idea from fastText; shared pooled feature embeddings are inspired by StarSpace. This is original code, not an exact reproduction of either implementation. A bag of features cannot encode arbitrary word order or reliably distinguish action opposites. Cosine similarity is not a probability.
22
+
23
+ ## Data
24
+
25
+ Original agent-authored synthetic fixtures: 120 training records, 80 development records, and 40 reserved synthetic test records. There are six menu families, six candidates per menu, 18 training concepts, 12 development concepts, and six reserved concepts. All judgments are **synthetic and unreviewed**. No external corpus or teacher was used. The actual counts are far below the proposed data targets.
26
+
27
+ The menu family is the split group, and all associated label, typo, and paraphrase examples remain together. Normalized queries do not overlap across splits. Shared ordinary English words still occur; zero literal overlap is not proof of semantic independence. The reserved synthetic test set was generated but not scored, is not genuinely sealed, and cannot replace a separately reviewed held-out set.
28
+
29
+ Every positive record has exactly one grade-3 destination; all remaining candidates are grade 0. No examples establish multi-positive relevance, grade-1 judgments, real product ambiguity, or context handling. No-match phrases are synthetic and mostly easy. These limitations prevent a release-quality conclusion even if the numeric target had been reached.
30
+
31
+ ## Training
32
+
33
+ Actual runs used Python 3.12.14, MLX 0.31.1 and Metal on an Apple M4 Max, macOS as recorded in `eval/feasibility.json`. Seeds: 17, 29, 43. AdamW, learning rate 0.003, weight decay 0.0001, temperature 0.15, global gradient clipping 1.0, maximum 100 epochs and development patience 10. The complete positive training set has 90 records, so the actual full batch is 90 (configuration upper bound 128).
34
+
35
+ Eighteen checkpoints cover pooled and projected 16-dimensional encoders, pooled margin loss with fixed 0.2 margin, a 24-dimensional projected variant, word-only pooled, and character-only pooled models. The margin was a fixed pilot choice; no margin sweep was run. Every checkpoint has raw history, materialized elapsed training time, settings and weights under `runs/`. Early stopping used pre-cutoff semantic development nDCG@5. The dataset has label-only candidates, so training does not exercise aliases/context composition.
36
+
37
+ ## Findings and deployment decision
38
+
39
+ At development cutoffs allowing at most one semantic false positive among 20 no-match records, learned semantic nDCG@5 ranged from 0 to 0.0833. All paired bootstrap 95% improvement intervals against lexical plus training aliases included zero. The character TF-IDF diagnostic scored 0.1462 on the same semantic slice. All 18 provisional numeric gates failed; there is also no reviewed final-test evidence.
40
+
41
+ Lexical plus the training alias map scored 0.4000 overall and 0 on the 36 semantic-only development queries. None of the train alias map's labels occur in development, so this is a weak test of a curated dictionary's usefulness. The model does not get development aliases either. This pilot suggests that the tiny synthetic data is inadequate; it does not demonstrate an architecture impossibility.
42
+
43
+ No release checkpoint, validated cosine cutoff, or trained WebGPU runtime was selected. Distillation, quantization optimization, and shader optimization stopped at the feasibility gate. The int8 artifact under `packages/model/experimental/` is an exporter smoke-test artifact from seed 17, with an explicit null validated cutoff. It is not a deployment recommendation. No int6 quality claim is made.
44
+
45
+ ## Missing evidence
46
+
47
+ Human-reviewed data and final evaluation; a pinned general embedding reference; real product-source licensing/curation; teacher distillation; alias/context, action-opposite and ambiguity evaluation; selected-model MLX-to-JavaScript-to-WebGPU numeric parity (the experimental CPU encoder is checked against exported NumPy fixtures); semantic transfer bytes and browser latency. Empty and long feature inputs have shared fixtures, but that does not establish retrieval quality for them.
docs/model-readiness-audit.md ADDED
@@ -0,0 +1,101 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Model readiness audit
2
+
3
+ The evidence does not support a production-ready or best-in-class learned search claim. The current lexical engine is useful, and the learned models are small, but the measured task mostly rewards recognizing long utterances for previously seen intent labels. The intended product is short-query search over arbitrary software menus. More optimization on the existing development set cannot establish that capability.
4
+
5
+ This audit is read-only. No dataset, weights, cutoff, or runtime changed. The previously opened 40-row synthetic test is now explicitly a consumed regression set. Its examples were inspected to diagnose failures, so it cannot be reused as independent final evidence. The primary data are in `eval/readiness-audit.json`; the Node CPU diagnostic is in `eval/readiness-cpu.json`.
6
+
7
+ Follow-up: the caching gap identified in section 4 has since been addressed by an experimental prepared-model index. [Prepared-index documentation](prepared-model-index.md) records its API, browser measurements, and unchanged weights. The audit measurements below describe the original uncached reference path; the quality findings and missing integrated semantic-search evidence remain applicable.
8
+
9
+ ## 1. The failure is representation shift as well as abstention
10
+
11
+ | Measurement | Expanded development | Consumed software-menu regression |
12
+ | --- | ---: | ---: |
13
+ | Median query length | 9 tokens | 2 tokens |
14
+ | Candidate count | 10 | 6 |
15
+ | Positive rows whose target label appeared in training | 876 / 1,120 (78.2%) | 0 / 30 |
16
+ | Query-token occurrences absent from training vocabulary | 804 / 14,469 (5.6%) | 42 / 80 (52.5%) |
17
+ | Bigram model median relevant cosine, semantic queries | 0.6751 | -0.0336 |
18
+ | Bigram raw top-one correct, semantic queries | 828 / 1,118 | 0 / 18 |
19
+
20
+ The rejected bigram model's frozen expanded-calibration cutoff is 0.6562. Its maximum relevant score on the consumed regression semantic slice is only 0.3286; every relevant candidate is below cutoff. The current deployed model similarly places every relevant regression semantic candidate below its 0.7004 cutoff, with raw top-one correctness only 2/18.
21
+
22
+ Lowering the threshold cannot repair incorrect raw ordering: the rejected model is wrong at rank one for all 18 semantic queries. This is not evidence that no semantic signal exists—lower ranks give nonzero nDCG—but it disproves a cutoff-only explanation. Those 18 correlated examples cover one publishing menu with six target labels, so they also cannot quantify broad product performance. The actionable response is independently judged short-query menu data across several software products, not threshold tuning against these examples.
23
+
24
+ The same symptom appears in the original 80-row development transfer probe: short queries, zero positive target labels seen in training, and much lower relevant scores. These are direct observed distribution differences. They do not isolate which causal factor—query length, domain, unseen labels, or menu composition—dominates. A new study should cross these factors while keeping judgments fixed.
25
+
26
+ ## 2. Calibration negatives need retrieval judgment, not class deletion
27
+
28
+ Every one of the 352 calibration no-match menus and 352 expanded-development no-match menus was constructed by removing the source's single positive intent. The remaining labels were assigned zero relevance without a human retrieval review.
29
+
30
+ A conservative lexical review flag—at least two shared label tokens and token-set Jaccard at least 0.5 between the removed label and a remaining label—identifies **106/352 calibration no-match menus (30.1%)**, **77/352 development menus (21.9%)**, and 161 training menus. Examples include:
31
+
32
+ - Removed **Schedule meeting**, remaining **Meeting schedule**, query “schedule my meeting with jim at 3pm.”
33
+ - Removed **Card not working**, remaining **Virtual card not working**, query “Please help. The card won't work.”
34
+
35
+ These flags are not proof that the remaining destination is relevant. They show why source intent classification does not determine whether an interface destination is useful. Some labels encode opposite actions; others may be ambiguous without context. Treating every alternative as a trusted negative can punish sensible retrieval and force a high abstention threshold.
36
+
37
+ There are no multi-positive menus in the training, calibration, or expanded-development data. There are also no exact-label contradictions within a menu. The problem is unmeasured semantic ambiguity, not a detected duplicate-ID or literal-label bug. Review flagged menus, permit multiple positive judgments, and leave genuinely unjudged candidates out of training loss and headline quality metrics. Do not automatically relabel the flags as positive or selectively remove difficult negatives after seeing model scores.
38
+
39
+ ## 3. The byte budget creates collision pressure, but causality remains unproven
40
+
41
+ Across distinct training query/label features:
42
+
43
+ | Feature table | Distinct feature preimages | Occupied buckets | Buckets with multiple preimages | Largest bucket |
44
+ | --- | ---: | ---: | ---: | ---: |
45
+ | Words | 5,179 | 1,019 / 1,024 | 983 | 14 |
46
+ | Character grams | 15,710 | 1,024 / 1,024 | 1,024 | 27 |
47
+ | Shared words plus bigrams | 27,282 | 1,024 / 1,024 | 1,024 | 47 |
48
+
49
+ The bigram variant adds order information while forcing substantially more distinct features into the same word table. A collision does not by itself establish a retrieval error; subword pooling can tolerate collisions. These counts justify matched-byte bucket/dimension ablations and error slices for unseen tokens. They do not justify claiming that larger tables will solve the domain problem, nor increasing the 32,768-byte weight budget without authorization. Existing same-size collision experiments should be judged on independent software-menu data rather than this consumed regression set.
50
+
51
+ ## 4. A production semantic index and its latency evidence are missing
52
+
53
+ At the time of the initial audit, the stable `createIndex` was lexical and the separate experimental `Model.score` recomputed label, alias, and context embeddings for every candidate on every query. The subsequent prepared API resolves candidate caching and per-index disposal. The stable `createIndex` still reports lexical degradation when semantics are requested; there is no integrated exact/lexical/semantic result API or validated semantic cutoff. The demo continues to expose raw scores.
54
+
55
+ A direct local Node diagnostic of the current trained CPU runtime on Apple M4 Max, Node 26.8.1, macOS 27.0.0 measured:
56
+
57
+ | Candidates, each with an alias and context | Raw semantic scoring p95 |
58
+ | --- | ---: |
59
+ | 10 | 0.485 ms |
60
+ | 100 | 4.403 ms |
61
+ | 1,000 | 43.987 ms |
62
+
63
+ The same candidate list was reused, with 30 warmups and 200 measured queries per size. Power state was uncontrolled. This is **Node diagnostic evidence**, not browser proof or a full hybrid-search benchmark; it excludes final lexical ranking. Even this partial operation is about 5.5 times the proposed 8 ms target at 1,000 candidates. Cached immutable candidate vectors are an actionable improvement before GPU work. Existing browser benchmarks measure lexical search and cannot be cited as learned-runtime latency proof.
64
+
65
+ The research build fits the size target: the bigram artifact has 32,768 raw weight bytes and its complete separately compressed demo resources total 36,243 Brotli bytes. That is size evidence for the recorded research build, not a quality or live deployment guarantee. WebGPU remains unimplemented.
66
+
67
+ ## New independent software-menu evaluation, before further training
68
+
69
+ Create a new versioned source manifest and register the protocol before training on any new extraction. The already consumed synthetic test and any inspected query/label families become regression material. No convenient relabeling of those records can make them fresh.
70
+
71
+ Candidate permissively licensed source projects include [Excalidraw (MIT)](https://github.com/excalidraw/excalidraw/blob/master/LICENSE), [JupyterLab (BSD)](https://github.com/jupyterlab/jupyterlab/blob/main/LICENSE), [Apache Airflow (Apache-2.0)](https://github.com/apache/airflow), and [Ghost (MIT)](https://github.com/TryGhost/Ghost/blob/main/LICENSE). License locations were checked for this proposal; exact commit SHAs, extracted paths, notice retention, and source-specific permissions must be recorded before extraction. These are proposals, not sources already ingested by this audit. Ghost's publishing domain overlaps the consumed regression concept family, so it should be development or an explicitly overlapping slice, not the headline independent holdout.
72
+
73
+ 1. **Capture actual menus and commands.** Preserve candidate ID, displayed short label, application section, contextual description, and source path at a pinned revision. Exclude inaccessible destinations before constructing the menu. Use real competing actions from the same interface, including import/export and enable/disable, rather than randomly combining unrelated labels.
74
+ 2. **Split by product and task family before authoring queries.** Keep all wording variants, aliases, and derived menus of a task together. Reserve at least two whole application families for final evaluation. Cross-check near-duplicate task descriptions and normalized query/token overlap with every prior dataset, including the consumed synthetic regression set. Generic shared words are unavoidable; report them rather than claiming perfect conceptual independence.
75
+ 3. **Target the product's queries.** Proposed primary distribution: 70% one-to-four tokens, 25% five-to-eight, 5% nine-to-twelve. Include exact, typo, prefix/acronym, low-overlap semantic, ambiguous/multi-positive, opposite-action, and genuine no-match cases. Keep representative 10-, 50-, 250-, and 1,000-candidate settings; report menu-size calibration transfer rather than fitting it from only ten-candidate menus.
76
+ 4. **Judge every candidate for each evaluation query.** Two reviewers independently assign grades 0–3, then resolve disagreements with application context. At least 15% of relevant queries should exercise legitimate multiple answers, and include at least 200 explicitly reviewed no-match menus. An annotation generated from a source class or a language model is a proposal until reviewed. Queries should describe user intent before showing the exact destination wording where feasible.
77
+ 5. **Keep fit, calibration, development, and final roles distinct.** Fit weights and aliases only on training. Fit abstention only on calibration. Use development for the registered finite model search. Freeze artifact, feature contract, aliases, threshold, and source hashes before opening the final judgments once. A failed final set becomes regression material; another round needs a genuinely new independent set.
78
+
79
+ If no human judgment resource is available, a source-derived benchmark still helps engineering, but the production-quality gate remains unresolved. Do not disguise weak labels as a reviewed test to meet a schedule.
80
+
81
+ ## Measurable readiness gates
82
+
83
+ The existing PRD efficacy and size gates should remain unchanged: at least +0.05 semantic nDCG@5 over the strongest fair alias/lexical baseline, paired interval excluding zero, no more than 0.01 overall regression, at most 5% semantic no-match returns with coverage reported, at most 0.01 quantization loss, weights no larger than the current 32,768 bytes, and complete CPU deployment at most 50 KiB Brotli. Compare nearest training paraphrase retrieval, train-fitted TF-IDF/BM25, and the pinned offline embedding reference on identical candidates and metadata. “Best in class” additionally needs a named comparison class, shared corpus, size accounting, and reproducible results; a development win alone is insufficient.
84
+
85
+ For the proposed new corpus, preregister additional anti-abstention and domain checks before scoring: for example at least 60% relevant semantic coverage overall, at least 50% in each sufficiently populated held-product slice, and explicit multi-positive recall. These coverage numbers are proposed prospective gates, not retroactive changes or measured successes. Report no-match confidence intervals and source-group bootstrap intervals; 0/10 observed errors is weak evidence, not a demonstrated universal 5% false-positive bound.
86
+
87
+ Before production, also require:
88
+
89
+ - One typed reusable semantic `SearchIndex` with cached candidate vectors, hard-tier invariants, reasons and raw score components, validation, immutable snapshots, abort, concurrency, and disposal; failure to load a model must explicitly degrade.
90
+ - Browser CPU end-to-end p95 at most 8 ms for 1,000 and 16 ms for 5,000 candidates on named hardware, using 30 warmups and 200 measured queries, including aliases/context, indexing cost, memory, and cold asset loading. Verify another browser and constrained device; do not infer this from Node or lexical-only timings.
91
+ - Actual exported-asset numerical parity and malformed-asset tests, query-privacy network audit, and separately compressed model/runtime/manifest size reports. WebGPU is optional until real end-to-end benefit is demonstrated; never claim acceleration without it.
92
+ - A documented supported domain, calibration limitations, source licenses, model/data hashes, and independently reviewed final evidence. Until these pass, retain the experimental label and the dependable lexical/alias path.
93
+
94
+ Reproduce the read-only diagnostics:
95
+
96
+ ```sh
97
+ uv run python -m eval.readiness-audit
98
+ node --import tsx eval/readiness-cpu.ts
99
+ ```
100
+
101
+ These commands intentionally inspect the **already consumed** synthetic regression data. They do not train weights, select a new model, alter data, or produce a new sealed-test claim.
docs/navigation-data.md ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Author-provided settings search metadata
2
+
3
+ This extraction preserves real software search metadata from three desktop providers. It contains **85 entries, 1,209 keyword-to-destination pairs, and 1,047 within-product query groups**. There are **128 groups with multiple inherited positive destinations**. It does not create training splits, candidate menus, relevance grades for other entries, or no-match judgments.
4
+
5
+ | Provider | Repository | Pinned revision | Entries | Keyword pairs |
6
+ |---|---|---|---:|---:|
7
+ | GNOME | [gnome-control-center](https://github.com/GNOME/gnome-control-center) | `cd79a897190989ca395c6f00962f0929734103c7` | 27 | 315 |
8
+ | KDE | [plasma-desktop](https://github.com/KDE/plasma-desktop), [plasma-workspace](https://github.com/KDE/plasma-workspace) | `19dae94ec1d563d095b08602dd5eed60d485a5e0`, `df3bba49a2c0657d4165f6511b1771138ccc75ea` | 40 | 754 |
9
+ | Xfce | [xfce4-settings](https://github.com/xfce-mirror/xfce4-settings) | `488c919d23c979c9b28abd192c05b9b2310fa732` | 18 | 140 |
10
+
11
+ The [freedesktop Desktop Entry specification](https://specifications.freedesktop.org/desktop-entry/latest/recognized-keys.html) describes `Keywords` as metadata useful for entry search. The [official KDE KCM documentation](https://develop.kde.org/docs/features/configuration/kcm/) identifies `X-KDE-Keywords` as terms used to search settings modules. These fields were authored for search, unlike documentation descriptions previously repurposed as queries. They remain inherited author choices, not new independently reviewed judgments or collected end-user queries.
12
+
13
+ ## Contents
14
+
15
+ - `data/navigation/raw/`: exact selected upstream metadata files, license texts, copyright notices and available SPDX/REUSE evidence. No source code is executed.
16
+ - `downloads.json`: pinned archive URL, archive hash, retained file paths, byte sizes and hashes.
17
+ - `panels.json`: original English label, description, keyword list, source path/line references, entry role, desktop visibility flags, immutable provider/product lineage and license evidence.
18
+ - `keyword-pairs.json`: exact authored query spelling and one inherited positive destination per row. It explicitly leaves other destinations unjudged.
19
+ - `query-groups.json`: groups case-insensitively within a product and unions all matching authored destinations. It never converts absent keywords into negative relevance. Original spelling remains in pairs.
20
+ - `manifest.json`: counts, pinned revisions, parser skips and explicit unresolved split/relevance policy.
21
+
22
+ Only unlocalized English fields are selected. Translation keys such as `Name[fr]` are ignored. Desktop keywords use semicolons; KDE keywords use commas. Desktop string escapes are decoded, but no paraphrases, spelling corrections, synonym guesses or descriptions-as-queries are generated. Exact duplicate pairs within one entry are removed.
23
+
24
+ There are 75 direct settings panels/modules, eight preferred-application launchers, one settings catalog and one background service. These roles are exposed so menu construction can exclude helper/service entries explicitly. Visibility flags alone cannot identify a settings panel: integrated GNOME entries may intentionally be hidden from application launchers.
25
+
26
+ ## Licensing and lineage
27
+
28
+ Raw metadata remains under its upstream terms. GNOME and Xfce preserve their repository `COPYING` GPL version 2 texts. KDE retains its mixed GPL/LGPL and other `LICENSES` texts plus available per-file notices and REUSE records. One extracted file has an explicit `GPL-3.0-or-later` SPDX identifier; the other 84 metadata files lack a direct file-level SPDX declaration. Their records retain the actual evidence and flag that absence. A repository license collection does not prove every metadata file has the same license, and this extraction does not invent a GPL version, an “or later” grant, or an MIT assignment.
29
+
30
+ The repository's MIT code license does not relicense these upstream files or derived metadata. Preserve original paths, full license notices and attribution when redistributing them. File-license resolution and any downstream model distribution decision must use the retained evidence rather than treating these records as MIT training assets.
31
+
32
+ Lineage IDs are fixed as `gnome:gnome-control-center`, `kde:plasma-desktop`, `kde:plasma-workspace`, and `xfce:xfce4-settings`. KDE products share a provider and may share concepts or wording; splitting them does not establish independent-provider generalization. Root coordination must freeze train/calibration/test lineage policy before consuming these extractions. Unlisted keywords are not evidence of irrelevance, especially for broad shared terms such as “screen” or “keyboard.”
33
+
34
+ ## Reproduction
35
+
36
+ ```sh
37
+ python -m training.prepare_navigation --download
38
+ python -m training.prepare_navigation
39
+ ```
40
+
41
+ The first command downloads exact pinned official repository archives and retains selected metadata/license files without extracting or executing application code. The second deterministically reproduces the extraction from checked-in raw files. Existing training datasets, models and evaluation splits are not modified.
42
+
43
+ Parser contracts run with `python -m unittest training.test_navigation -v`. They use temporary synthetic metadata, covering localized-field exclusion, shared-keyword positive unions, no invented negative grades, entry roles, preserved visibility/license evidence, desktop escapes and byte-identical reproduction. These tests do not train or score any provider, including GNOME.
docs/navigation-evaluation.md ADDED
@@ -0,0 +1,48 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Same-size navigation training
2
+
3
+ This evaluation selected the frozen `navigation-align0p5-seed29` experiment: **32,768 weight bytes**, exactly the previous model's size. It improves source-authored navigation retrieval and passes every existing release regression floor. It remains an experimental CPU model, not a production-approved semantic backend.
4
+
5
+ A subsequent [weight-rebalancing experiment](typo-weight-experiments.md) partially improved spelling and supplies the current demo artifact. The evidence below remains the original navigation evaluation.
6
+
7
+ ## What changed
8
+
9
+ We extracted real search keywords and destinations from pinned KDE, XFCE, and GNOME sources. KDE supplies training positives, XFCE supplies development selection, and GNOME was evaluated once after the candidate and manifest hashes were frozen. Only source-authored destinations are positive judgments; other menu items are unjudged, not proven negatives. [Data provenance and licenses](navigation-data.md) describe the sources and exclusions.
10
+
11
+ Twelve warm-start training runs tested navigation alignment, additional WordNet synonyms, and short-query augmentation. Navigation alignment won; the other additions were rejected. Nine separate experiments that traded embedding dimensions for more hash buckets also lost to matched controls. All artifacts and selection reports are retained for reproduction. No added dictionary, alias lookup, external API, or teacher model is shipped.
12
+
13
+ The selected model uses the existing feature contract and 16-dimensional int8 tables. Its SHA-256 is `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`. Selection chose epoch 2; additional epochs were not automatically better.
14
+
15
+ ## Untouched provider evaluation
16
+
17
+ GNOME contains 246 queries, including 222 without a relevant lexical match. The following values measure those 222 queries using raw model ranking and documented positives. They do not measure confidence, exhaustive relevance, or no-match safety.
18
+
19
+ | System | nDCG@5 | Documented hit in top 1 | Documented hit in top 5 |
20
+ | --- | ---: | ---: | ---: |
21
+ | Previous demo model, 32 KiB | 0.1016 | 4.95% | 21.17% |
22
+ | New navigation model, 32 KiB | 0.1705 | 7.21% | 32.43% |
23
+ | Train-paraphrase TF-IDF reference | 0.2065 | 16.22% | 31.53% |
24
+ | Offline MiniLM reference, ~91 MB weights | 0.3825 | 27.03% | 55.86% |
25
+
26
+ The new model's nDCG gain over the previous model is 0.0688. A bootstrap grouped by documented destination sets gives a 95% interval of [0.0260, 0.1102]. This supports an improvement on this provider, not a population-wide or best-in-class claim. The split contains both seen and new normalized training queries, reported separately in the full evidence.
27
+
28
+ With the existing calibration cutoff, the new model's semantic nDCG falls to 0.0063 and its top-five documented hit rate to 1.35%. These data contain no reviewed negative queries, so they cannot validate abstention. Raw model retrieval remains available in the demo’s inspection panel. Main results prioritize lexical matches and only use explicitly labeled model suggestions when no lexical matches exist; this presentation does not change the measured model quality.
29
+
30
+ ## Release decision and remaining work
31
+
32
+ All seven existing expanded-data and original-development transfer gates passed after selection, including coverage, no-match, and regression requirements; the raw model size also remains unchanged. The original consumed synthetic test was not reopened. We promoted the new weights only to the experimental demo; stable library behavior remains governed by its lexical contracts.
33
+
34
+ This model is **not production-ready or best-in-class**. It still loses to the training-example reference on ranking quality and has weak calibrated transfer. The next quality milestone requires independently reviewed product menus, multi-positive relevance and genuine no-match examples, with a new reserved evaluation before further tuning. The GNOME set is now consumed evaluation evidence and must not be reused as an untouched test. More epochs alone did not solve this gap.
35
+
36
+ Separately, [prepared candidate embeddings](prepared-model-index.md) remove repeated menu encoding from each query without altering scores or adding model bytes. Browser measurements distinguish one-time preparation from warm inference.
37
+
38
+ ## Reproduction and evidence
39
+
40
+ - `training/prepare_navigation.py` and `training/prepare_navigation_splits.py`: pinned source extraction and provider splits.
41
+ - `training/experiment_navigation.py` and `configs/navigation-*.json`: training configurations and follow-ups.
42
+ - `eval/navigation-training-selection.json`: development-only selection across all twelve runs.
43
+ - `eval/navigation-selection.json` and `eval/navigation-holdout-access.json`: exact frozen artifact and one-time evaluation receipt.
44
+ - `eval/navigation-evaluation.json`: per-query rankings, baselines, slices, hashes, and uncertainty.
45
+ - `eval/navigation-regressions.json`: all seven release regression gates passed.
46
+ - `eval/collision-summary.md`: rejected hash-bucket/dimension tradeoffs.
47
+
48
+ The evaluator defaults to development data. Holdout access requires a matching frozen record and an exclusive receipt; rerunning it does not create another independent test. The offline teacher is pinned and optional, and is never a browser dependency.
docs/optimization-experiments.md ADDED
@@ -0,0 +1,84 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Same-size optimization experiments
2
+
3
+ ## Decision
4
+
5
+ Keep the deployed `expanded-more-data-pooled-seed17` model from commit `c5a0a11`. Preserve the strongest new candidate for research, without promoting its weights. Its development gain did not transfer to the reserved synthetic evaluation. Neither model has passed the reviewed semantic release gate.
6
+
7
+ The work comprises 39 actual MLX training runs: 21 objective ablations, three wider-capacity runs, seven feature runs, three feature/objective interactions and five teacher-distillation runs. Ninety calibration configurations were evaluated on two fixed models. All training used the existing frozen training partition; calibration fitted thresholds; development selected checkpoints. These repeated development comparisons are exploratory, not independent evidence.
8
+
9
+ ## Selected research candidate
10
+
11
+ `packages/model/candidate-v2/` contains `features-bigram25-seed29-e10`, selected epoch 9, with 32,768 int8 weight bytes. It uses the same 1024-word/1024-character tables and 16 dimensions, adding ordered adjacent word-pair hashes to the existing word table. No vocabulary or teacher is shipped. Payload SHA-256: `166f105df6a3f4e0941828b1130d06671c2b5bbd759471811f607a2fa7ae30f8`.
12
+
13
+ The TypeScript research runtime is `packages/model/runtime-bigram.ts`, independently checked against the portable NumPy implementation. It is excluded from the live demo import graph. The standard loader retains its original feature contract and rejects the new feature version. The candidate's `validatedSemanticCutoff` remains null.
14
+
15
+ | Expanded development, same 1,118 semantic queries | Live baseline | Research candidate |
16
+ | --- | ---: | ---: |
17
+ | Calibrated semantic nDCG@5 | 0.46436 | 0.52254 |
18
+ | Semantic coverage | 49.55% | 55.81% |
19
+ | No-match semantic return rate, 352 menus | 4.83% | 4.83% |
20
+ | Raw semantic nDCG@5, no abstention | 0.84912 | 0.84910 |
21
+
22
+ The calibrated improvement is +0.05818, with descriptive paired query-bootstrap 95% interval [0.03250, 0.08399]. Raw ranking is essentially unchanged: the gain concerns which useful results survive calibration. Query-level intervals ignore family clustering and prior checkpoint selection. They do not establish a population-wide improvement.
23
+
24
+ Original 80-row synthetic development transfer passes the existing regression guards: raw overall nDCG rises from 0.64597 to 0.68489; calibrated overall nDCG rises from 0.41667 to 0.43333. No-match returns increase from 0/20 to 1/20. The unseen VS Code development slice still has 5/22 no-match errors (22.73%), despite passing the aggregate 5% budget. A global rate must not be advertised as a guarantee for each product family.
25
+
26
+ ## Reserved test, opened after freezing selection
27
+
28
+ The model and manifest hashes were frozen in `eval/optimized-selection.json` before the one-time access recorded in `eval/sealed-test-access.json`. The same expanded-calibration thresholds were applied to the reserved 40-row synthetic test. No threshold, checkpoint or training change followed these results.
29
+
30
+ | Reserved synthetic evaluation | Live baseline | Research candidate |
31
+ | --- | ---: | ---: |
32
+ | Calibrated semantic nDCG@5, 18 queries | 0 | 0 |
33
+ | Calibrated semantic coverage | 0% | 0% |
34
+ | Calibrated overall nDCG@5 | 0.40000 | 0.40000 |
35
+ | Raw semantic nDCG@5 | 0.38866 | 0.37082 |
36
+ | Raw overall nDCG@5 | 0.54524 | 0.53454 |
37
+ | Calibrated no-match return rate | 0% | 0% |
38
+
39
+ Abstaining on every semantic query is not useful semantic search, even with zero false positives. The small raw-score decrease is not evidence of a statistically established regression, but the test supplies no positive reason to replace the live model. Development and original-transfer floors passed; **the promotion decision is still no**. This test is synthetic, unreviewed and now consumed. Future tuning requires a new independent evaluation set; do not treat this set as fresh again.
40
+
41
+ The full `eval/optimized-evaluation.json` retains hashes, cutoffs, per-query top-five indices, slice counts and bootstrap comparisons. None of this is a reviewed final release result.
42
+
43
+ ## What was tried
44
+
45
+ | Experiment | Outcome |
46
+ | --- | --- |
47
+ | Ordered word pairs, unchanged 32 KiB | Selected research arm; feasible scores 0.5025 / 0.5225 / 0.5216 across three seeds. Equal-update gains are not independently significant for every seed. |
48
+ | More word weight, fewer character features | No eligible checkpoint in the bounded initial screen. |
49
+ | Explicit train no-match loss | More consistent across seeds, but no compelling gain over the selected word-pair model. |
50
+ | Hard-negative margin, lower temperature, source balancing | Mixed or weaker results; retained as ablations, not deployed changes. |
51
+ | Word pairs plus no-match loss | Best selected score 0.5311, only +0.00857 versus simpler word pairs; interval [-0.01614, 0.03274] includes zero, and worst-seed performance was weaker. Rejected. |
52
+ | 24 dimensions, int6 | 36 KiB raw weights. Quantization loss stayed below 0.01, but two of three seeds exceeded the no-match budget. Rejected. |
53
+ | Learned score/margin calibration | Ninety configurations across two fixed models; none improved over the original cosine cutoff. No runtime rule added. |
54
+ | Offline MiniLM teacher distillation | Four distilled students lost to the matched zero-distillation control. Absolute and centered cosine objectives were both tested. Rejected. |
55
+
56
+ Details: [training objectives and capacity](training-experiments.md), [features and interaction](feature-experiments.md), and [calibration](../eval/calibration-summary.md). Training timings materialize MLX work but may include contention from concurrent experiments; they are not controlled performance benchmarks.
57
+
58
+ ## Teacher provenance
59
+
60
+ The offline reference is [Sentence Transformers all-MiniLM-L6-v2](https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2), revision `1110a243fdf4706b3f48f1d95db1a4f5529b4d41`, licensed Apache-2.0 according to its pinned model card. Its safetensors file is 90,868,376 bytes, versus the student's 32,768 weight bytes. The teacher was loaded without remote code, using attention-mask mean pooling and normalized embeddings. Raw teacher semantic nDCG was 0.8678; calibrated nDCG was 0.1970 with 6.25% development no-match returns. A larger model alone did not solve abstention.
61
+
62
+ PyTorch 2.10.0, Transformers 4.57.6 and NumPy 2.4.3 were installed in an isolated environment. The teacher is not a production dependency. `runs/teacher-minilm/` contains the pinned model-card copy, source/model/score hashes and cached partition scores. Only `train-scores.npy` entered student losses. Calibration, development and original-transfer teacher embeddings were precomputed separately; their values were never training targets. No sealed teacher embeddings were computed.
63
+
64
+ The four distillation students and matched control used seed 17, 30 epochs, the same row order and optimizer settings, and unchanged feature/weight shape. The centered objective was an exploratory follow-up after the absolute-cosine screen. No benefit warranted further seeds. The cached teacher scores permit reproducing student training without downloading the teacher.
65
+
66
+ ## Size and verification
67
+
68
+ The complete local research demo, including UI, model metadata and weights, is **36,243 bytes Brotli (35.4 KiB)**; each resource is compressed separately. The actual live demo stays at approximately 35.5 KiB with its original weights. Both fit the 50 KiB CPU budget. Research byte counts do not imply deployment or a WebGPU implementation.
69
+
70
+ CI checks the lexical 8 KiB budget and complete CPU-demo 50 KiB budget, plus a separate isolated research build. Numerical checks cover real exported embeddings and scores, Unicode, token boundaries, ordered pairs, repeated features and truncation. Actual MLX objective checks agree with the portable reference within 2.69e-8; unjudged candidates have zero gradient in those checks. Research-quality regression tests read development only; CI does not reopen the consumed test.
71
+
72
+ ```sh
73
+ pnpm release:verify
74
+ pnpm bench:size:research
75
+ uv run python -m unittest training.test_export training.test_data_sources training.test_quality training.test_calibration training.test_experiment_objectives training.test_release_candidate
76
+
77
+ # Development and original-transfer reproduction; never reads the reserved test.
78
+ uv run python -m training.evaluate_release_candidate \
79
+ --candidate packages/model/candidate-v2 --output /tmp/gpu-search-research-evaluation.json
80
+ ```
81
+
82
+ See the linked experiment documents for immutable training commands. To reproduce a distillation run, choose a new seed not already present and use `uv run python -m training.experiment_distillation --seed 29 --weights 0 1 4`. The optional teacher preparation script requires the isolated dependencies listed above and a new `--output` directory; its pinned model downloads total roughly 87 MiB. Existing run directories are preserved rather than overwritten.
83
+
84
+ The next useful investment is independently reviewed software-navigation relevance: multiple valid destinations, realistic no-answer menus, action opposites and genuinely unseen product families. More optimization against the current development set would not resolve the observed transfer limitation.
docs/prepared-model-index.md ADDED
@@ -0,0 +1,58 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Reusing trained candidate embeddings
2
+
3
+ The experimental CPU model now supports an immutable prepared index. It validates and snapshots the menu, computes each candidate's label/alias/context vector once, and reuses those vectors as the query changes. This improves repeated inference without changing the model's 32,768 weight bytes, score ordering, or retrieval quality.
4
+
5
+ ```ts
6
+ // Repository-local module; the model runtime is not published as an npm package.
7
+ import { loadModel } from '../packages/model/runtime'
8
+
9
+ const model = await loadModel(manifest, weightsArrayBuffer)
10
+ const index = model.prepare([
11
+ { id: 'profile', label: 'Profile', aliases: ['my information'] },
12
+ { id: 'members', label: 'Members', context: 'organization settings' },
13
+ ])
14
+
15
+ const first = index.score('coworkers').slice(0, 5)
16
+ const second = index.score('personal information').slice(0, 5)
17
+
18
+ index.dispose()
19
+ ```
20
+
21
+ `prepare(candidates)` is synchronous and returns `{ size, score(query), dispose() }`. Preparation follows the core candidate count/text/alias limits and rejects malformed records or duplicate IDs with `SearchInputError`. Queries are also validated. Zero candidate embeddings are omitted. Scores are raw cosines, sorted descending with input order breaking exact ties; there is no learned relevance cutoff or cross-tier comparison.
22
+
23
+ Candidate records and aliases are copied. Later caller mutation cannot change the index or its results, and changing a returned result cannot mutate stored state. Each index owns its candidate vectors; there is no unbounded global cache. `dispose()` is idempotent, releases the index's vector and label storage, and future queries throw `ModelIndexDisposedError`. Disposing one prepared index leaves the model and other prepared indexes usable.
24
+
25
+ Rebuild when the menu changes, then dispose the old index. The demo prepares once after the model loads and again when an edited menu is accepted; ordinary query typing only encodes the query and performs cached dot products and sorting. The existing `model.score(query, candidates)` remains a backward-compatible raw reference path that prepares temporary vectors for that call. Prefer the prepared API for repeated queries.
26
+
27
+ This synchronous raw-scoring API is separate from the stable lexical `createIndex`. It does not add a validated semantic backend, query cancellation, or WebGPU to that API. The model remains experimental; faster inference does not resolve its generalization and abstention limitations.
28
+
29
+ ## Verification
30
+
31
+ Tests compare prepared and reference outputs exactly across multiple queries, aliases, and context; verify mutation isolation, stable ties, input limits, zero vectors, and independent disposal; and retain actual exported-weight numerical parity. The production demo browser suite verifies edited-menu inference, privacy, missing assets, and zero-weight behavior.
32
+
33
+ `bench/model-browser.ts` measures real trained CPU inference in Chromium and WebKit on the named local hardware. It reports asset fetch and validation/decode separately from menu preparation and warm queries at 100, 1,000, and 5,000 candidates. Every size uses 30 warmups and 200 measured queries; the 1,000-candidate case also times the uncached reference in the same browser. Full result sorting/materialization is included. Lexical ranking and a relevance cutoff are not included, so this is not a complete hybrid-search benchmark. Timer resolution can round very short calls to zero, and power state is uncontrolled.
34
+
35
+ These timings precede the word-scale rebalance; the report retains the measured artifact identity. Inference code is unchanged, but timings were not remeasured for the new weights. On 2026-09-13, the `navigation-align0p5-seed29` model ran on an Apple M4 Max with macOS 27.0.0. Times below are milliseconds. The uncached arm follows the prepared arm in the same browser session; these are local measurements, not guarantees across devices.
36
+
37
+ | Browser | Candidates | Preparation | Prepared p50 | Prepared p95 | Uncached p50 | Uncached p95 |
38
+ | --- | ---: | ---: | ---: | ---: | ---: | ---: |
39
+ | Chromium 153 | 100 | 6.9 | 0.0 | 0.1 | — | — |
40
+ | Chromium 153 | 1,000 | 42.9 | 0.2 | 0.3 | 37.1 | 39.1 |
41
+ | Chromium 153 | 5,000 | 188.1 | 1.0 | 1.1 | — | — |
42
+ | WebKit 26 | 100 | 7.0 | 0.0 | 0.0 | — | — |
43
+ | WebKit 26 | 1,000 | 36.0 | 0.0 | 1.0 | 18.0 | 19.0 |
44
+ | WebKit 26 | 5,000 | 106.0 | 1.0 | 1.0 | — | — |
45
+
46
+ Both browsers returned exactly the reference scores, made zero network requests during measured queries, and reported no errors. At 5,000 candidates the index stores 320,000 vector bytes plus candidate identifiers and labels. Preparation remains synchronous and can block the main thread for large menus; reuse removes that work from ordinary queries but does not eliminate the initial cost. Zero timings reflect browser timer resolution.
47
+
48
+ Run:
49
+
50
+ ```sh
51
+ node --import tsx --test fixtures/model.test.ts
52
+ node --import tsx bench/model-browser.ts
53
+ pnpm build
54
+ pnpm bench:size
55
+ pnpm test:demo
56
+ ```
57
+
58
+ Optional `CHROME_CHANNEL`, `WEBKIT_EXECUTABLE`, and `MODEL_BENCH_URL` overrides allow installed browser binaries and an already running root Vite server. Reports record overrides explicitly. See `bench/reports/model-browser.json` for current timings and model hashes, and `bench/reports/size.json` for complete separately compressed demo resources. The current rebalanced demo build is 36,843 bytes Brotli, below the 50 KiB CPU budget, with unchanged 32,768-byte weights.
docs/specification.md ADDED
@@ -0,0 +1,436 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # gpu search PRD and Technical Specification
2
+
3
+ Version 1.1 • 12 September 2026 • Owner Max Farrell • Audience lead implementation agent and delegated engineering agents
4
+
5
+ ## 1 Executive decision
6
+
7
+ Build a tiny, local, English-first search library for software interfaces. It must rank normalized exact matches first, strong lexical matches second, and learned semantic matches third. Train the semantic encoder with MLX on Apple Silicon. Deploy a custom TypeScript CPU implementation and optional specialized WebGPU kernels, with no general-purpose inference runtime in the browser.
8
+
9
+ Project name: `gpu-search`. Package-name availability is unverified. GPU acceleration is an implementation option, not the product's value proposition. The product is predictable, private retrieval of software concepts in a small download.
10
+
11
+ The lead agent should execute this specification, delegate independent work after fixing interfaces, integrate the results, and deliver a reproducible repository and evidence. This document authorizes implementation work; it does not authorize paid teacher API calls, publishing packages, or deploying public services. Use local or already authorized resources. Do not represent unrun training or benchmarks as completed.
12
+
13
+ **Central hypothesis:** a small domain-trained encoder can improve retrieval for unseen software interfaces enough to justify its bytes over lexical matching plus a curated alias dictionary. This is unproven. Prove it before optimizing shaders.
14
+
15
+ ## 2 Product requirements
16
+
17
+ ### Problem and intended users
18
+
19
+ Command palettes, settings search, navigation menus, and local action pickers often require users to know the application's exact terminology. A user may type “coworkers” when the destination is “Members,” or “my information” when it is “Profile.” Conventional spelling similarity cannot reliably bridge these expressions. General embedding runtimes may be disproportionate for a small menu.
20
+
21
+ The primary customer is a frontend developer supplying a local candidate list. The end user searches that list without sending queries to a server. The first release supports short English queries and labels, with optional developer-provided aliases and context.
22
+
23
+ ### Required behavior
24
+
25
+ | ID | Requirement | Acceptance evidence |
26
+ | --- | --- | --- |
27
+ | P1 | Exact label matches always precede eligible lexical matches; lexical matches always precede semantic matches | Comparator and adversarial ordering suite has zero violations |
28
+ | P2 | Correct common typos without requiring a model | Lexical-only benchmark and short-query safeguards |
29
+ | P3 | Retrieve related UI concepts with little character overlap | Held-out semantic evaluation against all baselines |
30
+ | P4 | Accept arbitrary candidate strings at runtime | No fixed concept ID or canonical-label inventory required |
31
+ | P5 | Run locally after assets load | Network audit finds no query transmission or telemetry |
32
+ | P6 | Work without WebGPU | CPU path returns the same tier decisions and numerically equivalent semantic results |
33
+ | P7 | Explain returned matches | Typed match reason, matched field, and separate score components |
34
+ | P8 | Avoid filling the list with unrelated results | Validated semantic cutoff and explicit no-match coverage |
35
+ | P9 | Support repeated interactive queries | Reusable immutable index, cached candidate embeddings, safe concurrent calls |
36
+ | P10 | Remain genuinely small | Reproducible complete transfer-size report, including model and optional runtime assets |
37
+
38
+ Initial scope is 10–5,000 candidates. Benchmark 50,000 as a stretch workload. Candidate labels are typically 1–8 words and queries 1–12 words. Support longer bounded inputs mechanically, but do not imply general document retrieval quality.
39
+
40
+ ### Examples and ambiguity
41
+
42
+ | Query | Candidates | Required relationship |
43
+ | --- | --- | --- |
44
+ | `profile` | Profile, Profiles, Account, Billing | Profile exact first; Profiles eligible lexical above Account semantic |
45
+ | `profle` | Profile, Account, Billing | Profile eligible typo above semantic results |
46
+ | `coworkers` | Members, Invoices, API Keys | Members relevant without relying on character overlap |
47
+ | `my information` | Profile, Personal Information, Billing | Multiple relevant answers allowed; precise order is evaluated |
48
+ | `plan` | Plan, Subscription, Billing | Exact Plan always first |
49
+ | `account` | Personal Account, Customer Accounts | Context-sensitive; do not declare either universally correct |
50
+ | `delete account` | Create Account, Delete Account | Exact destination first; opposite actions are difficult negatives |
51
+ | empty or whitespace | Any list | Return an empty result list |
52
+ | `xqzv` | Ordinary settings menu | Prefer no results over indiscriminate semantic matches |
53
+
54
+ “Profile ≈ account,” “organization ≈ workspace,” and “roles ≈ permissions” are context-dependent relationships, not universal equivalence classes. A model must distinguish “my profile” from “customer account” where supplied context permits it. Retrieval never executes actions or grants access. Applications must filter unauthorized candidates before indexing.
55
+
56
+ ### Non-goals
57
+
58
+ No generation, chatbot, autonomous navigation, document search, multilingual quality guarantee, automatic personalization, cloud inference, online learning, ANN dependency, or fixed database of all SaaS concepts. A Transformer is an experimental alternative only if simpler models fail a documented capability requirement. No universal semantic confidence probability is claimed.
59
+
60
+ ## 3 Success criteria and decision gates
61
+
62
+ All numbers below are proposed acceptance targets, not measured results. Register them before opening the final test set. Changes require a recorded decision, not silent threshold adjustment.
63
+
64
+ | Dimension | Initial release gate |
65
+ | --- | --- |
66
+ | Ordering | 100% exact and lexical tier invariants |
67
+ | Semantic improvement | At least +0.05 absolute nDCG@5 over the strongest lexical-plus-alias baseline on semantic-only held-out queries; paired bootstrap 95% interval for improvement excludes zero |
68
+ | Overall usefulness | No more than 0.01 absolute nDCG@5 regression on the complete held-out set |
69
+ | Typo quality | No degradation in lexical decisions relative to the same lexical engine alone |
70
+ | Abstention | At most 5% of no-match queries return any semantic result; report relevant-query coverage to prevent gaming by returning nothing |
71
+ | Quantization | At most 0.01 absolute nDCG@5 loss versus the selected float model |
72
+ | Download | Standard CPU semantic entry at most 50 KiB Brotli; CPU plus WebGPU total at most 64 KiB; lexical-only entry at most 8 KiB |
73
+ | Stretch download | Complete CPU plus WebGPU deployment at most 40 KiB Brotli |
74
+ | Warm latency | Reference Apple Silicon desktop: p95 at most 8 ms for 1,000 candidates and 16 ms for 5,000, full search including readback and ranking |
75
+ | Portability | Real Chromium WebGPU tested; another available browser/device tested; CPU fallback tested with unavailable and lost GPU |
76
+ | Reproducibility | Pinned dependencies, seeds, source manifests, dataset splits, export hashes, commands, and raw metrics |
77
+
78
+ Name the exact reference hardware and browser before collecting latency results. Cold initialization, indexing, shader compilation, and memory have mandatory reporting even though their first-release limits are not fixed here. A missed target must be reported. If no model beats the alias baseline, deliver that finding and the working baseline rather than advertise a successful semantic model.
79
+
80
+ ## 4 Public API contract
81
+
82
+ Use ESM and TypeScript declarations. Keep lexical-only imports independent of model assets. The API below is normative at the behavioral level; final names may change once before implementation with a recorded interface decision.
83
+
84
+ ```ts
85
+ type Backend = 'cpu' | 'webgpu' | 'auto';
86
+ type Candidate = {
87
+ id: string;
88
+ label: string;
89
+ aliases?: readonly string[];
90
+ context?: string; // e.g. "personal settings", not instructions
91
+ };
92
+ type Match = 'exact' | 'lexical' | 'semantic';
93
+ type SearchResult = {
94
+ id: string;
95
+ label: string;
96
+ match: Match;
97
+ reason: 'label-equality' | 'alias-equality' | 'prefix' |
98
+ 'token-prefix' | 'acronym' | 'typo' | 'concept';
99
+ matchedField: 'label' | 'alias' | 'semantic';
100
+ lexicalScore?: number; // [0, 1], meaningful only within lexical ranking
101
+ semanticScore?: number; // raw cosine [-1, 1], not probability
102
+ };
103
+ type SearchResponse = {
104
+ results: SearchResult[];
105
+ backend: 'cpu' | 'webgpu' | 'lexical';
106
+ degraded: boolean;
107
+ };
108
+ type SearchIndex = {
109
+ search(query: string, options?: {
110
+ limit?: number; signal?: AbortSignal;
111
+ }): Promise<SearchResponse>;
112
+ dispose(): void;
113
+ };
114
+ declare function createIndex(
115
+ candidates: readonly Candidate[],
116
+ options?: { backend?: Backend; semantic?: boolean }
117
+ ): Promise<SearchIndex>;
118
+ ```
119
+
120
+ Default limit is 10; limit zero returns no results; negative or noninteger limits throw. Duplicate IDs, empty labels, malformed records, and oversize inputs throw typed input errors. Duplicate labels with distinct IDs remain distinct and use insertion order as final tie breaker. Snapshot and validate inputs during construction. An empty candidate list is valid.
121
+
122
+ Default limits: 50,000 candidates; 256 Unicode scalar values per query or label; 8 aliases per candidate, each at most 128 scalars; context at most 256 scalars. Reject overflow rather than silently changing text. Document limits as configurable build constants, not training-performance guarantees.
123
+
124
+ Indexes are immutable in v1. Rebuild to change candidates. Abort must prevent stale responses from being delivered even when submitted GPU work cannot be cancelled. Queue GPU buffer use or provide per-call storage; concurrent searches must never overwrite one another. `dispose()` is idempotent; future searches reject. A lost GPU retries on CPU with the same weights and sets `degraded: true`. Failure to load valid model assets falls back to lexical with explicit degradation metadata. No unhandled async initialization races.
125
+
126
+ ## 5 Deterministic ranking specification
127
+
128
+ ### Normalization
129
+
130
+ Maintain the original text, a strict normalized key, and lexical/semantic tokens separately. Strict equality uses NFKC, ASCII A–Z lowercasing, Unicode whitespace collapse, and trim. Retain punctuation, digits, and diacritics. This English-first rule avoids Python versus JavaScript case-fold differences; expanding Unicode case behavior requires a feature-version change. Validate normalization with shared fixtures because Unicode versions can differ across runtimes.
131
+
132
+ For lexical and semantic tokens, split ASCII camelCase boundaries before lowercasing, then split on whitespace, underscore, and hyphen. Preserve meaningful symbols such as `+`, `#`, `@`, and digits. Exact comparison is never performed on the more aggressively split tokens. `C++`, `C#`, and `C` must remain distinct. Never stem or remove stopwords from the exact key.
133
+
134
+ ### Eligibility and sort key
135
+
136
+ Tier 3 is normalized equality with the label only. Alias equality is tier 2, so a developer's alias cannot displace another candidate's literal label. Context is never an exact or fuzzy matching field.
137
+
138
+ Tier 2 uses these initial lexical subclasses, in descending precedence:
139
+
140
+ 1. Alias equality.
141
+ 2. Full normalized field prefix, with at least two query scalars.
142
+ 3. Ordered token prefixes: each query token prefixes a distinct candidate token in order; every query token must be consumed, and the query has at least three non-space scalars.
143
+ 4. Acronym equality over token initials for a 2–6 character query and at least two candidate tokens.
144
+ 5. Strong typo match over the whole normalized label or alias.
145
+
146
+ For typo matching use optimal string alignment edit distance with adjacent transpositions, explicitly named OSA rather than claiming unrestricted Damerau–Levenshtein. Let L be the larger field/query scalar length. Require both lengths at least four, distance at most 1 when L is 4–7 or at most 2 when L is at least 8, and distance/L at most 0.25. Score is `1 - distance/L`. Do not use raw substring, unbounded subsequence, or n-gram similarity alone to elevate an item above semantic matches in v1. These create false lexical promotions for short queries.
147
+
148
+ For prefix subclasses use query non-space length divided by field non-space length as the score; acronym and alias equality score 1. Pick the best qualifying field/subclass per candidate, preferring label over alias on ties. Sort by descending tier, descending lexical subclass, descending lexical score, then insertion order. Exact ties use insertion order. Semantic-only candidates sort by descending cosine, then insertion order. Do not use a non-transitive epsilon comparator for near ties.
149
+
150
+ Tier 1 requires an available, nonzero embedding and cosine at or above a frozen model-specific cutoff. Below-cutoff candidates are omitted. For queries shorter than three scalars, disable learned semantic results in v1. Empty normalized queries return nothing. Only after all candidates have their best tier, apply top-k. Never prune to a lexical shortlist before semantic scoring; that would eliminate the intended synonym matches.
151
+
152
+ Semantic evidence must never promote a candidate into a lexical tier or change exact ordering. Scores across tiers are not comparable. Tighten lexical eligibility only through development evaluation and regression review; a falsely eligible lexical candidate will outrank a relevant semantic result by design.
153
+
154
+ ## 6 Model and feature contract
155
+
156
+ ### Architectural lineage and implementation decision
157
+
158
+ fastText informs the character-subword representation [S1]. StarSpace is the closer precedent for the retrieval task: representing entities through discrete features and learning task-dependent similarity in a shared embedding space, including information retrieval [S6]. Credit both; neither establishes that this project will meet its byte or quality targets.
159
+
160
+ The implementation is a custom MLX model inspired by these approaches. It does not require shipping or porting either original runtime. The proposed projections, multi-positive contrastive objective, optional teacher distillation, quantization, and deterministic ranking policy are project choices, not claims about the original StarSpace implementation. Verify original source and license before borrowing code.
161
+
162
+ Establish a StarSpace-style pooled shared-embedding baseline before adding nonlinear projections or distillation. Use the same hashed word/subword inputs, candidate composition, splits, and negative masks as the projected model; L2-normalize the pooled vector directly. First use the same contrastive loss to isolate the value of the projection. Then compare a sampled pairwise margin loss, `max(0, margin - sim(q,p) + sim(q,n))`, over judged positives and negatives, selecting margin on development. Label this an inspired baseline rather than an exact StarSpace reproduction. Reuse the pooled-model ablation already required below; do not create a duplicate workstream.
163
+
164
+ Select the simplest candidate that passes the quality and size gates. A projection or teacher is retained only with measured benefit. StarSpace is prior art and a baseline, not a new deployment dependency.
165
+
166
+ ### Starting architecture
167
+
168
+ Use a shared encoder for query and candidate strings. This is a small embedding network, not a Transformer. Hashed features avoid a shipped vocabulary, but collisions remain and unknown terms are not automatically understood. Character subword representations are established prior art; fastText documents representation through substrings [S1].
169
+
170
+ | Tensor | Shape | Parameters |
171
+ | --- | --- | ---: |
172
+ | Word embeddings | 1024 × 16 | 16,384 |
173
+ | Character embeddings | 1024 × 16 | 16,384 |
174
+ | First projection weights and bias | 24 × 16 plus 24 | 408 |
175
+ | Second projection weights and bias | 16 × 24 plus 16 | 400 |
176
+ | Total | | 33,576 |
177
+
178
+ Start without shape embeddings. Add shape or ordered token-bigram features only if ablations justify them. A pooled bag cannot reliably capture all word order, negation, or action direction. Include these limitations in the model card; add hard negatives such as enable/disable and import/export. Do not claim bag features solve those cases by construction.
179
+
180
+ Word features are each lexical token's UTF-8 bytes, prefixed by `w:`. Character features are within-token 2-, 3-, and 4-grams over Unicode scalars, with boundary markers added around each token, and namespace `c:`. Boundary markers are integer sentinels outside the scalar range, serialized through a specified tagged encoding so literal user text cannot imitate them. Freeze exact serialization in `feature-spec.json` before creating data.
181
+
182
+ Use FNV-1a 32-bit with offset 2166136261 and multiplier 16777619, unsigned wrap after each byte, and bucket `hash & 1023`. TypeScript must use `Math.imul`; Python masks with `0xffffffff`. No language-native hash functions. Preserve repeated feature occurrences and their weights; do not silently deduplicate hash collisions. Pool each feature family separately, then average the two family means. Empty family means are zero. Padding is excluded from counts and gradients. Include all feature IDs and counts in golden fixtures.
183
+
184
+ Bound model work to the first 32 tokens, each token's first 64 scalars, and at most 512 character features in deterministic token-then-n-gram-length-then-position order. These encoder bounds do not change full-string exact or lexical comparisons. Report truncation frequency in data and benchmarks.
185
+
186
+ For row-vector x, tensor storage uses output-major row-major matrices:
187
+
188
+ ```text
189
+ x = 0.5 * (mean(Eword[word_ids]) + mean(Echar[char_ids]))
190
+ h = tanh(x @ W1.T + b1)
191
+ z = h @ W2.T + b2
192
+ embedding = z / max(sqrt(sum(z*z)), 1e-8)
193
+ ```
194
+
195
+ If the pre-normalization norm is below 1e-8, mark the embedding invalid for semantic retrieval. Accumulate f32 on GPU and the reference path. Use explicit f32 rounding where necessary in the CPU parity implementation; optimized JS arithmetic may use double intermediates only when it passes tolerance tests.
196
+
197
+ Candidate representation: encode the label, each alias, and context separately with the same encoder. Average label and alias vectors with equal weight, add context vector with weight 0.25 when present, and normalize. Skip invalid vectors. This is the initial policy, not a claim that averaging synonyms is optimal. Train and evaluate with this exact composition. Benchmark label-only and alias/context ablations. Cache one vector per candidate, including a fingerprint of content, feature version, composition version, and model hash.
198
+
199
+ ### Architecture search
200
+
201
+ Compare pooled embeddings without projection, the network above, and a wider 24-dimensional model. Compare word-only and character-only variants. Start with three fixed random seeds, then run more only if uncertainty changes the decision. Measure hash bucket occupancy, collisions, and errors on unrelated brand names. Select on the quality/bytes/latency frontier, not minimum parameter count alone.
202
+
203
+ Raw storage arithmetic for the starting model is 33,576 bytes at int8 and 25,182 bytes at packed int6, excluding scales, metadata, alignment, and runtime. f32 weights occupy 134,304 bytes. At 5,000 candidates, 16-dimensional f32 candidate vectors occupy 320,000 bytes per copy. Browser memory and download size are different budgets. Brotli compression gains must be measured, not assumed.
204
+
205
+ ## 7 Data specification
206
+
207
+ ### Sources and limitations
208
+
209
+ Begin with a reviewed set of software navigation concepts and candidate menus from reusable public/open-source interfaces. Extract labels, headings, documented synonyms, paths, and task descriptions only where reuse is permitted. Public visibility alone is not a training-data license.
210
+
211
+ Mind2Web supplies web interaction task data [S2]; WorkArena evaluates knowledge work in ServiceNow [S3]. They are candidate sources of weak supervision, not ready-made short-query synonym datasets. An entire task description is not automatically a positive label for every clicked element. A task such as changing a billing address may involve navigation, search, and form submission. Retain a task-to-label pair only when the local action supports the association. Inspect source licenses and dependencies before ingestion. Do not assume WorkArena's code license covers every underlying asset.
212
+
213
+ Avoid relying on unavailable external corpora to bootstrap. Create a clearly labeled synthetic seed suite and a separately reviewed development set. Teacher-generated paraphrases may enlarge training data after filtering, but generated examples alone cannot validate usefulness.
214
+
215
+ Initial planning target: 15–30 product families, 200–500 concepts, 10,000–50,000 accepted query/menu records after augmentation. Quality and diversity matter more than hitting these counts. Aim for at least 1,000 evaluated queries with at least 200 semantic-only and 200 no-match examples. These are work targets; disclose actual counts.
216
+
217
+ ### Record format
218
+
219
+ ```json
220
+ {
221
+ "id": "record-00001",
222
+ "product_family": "example-suite",
223
+ "source_id": "manifest-entry-001",
224
+ "query": "coworkers",
225
+ "candidates": [
226
+ {"id": "members", "label": "Members", "context": "workspace access"},
227
+ {"id": "billing", "label": "Billing"}
228
+ ],
229
+ "relevance": {"members": 3, "billing": 0},
230
+ "judgment_status": "reviewed",
231
+ "query_kind": "semantic",
232
+ "augmentation_parent": null,
233
+ "split_group": "example-suite-template-family-1"
234
+ }
235
+ ```
236
+
237
+ Relevance grades: 3 directly satisfies intent; 2 useful destination; 1 related but insufficient; 0 irrelevant. Unjudged candidates are absent from the map and must not silently become negatives. Multiple positives are legal. No-match examples have all candidates judged zero. Add provenance for source URL, retrieval date, revision/hash, applicable license, extracted fields, transformation script version, teacher/model revision where used, and reviewer status. Keep sensitive user/account data out of distributed examples.
238
+
239
+ ### Splits and negative mining
240
+
241
+ Split by product family and source/template lineage before augmentation. Keep aliases, paraphrases, typo children, and near-duplicate menus together. Hold out additional wording/concept families to test transfer beyond memorized vocabulary. Product-only splitting does not eliminate shared-phrase leakage; measure overlap and report it.
242
+
243
+ Use train, development, and sealed final-test sets. Mining operates on train only. Development selects architectures, thresholds, QAT choices, and stopping points. The final set is opened once for a selected release candidate. After repeated inspection it becomes a regression set; a fresh sealed set is needed for new generalization claims. Autonomous agents may not use final-test examples to generate training data.
244
+
245
+ Siblings are potential hard negatives, not automatic negatives: Members and Users can both be relevant. Mask known positives, aliases, and uncertain candidates from negative denominators. Include realistic no-answer menus, ambiguous account meanings, visually similar unrelated labels, action opposites, number-bearing labels, and plausible incorrect destinations. Oversample observed failures within train rather than accumulating easy random negatives.
246
+
247
+ ## 8 MLX training and distillation
248
+
249
+ MLX is the training framework. Its official documentation provides automatic differentiation, array operations, saving/loading, and evaluation controls [S4]. Pin a tested version and Apple Silicon environment. Do not estimate training time from another project's epoch count. Teacher generation and data curation can dominate compute cost.
250
+
251
+ Implement deterministic feature fixtures in Python and TypeScript first. Cache prepared batches. Train the float reference with AdamW, recording learning rate, weight decay, batch size, temperature, clipping, seeds, and stopping rule. Initial search: learning rates 1e-3 and 3e-3, batch size 128 where feasible, temperature 0.07 and 0.15, up to 100 epochs with development patience 10. These are starting values, not mandatory exhaustive grid combinations. Materialize MLX lazy computations before stopping timers or saving results.
252
+
253
+ Use query-to-menu multi-positive contrastive loss. For each query, probability mass assigned to all judged relevant candidates is the numerator and all eligible judged candidates form the denominator. Exclude unjudged or conflicting negatives. Optionally weight positive grades. Do not mechanically use symmetric query/candidate loss: navigation relevance can be directional and many-to-many.
254
+
255
+ ```text
256
+ Lcontrastive = -log(sum(exp(sim(q,p)/T), p in positives)
257
+ / sum(exp(sim(q,c)/T), c in eligible candidates))
258
+ ```
259
+
260
+ Use numerically stable logsumexp. All-zero relevance menus do not use an empty positive numerator; use them for abstention calibration and, if justified, a separately specified margin loss. Avoid forcing an arbitrary absolute cosine target before measuring the distribution.
261
+
262
+ Train a supervised-only baseline before teacher use. If distillation is beneficial, select a pinned locally available embedding teacher or reranker and cache menu-level scores offline. Validate the teacher against reviewed UI examples first. Generic semantic similarity can confuse related but operationally different settings. Distill distributions over the same candidate menu with KL divergence and a development-selected weight; do not copy illustrative cosine values from the prior discussion as labels. Teacher scores are not ground truth or probabilities of correct navigation.
263
+
264
+ No teacher is shipped. Record model license, revision, prompts if applicable, temperatures, inference settings, and filtering rules. If access or budget blocks teacher inference, complete the supervised pipeline and document the missing experiment.
265
+
266
+ ## 9 Quantization and export
267
+
268
+ Try post-training int8 first. Compare float, int8, then int6; int4 is optional. Add quantization-aware training only when it resolves a measured quality loss. Choose the last 10–20% of training steps as an initial QAT phase, with development-based selection rather than copying another project's final-epoch count.
269
+
270
+ Define one deployment-matched symmetric per-tensor scheme:
271
+
272
+ ```text
273
+ Q = 2^(bits-1) - 1
274
+ scale = max(abs(W)) / Q, or 1 for an all-zero tensor
275
+ round_away(t) = sign(t) * floor(abs(t) + 0.5)
276
+ codes = clip(round_away(W / scale), -Q, Q)
277
+ Wq = codes * scale
278
+ Wfake = W + stop_gradient(Wq - W)
279
+ ```
280
+
281
+ The straight-through estimator is necessary: bare round/clip is not a sufficient differentiable training recipe. Freeze scales from float master tensors for each forward evaluation, stopping gradient through the quantization path. Record scale f32 serialization. Include small bias tensors in the quantization scheme initially; retain f32 biases only as an explicitly measured format change.
282
+
283
+ Export a versioned manifest and packed binary payload. Required metadata: format version, feature version, normalization version, candidate-composition version, model architecture, ordered tensor names, shapes, element counts, quantization bits, f32 scales, byte offsets/lengths, payload SHA-256, model ID, and validated semantic cutoff. Offsets are bytes into the actual packed payload, never float element counts.
284
+
285
+ For int6, map signed code q in [-31,31] to unsigned `u=q+31` in [0,62]; 63 is invalid. Pack codes least-significant-bit first into a byte stream; zero unused trailing bits; start each tensor at a byte boundary. Declare byte order for header numbers. Reject invalid codes, nonfinite scales, overflow, overlapping ranges, unexpected shapes, and feature-version mismatches. Test pack/unpack round trips at extrema and non-byte-aligned counts.
286
+
287
+ Generate export fixtures using unpacked, quantized weights and the exact deployment computation. Save feature IDs, pooled vectors, projection outputs, norms, final vectors, and menu scores. Compare float-versus-quantized quality separately from quantized-reference-versus-browser parity.
288
+
289
+ ## 10 Browser runtime
290
+
291
+ CPU implementation comes first and is the fallback. Unpack and dequantize weights once to f32 buffers, cache candidate embeddings, and compute query embedding plus dot products at search time. No ONNX, Transformers.js, or other general runtime in the production bundle. Such tools may be used as offline baselines.
292
+
293
+ Implement the same mathematical operations in WGSL, whose normative language specification is maintained by W3C [S5]. Specialized kernels cover embedding gather/reduction, projection, normalization, and candidate dot products. Begin with straightforward f32 kernels; do not require shader-f16 or attempt native six-bit arithmetic. Packed weights reduce transfer size, while runtime compute remains float.
294
+
295
+ Feature extraction, exact/fuzzy checks, and final tier ordering stay on CPU. Candidate embeddings stay resident on GPU when available. Return score buffers for deterministic final sorting. For 5,000 candidates the readback is only 20,000 bytes, but dispatch/synchronization can still dominate. Measure the complete operation.
296
+
297
+ Initialize adapter/device and pipelines once per runtime. Reuse bounded growable buffers, handle device loss, honor device limits, and dispose allocations. Parallel GPU reductions must use valid workgroup barriers and no cross-workgroup synchronization assumptions. Keep weights shareable between indexes if this does not complicate lifecycle correctness.
298
+
299
+ `auto` initially chooses CPU. Enable a GPU crossover only after end-to-end benchmarks on named devices show improvement. Candidate count alone may not capture query length, cold state, or index residency. Forced WebGPU selects it when available but falls back explicitly when unavailable. Importing the package must not request a GPU or fetch assets as a side effect.
300
+
301
+ Model assets may be bundled or explicitly loaded by the host. Document both approaches and count all resources in size reports. The default production library emits no telemetry. Render labels as text in the demo, not raw HTML.
302
+
303
+ ## 11 Evaluation and verification
304
+
305
+ ### Baselines
306
+
307
+ Evaluate exactly the same menus, queries, aliases, and available context with:
308
+
309
+ 1. The proposed lexical engine alone.
310
+ 2. Lexical engine plus a compact, curated alias map built using train/development only.
311
+ 3. Character n-gram TF-IDF cosine retrieval combined with the same hard tiers.
312
+ 4. A pinned general embedding model as an offline quality reference, with its actual model size reported.
313
+ 5. The StarSpace-style pooled shared-embedding baseline specified in Section 6, including the controlled loss comparison.
314
+ 6. Each projected tiny learned variant, with and without distillation and quantization.
315
+
316
+ An alias dictionary is a serious competing product. Include its complete compressed bytes. Keep host-supplied aliases identical across systems; a learned system must not receive metadata withheld from baselines. The general model is a reference, not a guaranteed upper bound.
317
+
318
+ ### Metrics
319
+
320
+ Report nDCG@5 using gain `2^grade-1`, MRR using relevance at least 2, and Recall@1/3/5 over the same threshold. Distinguish hit rate from recall when there are multiple positives. Report conditional semantic quality both before abstention and after the cutoff, and coverage. For no-match menus report the fraction returning any semantic item and the fraction returning any item, including lexical false positives.
321
+
322
+ Slices: exact, typo, prefix, acronym, semantic-only, multiword intent, ambiguous label, action opposites, unseen products, unseen wording/concept families, no-match, and encoder truncation. Semantic-only excludes queries with a relevant eligible lexical match under the frozen matcher. Show counts and confidence intervals; tiny slices are descriptive, not definitive. Choose a global cosine cutoff on development to meet the false-positive budget; if inadequate across menu sizes, document failure before adding menu-size or margin conditioning.
323
+
324
+ ### Parity and runtime tests
325
+
326
+ Feature IDs, normalization outputs, packed codes, and counts must match exactly. Starting numeric acceptance is maximum absolute embedding-component error at most 1e-4 and score error at most 2e-4 between quantized MLX, CPU, and WebGPU. Investigate deviations by stage rather than widening tolerance blindly. Rankings must agree for comparisons whose reference score gap exceeds twice the score-error tolerance. Near-tie inversions are reported separately; hard tier decisions must still agree exactly.
327
+
328
+ Include empty/zero embeddings, long inputs, repeated tokens, Unicode normalization, punctuation, digit labels, all quantizer extrema, malformed payloads, duplicate IDs, duplicate labels, abort, concurrent searches, disposal, missing GPU, lost GPU, failed asset loads, and prefix false positives. Finite outputs and deterministic tie ordering are release gates.
329
+
330
+ ### Performance and size methodology
331
+
332
+ Run candidate sizes 10, 50, 250, 1,000, 5,000, and 50,000. Record cold initialization, initial index encoding, warm queries, full CPU and GPU timings, p50/p95, JS main-thread time, and memory. Include realistic labels, contexts, and aliases rather than only cached synthetic vectors. Use at least 30 untimed warmups and 200 measured queries for each warm configuration; randomize query order and record power state and browser version. Use GPU timestamps only as supplementary diagnostics, not user-visible latency.
333
+
334
+ Measure the minified production import graph and model resources as served. Pin minifier and Brotli implementation and settings; report Brotli quality 11/window 22 plus raw and gzip sizes. Compress separate network resources separately and sum them, rather than claiming the size of an unrealistically concatenated archive. Report lexical-only, CPU semantic, optional WebGPU increment, and total. Exclude source maps from deployment totals only if they are actually not served; disclose the exclusion.
335
+
336
+ ## 12 Repository and reproducible commands
337
+
338
+ ```text
339
+ packages/core/src/ API, normalization, lexical ranking, CPU backend
340
+ packages/webgpu/src/ GPU lifecycle and WGSL
341
+ packages/model/ generated manifest and quantized model assets
342
+ training/ MLX model, losses, experiments, QAT, exporter
343
+ data/ source manifests, split definitions, permitted fixtures
344
+ eval/ metrics, baseline adapters, sealed-test runner
345
+ bench/ real-browser timing and size reports
346
+ fixtures/ shared feature and numerical parity fixtures
347
+ apps/demo/ minimal command palette and benchmark view
348
+ docs/ feature contract, model card, decisions, reproduction
349
+ ```
350
+
351
+ Use one JS lockfile and one pinned Python environment. Choose concrete scripts implementing this command contract:
352
+
353
+ ```sh
354
+ pnpm install --frozen-lockfile
355
+ uv sync --frozen
356
+ pnpm test:contracts
357
+ uv run python -m training.prepare --config configs/data.yaml
358
+ uv run python -m training.train --config configs/base.yaml
359
+ uv run python -m training.evaluate --split dev --checkpoint runs/selected
360
+ uv run python -m training.export --checkpoint runs/selected --bits 8
361
+ pnpm test:parity
362
+ pnpm bench:browser
363
+ pnpm bench:size
364
+ pnpm release:verify
365
+ ```
366
+
367
+ These are required future entry points, not commands already run. Training requires an appropriate MLX environment. Ordinary CI must still run CPU contracts, export-format tests, and fixed-fixture parity without retraining. Hardware-dependent jobs must report skipped status accurately.
368
+
369
+ ## 13 Delegation plan for the lead agent
370
+
371
+ The lead agent owns product scope, the feature/API freeze, decisions, integration, and release evidence. Delegate bounded work by owned paths. Agents may propose interface changes, but must not independently invent incompatible hashing, score meanings, tensor layouts, or export formats.
372
+
373
+ | Workstream | Ownership and output | Dependency |
374
+ | --- | --- | --- |
375
+ | Data and provenance agent | `data/`, source adapter proposals, leakage report, reviewed menus, licenses | Record schema and split policy |
376
+ | Lexical and API agent | `packages/core/` normalization, tiers, API lifecycle and tests | Frozen normalization and API |
377
+ | MLX modeling agent | `training/`, trained candidates, ablations, quantized export proposal | Shared feature fixtures and train/dev data |
378
+ | Evaluation agent | `eval/`, baselines, sealed-test process, scorecards | Frozen relevance and metric definitions |
379
+ | WebGPU agent | `packages/webgpu/`, parity and loss recovery | Selected quantized export and CPU reference |
380
+ | Packaging and demo agent | `bench/`, `apps/demo/`, size automation, browser evidence | Stable API and model loading contract |
381
+
382
+ With limited agent slots, combine packaging/demo and evaluation or defer them. The lead should continue useful integration work while agents run. A data agent and evaluation agent can work independently after the schema freezes. Do not start shader optimization while the architecture is changing.
383
+
384
+ Each assignment must include owned files, inputs, output interfaces, acceptance tests, constraints, and a completion report containing changes, commands actually run, outcomes, and blockers. Agents must not write each other's owned files without coordination. The lead resolves shared-file conflicts and runs the integrated release checks. Independent evaluation review should inspect leakage and baseline fairness before final-test execution.
385
+
386
+ ## 14 Milestones and stop conditions
387
+
388
+ **M0 Contracts and evidence.** Inspect reference repositories at pinned commits, verify licenses, freeze API/features/format, prepare golden fixtures, and register evaluation targets. Confirm Apple Silicon training and real-browser test access. If unavailable, implement portable pieces and label hardware-dependent work blocked.
389
+
390
+ **M1 Baselines and data.** Deliver the lexical library, alias baseline, source manifest, train/dev/test definitions, and first evaluation report. This is a useful product even before ML.
391
+
392
+ **M2 Semantic feasibility.** Train the StarSpace-style pooled baseline and projected MLX encoders and compare against aliases. Inspect errors and run the finite architecture/feature ablations. If two materially different small models fail the semantic improvement gate, stop compression work and report whether data quality, capacity, or weak product benefit appears limiting. Do not spend indefinitely optimizing failed models.
393
+
394
+ **M3 Deployable CPU model.** Select quantization, export deterministic assets, pass CPU parity, implement abstention and lifecycle behavior, and meet the CPU transfer budget.
395
+
396
+ **M4 WebGPU parity and benefit.** Implement kernels, verify actual browser execution and fallback, and measure crossover. If GPU provides no useful speedup, retain it as an explicit experimental backend and keep auto on CPU. Do not claim acceleration from kernel-only timings.
397
+
398
+ **M5 Release candidate.** Integrate demo, run sealed evaluation once, generate model card and measured size/latency reports, and produce a local reviewable package. Publication is a separate user-authorized action.
399
+
400
+ The demo must show an editable query, candidate menu, match reasons, backend, and comparable lexical-only results. Include representative semantic and typo cases without claiming curated examples prove generalization. No unrelated dashboard or marketing work.
401
+
402
+ ## 15 Final handoff checklist
403
+
404
+ - Working local repository with clean reproduction instructions and pinned dependencies.
405
+ - Exact/fuzzy/semantic behavior and typed API documented with runnable examples.
406
+ - MLX source, configuration, selected float checkpoint, and exported quantized artifact where training ran.
407
+ - Data provenance and license manifest; split leakage audit; clear synthetic versus reviewed counts.
408
+ - Baseline comparison and ablation table with actual numbers, uncertainty, and failures.
409
+ - Feature and quantized numerical fixtures; CPU/WebGPU parity results on named devices.
410
+ - Complete transfer bytes, runtime memory, and end-to-end timings.
411
+ - Model card stating intended domain, ambiguity limitations, quantization, evaluation scope, and no-confidence-probability caveat.
412
+ - Decision log explaining model choice, alias comparison, quantization, and auto-backend policy.
413
+ - Accurate list of skipped or blocked work; no fabricated training success or browser test results.
414
+
415
+ ## 16 Reference evidence and corrections to the earlier discussion
416
+
417
+ The architecture and gates in this document are design decisions. The prior conversation's claims about particular model sizes, six-bit schemes, training seconds, and repository internals were not independently verified in this authoring pass because the referenced GitHub pages could not be fetched. Do not copy their numbers into a README as facts. Before borrowing code, inspect the actual repository and license at a pinned commit. There is no requirement to fork either repository; a small clean implementation may be easier to audit.
418
+
419
+ The useful proposed pattern is MLX training → deployment-matched quantization → custom weight export → CPU/WGSL implementations → numerical parity → size and retrieval gates. MLX preference does not establish that training will take seconds. Tiny weights do not establish useful semantic quality. Six-bit storage does not imply six-bit compute. Prefixes are lexical, not exact. Cosine is not calibrated confidence. Teacher outputs and UI hierarchies require relevance judgment.
420
+
421
+ | ID | Primary reference | Use and verification status |
422
+ | --- | --- | --- |
423
+ | S1 | [fastText word representations](https://fasttext.cc/docs/en/unsupervised-tutorial.html) | Accessed; subword embedding prior art, not evidence that this size target will succeed |
424
+ | S2 | [Mind2Web project](https://osu-nlp-group.github.io/Mind2Web/) | Accessed; candidate web-task source requiring extraction and license review |
425
+ | S3 | [WorkArena project](https://servicenow.github.io/WorkArena/) | Accessed; enterprise task benchmark, not a turnkey synonym dataset |
426
+ | S4 | [MLX documentation](https://ml-explore.github.io/mlx/build/html/index.html) | Accessed; training framework reference; pin the version actually used |
427
+ | S5 | [WebGPU Shading Language specification](https://www.w3.org/TR/WGSL/) | Accessed; language/operator reference for browser kernels |
428
+ | S6 | [StarSpace paper](https://arxiv.org/abs/1709.03856) | Accessed for version 1.1; shared feature-based embeddings and task-dependent retrieval similarity |
429
+ | R1 | [gpu lexer architecture](https://github.com/vercel-labs/gpu-lexer/blob/main/architecture.md) | Supplied reference; fetch failed in this pass; verify before reuse |
430
+ | R2 | [gpu cron model card](https://github.com/manuschillerdev/gpu-cron/blob/main/MODEL_CARD.md) | Supplied reference; fetch failed in this pass; verify before reuse |
431
+ | R3 | [Manu Schiller post](https://x.com/manuschiller/status/2098801726023139403) | User-supplied provenance; not independently verified here |
432
+
433
+
434
+ ## 17 Revision history
435
+
436
+ Version 1.1 names the project gpu-search, credits StarSpace alongside fastText, makes the pooled retrieval baseline explicit, and requires controlled comparison before retaining projections or distillation. Core product scope, MLX deployment approach, and exact/lexical/semantic tier precedence are unchanged.
docs/training-experiments.md ADDED
@@ -0,0 +1,56 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Training objective and capacity experiments
2
+
3
+ Twenty-one 16-dimensional runs compared seven fixed training choices over seeds 17, 29 and 43. A three-run 24-dimensional follow-up tested capacity and int6 storage. These are weak-label development experiments; none used the original development set for training or selection, and none establish reviewed final-test quality.
4
+
5
+ The existing runtime, feature contract and deployed model were preserved. The 16-dimensional experiments use two 1024×16 embedding tables, totaling 32,768 int8 weight bytes. The 24-dimensional alternative requires 49,152 int8 or 36,864 packed int6 weight bytes, before metadata and runtime code.
6
+
7
+ ## Controlled setup
8
+
9
+ All arms use the frozen expanded dataset, 7,165 sampled records and 56 optimizer steps per epoch, AdamW learning rate 0.003, and three fixed seeds. The original arm uses temperature 0.15. Training stops after development patience 10, with at least 10 and at most 60 epochs; capacity is capped at 30. All corresponding selected 16-dimensional no-match epochs were below 30.
10
+
11
+ Cutoffs are fitted exclusively to calibration no-match records. Checkpoint selection first requires at most 5% development semantic false positives, then maximizes post-cutoff semantic nDCG@5 and coverage. If no epoch satisfies that gate, the retained checkpoint is explicitly **NOT SELECTED**, with a false gate flag. Such weights are failure-analysis artifacts, never automatically deployable candidates.
12
+
13
+ The no-match arm adds a weight-0.25 squared hinge above cosine 0.35 for TRAIN all-zero menus. It never creates an empty positive contrastive numerator. The hard-negative arm adds a weight-0.25 margin of 0.15 against the hardest judged negative already in that training menu; it does not infer global negative labels. Source balancing samples sources uniformly, then rows within source. These choices were fixed before their runs.
14
+
15
+ The fixed total sample budget means no-match arms see fewer positive examples per epoch, and balanced arms repeat small sources more often. Actual positive/no-match and source counts are logged. This is equal optimizer-budget experimentation, not equal unique-example exposure.
16
+
17
+ ## Actual 32 KiB results
18
+
19
+ | Arm | Seeds meeting development false-positive gate | Mean semantic nDCG@5 | Mean held-family nDCG@5 | Mean unseen UI namespace nDCG@5 |
20
+ | --- | ---: | ---: | ---: | ---: |
21
+ | Control | 2/3 | 0.4162 | 0.1265 | 0.1360 |
22
+ | Temperature 0.07 | 3/3 | 0.3341 | 0.0929 | 0.1353 |
23
+ | Train no-match loss | 3/3 | 0.4598 | 0.1733 | 0.2399 |
24
+ | Hard-negative margin | 2/3 | 0.4157 | 0.1421 | 0.1206 |
25
+ | Source balanced | 3/3 | 0.4473 | 0.1610 | 0.2969 |
26
+ | Balanced plus no-match | 2/3 | 0.4441 | 0.1382 | 0.3636 |
27
+ | Balanced plus hard margin, temperature 0.07 | 3/3 | 0.3684 | 0.1089 | 0.1058 |
28
+
29
+ Means include the explicitly rejected retained checkpoint where a seed failed; the gate column is essential. No-match loss had the strongest consistent result in this objective screen: worst-seed nDCG 0.4501 and all three development false-positive rates below 5%. Source balancing helped the small UI slice, but reduced aggregate consistency. These findings do not rank separately run feature-representation experiments.
30
+
31
+ The exported no-match seed-17 candidate is at `packages/model/experiments/no-match-seed17/`. It remains explicitly experimental with no validated deployment cutoff. A later independent quantized evaluation and comparison against other workstreams must determine any promotion.
32
+
33
+ ## Capacity decision
34
+
35
+ Using the same no-match objective, the 24-dimensional seed-17 float checkpoint reached 0.4975 semantic nDCG@5 and 4.83% development false positives. At its **unchanged float calibration cutoff**, int8 scored 0.4964 with loss 0.0011; int6 scored 0.4936 with loss 0.0039 and 4.26% false positives.
36
+
37
+ The other two capacity seeds failed the false-positive gate: int6 rates were 8.52% and 9.09%. All three int6 quantization losses were below 0.01, but capacity did not demonstrate consistent safe coverage. The roughly 0.025 seed-17 improvement over the smaller no-match model does not justify a larger payload and decoder/runtime change on this evidence. The capacity alternative was rejected; its artifacts remain isolated for analysis.
38
+
39
+ ## Evidence and verification
40
+
41
+ `eval/training-experiments.json` contains all 21 histories' selected metrics and slice summaries. `eval/training-experiments-capacity.json` records the three capacity checkpoints and fixed-cutoff float/int8/int6 comparisons. Each `runs/experiment-*/` directory retains full histories, epoch-1/10 checkpoints, configuration, source hashes and sampled counts.
42
+
43
+ Training ran on Apple M4 Max using MLX 0.31.1 Metal. The separate feature workstream sometimes used the GPU concurrently, so timing is contention-affected and cannot compare kernels or architectures fairly. Recorded loop times include checkpointing and per-epoch development evaluation but exclude initial data/feature/cache preparation; optimizer time is recorded separately. No latency claim follows from these training times.
44
+
45
+ Eleven portable NumPy tests verify probability-mass behavior, positive/negative monotonicity, multiple positives, padding and unjudged masks, no-match and mixed batches, finite-weight rejection, and truthful saved gate flags. A separate actual MLX check covered 15 objective/mask cases: maximum objective error versus NumPy was 2.69e-08, unjudged gradients were exactly zero, and every checked gradient was finite.
46
+
47
+ ```sh
48
+ uv sync --frozen
49
+ uv run python -m unittest training.test_experiment_objectives
50
+ # Requires Apple Silicon Metal; tags preserve all existing checkpoints.
51
+ uv run python -m training.experiment_training --tag reproduction
52
+ uv run python -m training.experiment_capacity --tag reproduction
53
+ uv run python -m training.verify_experiment_objectives
54
+ ```
55
+
56
+ Use `--arm no-match --seed 17 --tag another-run` for a bounded objective rerun. Existing checkpoint directories are never overwritten. Development is already used for model selection; new product-quality claims require new reviewed held-out data.
docs/typo-training.md ADDED
@@ -0,0 +1,53 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Weight-only typo consistency experiment
2
+
3
+ None of the eleven training runs was promoted. The best passing exported model improved the mixed typo development score from 0.522949 to 0.531955, below the separately evaluated word-table rebalancing candidate's 0.550897. Gains were small and inconsistent across seeds. All models retain the v1 pooled16 architecture and 32,768-byte int8 payload; this experiment adds no runtime dictionary, feature rule, or decoder change.
4
+
5
+ The [feature diagnosis](typo-weight-diagnosis.md) found that the existing character representation favors the intended spelling, while the word hash vector dominates the pooled result. We tested learned consistency rather than changing any individual query's ranking.
6
+
7
+ ## Frozen scope and objective
8
+
9
+ Warm start: `packages/model/experiments/navigation-align0p5-seed29`, payload SHA-256 `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`, dequantized before optimization. Training used 7,165 positive expanded-training menus, 750 explicitly documented KDE navigation positive pairs, and 4,212 synthetic noisy/clean pairs from 351 training-derived labels. Typo label groups were split before corruption. These new augmentation groups are disjoint, but the base model may already have seen their clean labels. Synthetic corruption is not human query evidence.
10
+
11
+ The objective combines existing judged-menu contrastive cross entropy (temperature 0.15), KDE positive alignment `0.5 * mean(max(0, 0.7 - cosine)^2)`, typo consistency `weight * mean((1 - cosine(noisy, stop_gradient(clean)))^2)`, and frozen clean-teacher retention `mean((1 - cosine(current, initial))^2)`. Teacher observations include sampled anchor text, navigation positive text, and canonical typo labels. There is no explicit teacher matching for noisy variants. Missing navigation judgments never become negatives.
12
+
13
+ Every run used eight epochs, batch size 128, AdamW decay 0.0001, and gradient norm clipping at 1.0. The first seven runs used learning rate 0.001. After all seven retained the initial checkpoint, four explicitly exploratory followups changed only learning rate to 0.0001 and tested typo weights 0.5 and 2. Both stages include epoch zero for checkpoint selection, with earliest ties retained. Seeds control minibatch and positive-pair sampling; all share the same warm start.
14
+
15
+ Selection used the mean of raw typo top-1 on all 1,248 development queries and on 300 queries whose clean label has at most two tokens. The remaining 948 queries have longer labels. This prevents long setting names from hiding short-query failures. Checkpoints must also satisfy expanded overall nDCG@5 ≥ 0.447708, semantic nDCG@5 ≥ 0.45, **semantic** coverage ≥ 0.48, semantic false-positive rate on no-match menus ≤ 0.05, and XFCE raw semantic nDCG@5 ≥ 0.5113. Thresholds use expanded calibration only; expanded and XFCE development are selection gates. XFCE scores measure documented positives with incomplete judgments, not exhaustive relevance.
16
+
17
+ No typo holdout, Profile diagnostic, consumed GNOME test, or original development examples were read or scored by this trainer. Input hashes are recorded and checked after each run. The authorized initial named-query feature diagnosis is separate from selection; new checkpoints were not tested against that diagnostic here.
18
+
19
+ ## Exported results
20
+
21
+ All numbers below come from independently recomputed NumPy evaluation of the exported int8 weights. A passing float checkpoint can fail after quantization and calibration; the export must pass again.
22
+
23
+ | LR | Typo weight | Seed | Selected epoch | Short top-1 | All top-1 | Long top-1 | Mixed score | Export passes gates |
24
+ | --- | --- | --- | --- | --- | --- | --- | --- | --- |
25
+ | .001 | 0 | 17 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
26
+ | .001 | .5 | 17 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
27
+ | .001 | .5 | 29 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
28
+ | .001 | 2 | 17 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
29
+ | .001 | 2 | 29 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
30
+ | .001 | 5 | 17 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
31
+ | .001 | 5 | 29 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
32
+ | .0001 | .5 | 17 | 1 | .366667 | .691506 | .794304 | .529087 | yes |
33
+ | .0001 | .5 | 29 | 4 | .370000 | .695513 | .798523 | .532756 | **no** |
34
+ | .0001 | 2 | 17 | 0 | .360000 | .685897 | .789030 | .522949 | yes, unchanged |
35
+ | .0001 | 2 | 29 | 2 | .370000 | .693910 | .796414 | .531955 | yes |
36
+
37
+ The failed weight-0.5, seed-29, lower-LR export has semantic nDCG@5 0.448420 and semantic coverage 0.477639, below both floors. Its float checkpoint passed; its exported weights do not. It is retained as a failed research artifact with `int8PassesGates: false`, and must not be promoted based on its higher typo score.
38
+
39
+ The best passing weight-2, seed-29, lower-LR export has expanded overall nDCG@5 0.454315, semantic nDCG@5 0.453339, semantic coverage 0.482111, no-match false-positive rate 0.036932, and XFCE raw semantic nDCG@5 0.512756. It adds only three correct short-query predictions over the starting model (111/300 versus 108/300), while reducing expanded semantic quality and approaching the XFCE floor. The other seed at weight 2 selected epoch zero. These results do not establish a robust learned improvement over the simpler rebalancing alternative.
40
+
41
+ ## Evidence and reproduction
42
+
43
+ Full configuration, input hashes, selected metrics, export hashes and independent evaluations are in `eval/typo-training.json` and `eval/typo-training-lowlr.json`. Per-epoch histories and immutable selected checkpoints are in `runs/typo-consistency*`; research exports are in matching directories under `packages/model/experiments/`. No production artifact was changed by this trainer.
44
+
45
+ Using the pinned uv environment on Apple Silicon with MLX Metal:
46
+
47
+ ```sh
48
+ uv sync --frozen
49
+ uv run python -m training.experiment_typo --config configs/typo-training.json
50
+ uv run python -m training.experiment_typo --config configs/typo-training-lowlr.json --tag lowlr
51
+ ```
52
+
53
+ The trainer refuses to overwrite existing run or artifact directories. Reproduction therefore requires a fresh output workspace. Recorded per-run loop times were 1.47–1.54 seconds on this M4 Max, covering training and in-loop MLX development evaluation but excluding cache construction and independent post-export evaluation. These local times are not a portable throughput benchmark. Training stopped after the authorized eleven runs; further holdout evaluation and any release decision belong to the separately frozen candidate workflow.
docs/typo-weight-diagnosis.md ADDED
@@ -0,0 +1,28 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Why the learned encoder confuses `profle`
2
+
3
+ The current navigation checkpoint scores `profle` against `Profile` at **0.2390**, but against `Invoices` at **0.7065**. This is a weight/representation problem even though the library's separate lexical tier can recognize the typo. No runtime or ranking changes were made for this diagnosis.
4
+
5
+ The character features already provide the right signal: character-only cosine is **0.8080** for `profle`/`profile` and **−0.1876** for `profle`/`invoices`. The typo retains 12 shared character buckets with `profile`.
6
+
7
+ The word and character families have equal *coefficients*, but unequal vector magnitudes. For `profle`, the single word embedding has norm **0.8285**, while the average of 18 character features has norm **0.2646**—a 3.13× ratio. Averaging many character rows reduces their magnitude; equal family coefficients do not imply equal influence after normalization.
8
+
9
+ | Query/candidate | Full cosine | Query word contribution | Query character contribution |
10
+ | --- | ---: | ---: | ---: |
11
+ | `profle` → `Profile` | 0.2390 | 0.0444 | 0.1946 |
12
+ | `profle` → `Invoices` | 0.7065 | 0.6642 | 0.0423 |
13
+
14
+ The contributions sum to full cosine and include the candidate's complete embedding. They are not the same as independently normalized word-only or character-only similarity.
15
+
16
+ The word hash makes an unseen spelling select an unrelated learned row:
17
+
18
+ - `profle` hashes to word bucket 208, shared by TRAIN tokens including `shift`, `trim`, `350` and `transferred.`.
19
+ - `profile` hashes to bucket 17, shared with tokens including `dictation` and `fields`.
20
+ - `invoices` hashes to bucket 84, shared with `accessibility`, `tree` and others. `accessibility` occurs 19 times across the inventoried unique TRAIN texts.
21
+
22
+ These are actual feature aliases; they do not prove which particular training example caused the final vectors. The typo and `Invoices` do not directly share their word bucket. Rather, the typo receives an unrelated word vector whose direction dominates its useful character signal.
23
+
24
+ A weight-level experiment should therefore add **train-only clean/noisy embedding consistency**, alongside the existing positive-menu and navigation objectives, while preserving clean embeddings with a frozen clean teacher. This asks the model to compensate for noisy word hashes using retained subword information. It adds no inference bytes. Broad typo-label training is preferable to a special case for `profle`; the named Profile example must remain outside training and checkpoint selection.
25
+
26
+ Do not infer that stronger consistency is always better: shared word rows also represent legitimate unrelated terms, so excessive alignment can damage clean retrieval. Keep clean-quality, coverage and no-match gates, and evaluate corruptions grouped by previously unused clean labels. Family renormalization or a character-heavy inference mix would change runtime behavior and is outside this weight-only experiment.
27
+
28
+ The exact vectors, feature IDs, norms, collision terms and source hashes are in `eval/typo-weight-diagnosis.json`. The artifact payload hash is `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`; only existing TRAIN corpora were scanned. No consumed source test or typo holdout was read.
docs/typo-weight-experiments.md ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Investigating the model's typo failure
2
+
3
+ We found a **partial weight-level improvement**, not a full fix. The demo now uses `navigation-align0p5-seed29-word085`: the word-table dequantization scale is 15% smaller, while the character table is unchanged. The int8 payload remains exactly **32,768 bytes**. This changes the effective embedding weights; it adds no dictionary, query correction, ranking exception, or inference code.
4
+
5
+ The raw model still ranks `profle` against **Invoices** above **Profile**. The main search correctly returns Profile through its separate text-matching tier. These are different claims: safer application behavior does not establish reliable learned spelling.
6
+
7
+ ## Cause
8
+
9
+ [Feature diagnosis](typo-weight-diagnosis.md) showed that the character representation already favors Profile. The unrelated word-hash vector selected by `profle` dominates the pooled representation: its norm is about 3.13 times the mean character vector's norm. Equal family coefficients do not mean equal influence. Hash buckets share rows across unrelated words, limiting what a single weight update can repair without affecting other inputs.
10
+
11
+ ## Experiments and selection
12
+
13
+ We generated single-edit spelling variants from existing training labels, splitting clean labels before augmentation. Training had 4,212 pairs from 351 labels, development 1,248 cases from 104 labels, and the reserved evaluation 600 cases from 50 labels. Profile was excluded from augmentation training and development selection. The baseline may have seen the clean labels during its original training: this is a reserved test of new spelling robustness, not fresh evidence of semantic generalization.
14
+
15
+ Eight fixed word-table scales and eleven consistency-training runs were compared. Training aligned noisy and clean embeddings, retained frozen clean representations, and preserved existing semantic objectives. Seven initial runs and four smaller-learning-rate follow-ups failed to beat the eligible scale adjustment. Stronger spelling improvements commonly reduced semantic quality or coverage. One selected floating-point checkpoint failed its gates after quantization and was rejected.
16
+
17
+ Selection used the mean of short-label and overall development top-one accuracy, subject to fixed semantic, coverage, no-match, and XFCE regression floors. The 0.85 scale also passed all seven existing release regression gates before the reserved spelling test was opened. No gate was relaxed to promote it.
18
+
19
+ ## Reserved evaluation
20
+
21
+ | Measure | Previous weights | Rebalanced weights |
22
+ | --- | ---: | ---: |
23
+ | Top-one accuracy, 600 synthetic typo cases | 64.33% | 67.33% |
24
+ | Top-one accuracy, 144 short-label cases | 42.36% | 45.83% |
25
+ | Raw int8 payload | 32 KiB | 32 KiB |
26
+
27
+ The rebalanced model fixes 20 previously wrong cases and regresses 2 previously correct cases. The overall gain is **3.0 percentage points**. A bootstrap clustered by clean label gives a 95% interval of **[1.5, 4.67] percentage points**. NumPy and the actual TypeScript runtime agree on these aggregate held-out values. They were evaluated as the same frozen event, not independent replications. Small floating-point differences affect a few near-tied development rankings; both implementations are recorded.
28
+
29
+ The known demo checks still return Members for `coworkers` and Profile for `my information`. However, `profle` still returns Invoices in raw inference, with cosine 0.7021. The scale adjustment is therefore a measured broad improvement, not resolution of the reported model bug.
30
+
31
+ ## What would constitute a fuller fix?
32
+
33
+ These results support further investigation of how the model allocates capacity between word and character features. A future experiment could constrain word-vector dominance while teaching the character branch semantic retrieval, potentially changing pooling while keeping the same weight budget. That proposal is untested. Simply increasing epochs, strengthening consistency loss, or reducing word influence further did not satisfy the current gates.
34
+
35
+ No production-ready or best-in-class claim follows from this synthetic spelling benchmark. A stronger model needs retained semantic accuracy, reliable spelling on short unseen labels, reviewed ambiguous/no-match queries, and a new reserved evaluation. This spelling holdout is now consumed and cannot be represented as untouched in later selection.
36
+
37
+ ## Reproduction and evidence
38
+
39
+ - `training/prepare_typo.py`, `data/typo/manifest.json`: deterministic perturbations, label separation and source hashes. Ambiguous shared corruptions and spellings equal to another clean label are excluded. Judgments are synthetic, not human-reviewed.
40
+ - `configs/typo-training*.json`, `training/experiment_typo.py`, `eval/typo-training*.json`: bounded training runs and exact-export gates. [Training findings](typo-training.md).
41
+ - `eval/typo-word-scaling*.json`: initial sweep, bounded follow-up and exact exported-model verification. [Scaling findings](typo-word-scaling.md).
42
+ - `eval/typo-selection.json`, `eval/typo-holdout-access.json`: frozen payload and manifest hashes, candidate selection and one-time evaluation receipt.
43
+ - `eval/typo-evaluation.json`, `eval/typo-runtime-holdout.json`: reserved NumPy and actual serving-runtime evidence. `eval/typo-known-diagnostic.json` separately records the already-known failures.
44
+ - `eval/typo-release-regressions.json`: all seven existing release regression gates passed. Original consumed semantic tests were not reopened.
45
+
46
+ The training-only consistency approach was informed by [CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos](https://arxiv.org/abs/2204.00716). Our small hashed encoder and synthetic label task differ substantially from that work; its results are not evidence that this implementation will match them.
docs/typo-word-scaling.md ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Fixed word-table scaling ablation
2
+
3
+ The five predeclared scales do not produce an eligible replacement for `navigation-align0p5-seed29`. Reducing the word-table contribution improves synthetic typo retrieval while reducing broader semantic coverage. The unchanged scale is the only arm that passes every existing floor; no new model artifact was exported.
4
+
5
+ | Word scale | Typo top-1 | Typo top-5 | Expanded semantic nDCG@5 | Expanded coverage | Expanded no-match | Expanded overall nDCG@5 | XFCE semantic raw nDCG@5 | All floors |
6
+ | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
7
+ | 1.00 | .6827 | .8502 | .4567 | .4866 | .0398 | .4577 | .5313 | Pass |
8
+ | 0.75 | .7276 | .8790 | .4532 | .4785 | .0455 | .4542 | .5327 | Fail: coverage |
9
+ | 0.50 | .7804 | .9183 | .4180 | .4463 | .0455 | .4190 | .5116 | Fail |
10
+ | 0.25 | .8790 | .9639 | .2536 | .2791 | .0398 | .2550 | .5002 | Fail |
11
+ | 0.00 | .9319 | .9856 | .1063 | .1261 | .0199 | .1079 | .4694 | Fail |
12
+
13
+ The frozen floors were expanded semantic nDCG@5 ≥ .45, semantic coverage ≥ .48, no-match semantic rate ≤ .05, overall nDCG@5 ≥ .447708, and XFCE semantic raw nDCG@5 ≥ .5113. Each ratio used its own cutoff fitted only to the expanded calibration partition. Typo metrics use full raw cosine order without lexical ranking or a cutoff. The 0.75 arm misses the coverage floor by .001467; the floor was not relaxed to accept it.
14
+
15
+ The NumPy evaluator multiplies the dequantized word table by each fixed ratio and leaves the character table unchanged. Existing query normalization and candidate label/alias/context composition remain identical. This suggests the two feature branches trade typo robustness against semantic retrieval in this checkpoint, rather than establishing that character-only inference is generally better. Retraining may find a better tradeoff; this experiment provides no evidence that it will.
16
+
17
+ Only `data/typo/dev.jsonl`, `data/expanded/{calibration,dev}.jsonl`, and `data/navigation-splits/dev.jsonl` were read. There was no access to typo holdout, diagnostic, custom queries, original test, or navigation holdout. Typo judgments are synthetic one-edit origins; XFCE judgments contain documented positives with other labels unjudged. Results are development evidence, not reviewed independent quality claims.
18
+
19
+ Positive scaling can be represented through the word tensor's float scale with unchanged int8 codes and the same 32,768-byte payload. The zero arm would require zero word codes and a valid positive scale, still occupying the same payload size. No exported artifact, runtime, demo, or dataset was changed.
20
+
21
+ Reproduce with `.venv/bin/python -m eval.typo-word-scaling`. [The machine-readable report](../eval/typo-word-scaling.json) records source hashes, model hash, exact metrics, and pass/fail for every floor.
22
+
23
+ ## Authorized bounded follow-up
24
+
25
+ After the initial scan, a separate follow-up froze three intermediate ratios (.8, .85, .9) plus the unchanged control. It added the requested short-label slice (at most two whitespace-separated tokens in the clean label, 300 development queries) and selected the highest mean of short-label and all-query top-1 accuracy among arms passing every floor. The initial report remains unchanged.
26
+
27
+ | Word scale | All typo top-1 | Short-label top-1 | Mean selection score | Expanded semantic nDCG@5 | Expanded coverage | XFCE semantic raw nDCG@5 | All floors |
28
+ | ---: | ---: | ---: | ---: | ---: | ---: | ---: | :--- |
29
+ | 1.00 | .6827 | .3467 | .5147 | .4567 | .4866 | .5313 | Pass |
30
+ | 0.80 | .7147 | .4133 | .5640 | .4457 | .4714 | .5392 | Fail |
31
+ | 0.85 | .7035 | .3900 | .5468 | .4654 | .4946 | .5300 | Pass |
32
+ | 0.90 | .6979 | .3900 | .5440 | .4699 | .5018 | .5324 | Pass |
33
+
34
+ The selected .85 artifact is `packages/model/experiments/navigation-align0p5-seed29-word085`. Only its word tensor's dequantization scale changes, stored as float32 and matching little-endian bytes; the 32,768 weight bytes remain byte-identical. Its manifest hash therefore distinguishes it from its parent even though the payload hash is shared. The existing TypeScript loader accepts the exported artifact.
35
+
36
+ Reloading the exact exported metadata and bytes passes every floor: semantic nDCG .465414, coverage .494633, no-match rate .039773, overall nDCG .466369, and XFCE .529963. Exact exported all-query typo top-1 is .705128; short-label top-1 is .396667. Two additional short-label queries rank correctly in the exported evaluation. The export uses a different float32 operation order (scaling before decoding versus scaling decoded values), which can change near-tied ordering; selection evidence and exported verification are reported separately. There is no further ratio tuning.
37
+
38
+ This is a development-selected candidate for comparison with the parallel training experiments, not a deployment or independent holdout result. No demo/runtime change, diagnostic access, custom-query scoring, or holdout access occurred. [Follow-up results](../eval/typo-word-scaling-followup.json) and [exact export verification](../eval/typo-word-scaling-export-verification.json) retain their own source hashes. Reproduce with `.venv/bin/python -m eval.typo-word-scaling-followup` followed by `.venv/bin/python -m eval.export-word-scaling`.
docs/verification.md ADDED
@@ -0,0 +1,39 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Release evidence and limits
2
+
3
+ This is the baseline milestone allowed by the specification's semantic stop condition, not completion of its learned-search release criteria.
4
+
5
+ - TypeScript contracts cover exact/alias precedence, every lexical subclass, OSA transpositions and safeguards, Unicode and symbols, input bounds, snapshots, abort/concurrency/disposal, ties, empty queries, fallback metadata, and 200 seeded adversarial menus.
6
+ - Python/TypeScript feature IDs and normalization agree on golden fixtures. Both lexical evaluation baselines agree with the shipping engine on all development fixtures plus adversarial camelCase cases.
7
+ - Five Python export-format tests cover quantizer rounding/extrema, int6 packing, reserved codes, trailing bits and malformed byte lengths. The research exporter is not a validated browser asset loader.
8
+ - Revised Chrome production-build checks verify actual manifest/weight fetches, model scores for six queries against the real encoder, arbitrary edited candidates, invalid JSON, literal HTML-like labels, 390px containment, and Auto/Light/Dark persistence and system changes. Blocking weights shows Model unavailable; substituting valid zero weights eliminates model results while lexical search still works. Zero requests were observed during query/edit interactions after load. See `bench/reports/demo.json`.
9
+ - The custom experimental CPU runtime verifies SHA-256 and strict int8 tensor metadata. Exported NumPy fixture parity has maximum embedding-component error 8.94e-8 and cosine error 1.70e-7. Tests also check malformed assets, mutation isolation, actual weight perturbation, and alias/context composition. This verifies execution, not model usefulness.
10
+ - Full lexical search ran in Chrome 153.0.8010.36 and Playwright WebKit 26.0 on Apple M4 Max, macOS/Darwin 27.0.0. At 1,000/5,000 candidates Chrome p95 was 2.1/10.1 ms; WebKit was 2/11 ms. The 50,000 stretch workload reached 62.1/76 ms p95. See all samples and indexing timings in `bench/reports/browser.json`.
11
+ - Measurements used 30 warmups and 200 seeded randomized queries at each size. Labels have repeated synthetic families with suffixes, aliases, and contexts; this is not representative product corpus performance. Timers can round short calls to zero. Power state was not controlled. Heap reporting was available only in Chrome and is a whole-page estimate after the sweep, not per-index allocation accounting.
12
+ - The local browser benchmark used installed Chrome (`CHROME_CHANNEL=chrome`) and WebKit revision 2359 (`WEBKIT_EXECUTABLE` override), recorded in the report. Default reproduction downloads Playwright's pinned browsers; results will differ by version/device.
13
+ - The size report separately compresses all emitted resources with Node's zlib Brotli quality11/window22 and gzip level9. The lexical library meets the 8KiB target. The revised demo transfer includes the actual experimental manifest and int8 model weights; no source maps are served. Historical latency results above measure only the lexical engine.
14
+
15
+ Not claimed: reviewed held-out semantic quality, actual WebGPU kernels, device-loss recovery, semantic latency guarantees, or production model readiness. Later updates below add experimental-model parity, measured CPU transfer budgets and an offline general embedding reference; none establishes a production-ready semantic model.
16
+
17
+ ## Expanded-data update (2026-09-13)
18
+
19
+ The active demo artifact is now `packages/model/candidate/` (`expanded-more-data-pooled-seed17`), with the pilot preserved for comparisons. The full pipeline passed 15 Python tests covering exporter contracts, data separation and source-label quality floors; JavaScript tests additionally cover candidate NumPy parity and the training-seen Members/Profile sanity cases. Browser verification checked actual model bytes, arbitrary candidate edits, no query transmission, themes, and missing/zero weights with the new artifact. See `docs/expanded-evaluation.md` for numerical development gates, transfer limits, and the still-unmet reviewed-final-test requirement. These results supersede the original model-quality stop finding for this expanded experiment only.
20
+
21
+ ## Optimization follow-up (2026-09-13)
22
+
23
+ The same live weights were retained after 39 additional training runs and 90 calibration configurations. The strongest same-size research candidate improved development but did not improve calibrated reserved-test semantics. Its final synthetic test was opened once after model/manifest selection was frozen; the test is now consumed and remains unreviewed. See [the optimization decision](optimization-experiments.md) and `eval/optimized-evaluation.json`.
24
+
25
+ All 37 portable Python tests and 24 JavaScript tests pass, together with types, production build, complete CPU transfer-size gates, and real Chrome demo inference/failure/privacy/theme checks. The standalone research runtime also passes independent NumPy parity on real weights and Unicode/ordered-feature fixtures. Neither its runtime nor its weights enter the live demo import graph. The original CPU demo is 36,326 bytes Brotli; the isolated research build is 36,243 bytes. CI enforces both 50 KiB CPU budgets and the lexical 8 KiB limit. No training or reserved-test access occurs in CI.
26
+
27
+ ## Navigation and inference reuse (2026-09-13)
28
+
29
+ The experimental demo now uses `navigation-align0p5-seed29`, retaining 32,768 raw weight bytes. Twelve navigation/data follow-ups and nine collision-capacity runs preceded a frozen one-time GNOME provider evaluation. The selected model improved raw semantic known-positive nDCG from 0.1016 to 0.1705 and passed all seven existing regression floors. It still trails the train-paraphrase TF-IDF reference on this measure and does not establish safe semantic abstention. See [the complete evaluation and release decision](navigation-evaluation.md).
30
+
31
+ Prepared candidate embeddings reuse menu vectors across queries with exact reference score parity. The browser report separates synchronous setup and warm query latency; see [the prepared API and measurements](prepared-model-index.md). Portable tests cover navigation metadata parsing, positive-only augmentation, evaluation judgments, and runtime isolation. CI does not train models or reopen the provider holdout.
32
+
33
+ ## Typo-result precedence fix
34
+
35
+ Main demo results now use exact/fuzzy lexical results whenever present and otherwise show explicitly labeled real model suggestions. They do not pad successful text matches with unrelated cosine neighbors. Raw scores remain separately inspectable. Real-weight regression tests protect `profle → Profile`, exact labels, prefixes and aliases, including a fixture where the model itself ranks the wrong destination first. Browser checks cover main-result precedence and typo search with unavailable or zeroed weights. No weights were retrained or changed by this fix.
36
+
37
+ ## Weight-level typo follow-up
38
+
39
+ The demo uses a 0.85 word-table scale with unchanged int8 payload size. Eight scaling values and eleven bounded retraining runs were compared before freezing the selected manifest. On 600 reserved synthetic typo cases, top-one accuracy increased from 64.33% to 67.33%; the actual TypeScript runtime confirms the same aggregate result. All seven existing release regression gates pass. The raw `profle` failure remains, so this is a partial robustness improvement. See [the complete experiment](typo-weight-experiments.md) for frozen hashes, uncertainty, source boundaries and rejected runs. No original consumed semantic holdout was reopened.
eval/typo-evaluation.json ADDED
The diff for this file is too large to render. See raw diff
 
eval/typo-release-regressions.json ADDED
The diff for this file is too large to render. See raw diff
 
eval/typo-runtime-holdout.json ADDED
The diff for this file is too large to render. See raw diff
 
eval/typo-selection.json ADDED
@@ -0,0 +1,26 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "directory": "packages/model/experiments/navigation-align0p5-seed29-word085",
3
+ "modelId": "navigation-align0p5-seed29-word085",
4
+ "manifestSha256": "019b113275824658b6ac18aa40afb7a28187ce5855620f5be393bdeda2e24388",
5
+ "payloadSha256": "fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2",
6
+ "payloadBytes": 32768,
7
+ "featureVersion": "gpu-search-features-v1",
8
+ "dimension": 16,
9
+ "quantizationBits": [
10
+ 8
11
+ ],
12
+ "selectionFrozen": true,
13
+ "selectionBasis": "Highest exact-export mean(short/all) typo DEV top1 among8scalevalues and11trainingruns subject expanded/XFCE gates; all7release regressiongates passed. Selection completed before reserved spelling evaluation.",
14
+ "selectionScore": 0.5508974358974359,
15
+ "runtimeDevSelectionScore": 0.5529647435897436,
16
+ "evaluationEvent": "One frozen synthetic spelling evaluation with NumPy and actual TypeScript runtime verification; neither is an independent replication of the other. Profile knownbug separately diagnostic only.",
17
+ "trainingReports": [
18
+ "eval/typo-training.json",
19
+ "eval/typo-training-lowlr.json"
20
+ ],
21
+ "scalingReports": [
22
+ "eval/typo-word-scaling.json",
23
+ "eval/typo-word-scaling-followup.json",
24
+ "eval/typo-word-scaling-export-verification.json"
25
+ ]
26
+ }
manifest.json ADDED
@@ -0,0 +1,43 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "formatVersion": "gpu-search-experimental-v1",
3
+ "featureVersion": "gpu-search-features-v1",
4
+ "normalizationVersion": "nfkc-ascii-v1",
5
+ "candidateCompositionVersion": "mean-alias-context025-v1",
6
+ "architecture": "pooled",
7
+ "dimension": 16,
8
+ "featureFamily": "both",
9
+ "modelId": "navigation-align0p5-seed29-word085",
10
+ "byteOrder": "little-endian",
11
+ "tensors": [
12
+ {
13
+ "name": "word",
14
+ "shape": [
15
+ 1024,
16
+ 16
17
+ ],
18
+ "elementCount": 16384,
19
+ "bits": 8,
20
+ "scale": 0.007353196851909161,
21
+ "scaleF32LE": "16f3f03b",
22
+ "byteOffset": 0,
23
+ "byteLength": 16384
24
+ },
25
+ {
26
+ "name": "char",
27
+ "shape": [
28
+ 1024,
29
+ 16
30
+ ],
31
+ "elementCount": 16384,
32
+ "bits": 8,
33
+ "scale": 0.008905642665922642,
34
+ "scaleF32LE": "f9e8113c",
35
+ "byteOffset": 16384,
36
+ "byteLength": 16384
37
+ }
38
+ ],
39
+ "payloadSha256": "fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2",
40
+ "payloadBytes": 32768,
41
+ "validatedSemanticCutoff": null,
42
+ "status": "EXPERIMENTAL ONLY; release quality gate not passed; do not load in default search"
43
+ }
package.json ADDED
@@ -0,0 +1,5 @@
 
 
 
 
 
 
1
+ {
2
+ "name": "gpu-search-model",
3
+ "private": true,
4
+ "type": "module"
5
+ }
packages/core/src/index.ts ADDED
@@ -0,0 +1,142 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ /** Deterministic, local retrieval. No assets or GPU are accessed at import time. */
2
+ export type Backend = 'cpu' | 'webgpu' | 'auto';
3
+ export type Candidate = { id: string; label: string; aliases?: readonly string[]; context?: string };
4
+ export type Match = 'exact' | 'lexical' | 'semantic';
5
+ export type SearchResult = {
6
+ id: string; label: string; match: Match;
7
+ reason: 'label-equality' | 'alias-equality' | 'prefix' | 'token-prefix' | 'acronym' | 'typo' | 'concept';
8
+ matchedField: 'label' | 'alias' | 'semantic';
9
+ lexicalScore?: number; semanticScore?: number;
10
+ };
11
+ export type SearchResponse = { results: SearchResult[]; backend: 'cpu' | 'webgpu' | 'lexical'; degraded: boolean };
12
+ export type SearchOptions = { limit?: number; signal?: AbortSignal };
13
+ export type IndexOptions = { backend?: Backend; semantic?: boolean };
14
+ export type SearchIndex = { search(query: string, options?: SearchOptions): Promise<SearchResponse>; dispose(): void };
15
+
16
+ /** Build constants, not model quality or latency guarantees. */
17
+ export const INPUT_LIMITS = Object.freeze({ candidates: 50_000, label: 256, query: 256, aliases: 8, alias: 128, context: 256 });
18
+ export class SearchInputError extends TypeError {
19
+ readonly code = 'ERR_SEARCH_INPUT';
20
+ constructor(message: string, readonly path: string) { super(`${path}: ${message}`); this.name = 'SearchInputError'; }
21
+ }
22
+ export class IndexDisposedError extends Error {
23
+ readonly code = 'ERR_INDEX_DISPOSED';
24
+ constructor() { super('Search index has been disposed'); this.name = 'IndexDisposedError'; }
25
+ }
26
+
27
+ const asciiLower = (text: string): string => text.replace(/[A-Z]/g, c => c.toLowerCase());
28
+ /** NFKC, ASCII case only, Unicode White_Space collapse. Punctuation is retained. */
29
+ export const normalizeKey = (text: string): string => asciiLower(text.normalize('NFKC')).replace(/\p{White_Space}+/gu, ' ').replace(/^ +| +$/g, '');
30
+ /** Token splitting is deliberately separate from strict equality. */
31
+ export function tokenize(text: string): string[] {
32
+ return asciiLower(text.normalize('NFKC').replace(/([a-z0-9])([A-Z])/g, '$1 $2'))
33
+ .split(/[\p{White_Space}_-]+/u).filter(Boolean);
34
+ }
35
+ const length = (text: string): number => Array.from(text).length;
36
+ const nonSpaceLength = (text: string): number => length(text.replace(/\p{White_Space}/gu, ''));
37
+ function validateText(value: unknown, path: string, max: number, nonempty = false): asserts value is string {
38
+ if (typeof value !== 'string') throw new SearchInputError('expected a string', path);
39
+ if (/[\uD800-\uDBFF](?![\uDC00-\uDFFF])|(?<![\uD800-\uDBFF])[\uDC00-\uDFFF]/u.test(value)) throw new SearchInputError('unpaired surrogate is not a Unicode scalar', path);
40
+ if (length(value) > max) throw new SearchInputError(`exceeds ${max} Unicode scalars`, path);
41
+ if (nonempty && !normalizeKey(value)) throw new SearchInputError('must not be empty', path);
42
+ }
43
+ type Field = { key: string; tokens: string[]; kind: 'label' | 'alias' };
44
+ type Entry = { id: string; label: string; fields: Field[]; order: number };
45
+ type Ranked = { result: SearchResult; tier: number; subclass: number; score: number; order: number };
46
+ function field(text: string, kind: Field['kind']): Field { return { key: normalizeKey(text), tokens: tokenize(text), kind }; }
47
+
48
+ /** Optimal string alignment distance, including adjacent transpositions (not unrestricted DL). */
49
+ export function osaDistance(left: string, right: string): number {
50
+ const a = Array.from(left), b = Array.from(right);
51
+ let previous = Array.from({ length: b.length + 1 }, (_, i) => i);
52
+ let beforePrevious = previous;
53
+ for (let i = 1; i <= a.length; i++) {
54
+ const current = new Array<number>(b.length + 1); current[0] = i;
55
+ for (let j = 1; j <= b.length; j++) {
56
+ current[j] = Math.min(previous[j]! + 1, current[j - 1]! + 1, previous[j - 1]! + (a[i - 1] === b[j - 1] ? 0 : 1));
57
+ if (i > 1 && j > 1 && a[i - 1] === b[j - 2] && a[i - 2] === b[j - 1]) current[j] = Math.min(current[j]!, beforePrevious[j - 2]! + 1);
58
+ }
59
+ beforePrevious = previous; previous = current;
60
+ }
61
+ return previous[b.length]!;
62
+ }
63
+ function qualifies(query: string, tokens: string[], candidate: Field): { subclass: number; score: number; reason: SearchResult['reason'] } | undefined {
64
+ const qLength = length(query), fLength = length(candidate.key);
65
+ if (candidate.kind === 'alias' && query === candidate.key) return { subclass: 5, score: 1, reason: 'alias-equality' };
66
+ const prefixScore = Math.min(1, nonSpaceLength(query) / Math.max(1, nonSpaceLength(candidate.key)));
67
+ if (qLength >= 2 && candidate.key.startsWith(query)) return { subclass: 4, score: prefixScore, reason: 'prefix' };
68
+ if (nonSpaceLength(query) >= 3 && tokens.length > 0) {
69
+ let next = 0;
70
+ for (const token of candidate.tokens) if (next < tokens.length && token.startsWith(tokens[next]!)) next++;
71
+ if (next === tokens.length) return { subclass: 3, score: prefixScore, reason: 'token-prefix' };
72
+ }
73
+ if (qLength >= 2 && qLength <= 6 && candidate.tokens.length >= 2 && query === candidate.tokens.map(t => Array.from(t)[0]).join('')) return { subclass: 2, score: 1, reason: 'acronym' };
74
+ const longest = Math.max(qLength, fLength), maxDistance = longest >= 8 ? 2 : 1;
75
+ if (qLength >= 4 && fLength >= 4 && Math.abs(qLength - fLength) <= maxDistance) {
76
+ const distance = osaDistance(query, candidate.key);
77
+ if (distance <= maxDistance && distance / longest <= 0.25) return { subclass: 1, score: 1 - distance / longest, reason: 'typo' };
78
+ }
79
+ return undefined;
80
+ }
81
+ function compare(a: Ranked, b: Ranked): number { return b.tier - a.tier || b.subclass - a.subclass || b.score - a.score || a.order - b.order; }
82
+
83
+ export async function createIndex(candidates: readonly Candidate[], options: IndexOptions = {}): Promise<SearchIndex> {
84
+ if (!Array.isArray(candidates)) throw new SearchInputError('expected an array', 'candidates');
85
+ if (candidates.length > INPUT_LIMITS.candidates) throw new SearchInputError('too many candidates', 'candidates');
86
+ if (!options || typeof options !== 'object' || Array.isArray(options)) throw new SearchInputError('expected an options object', 'options');
87
+ if (options.backend !== undefined && !['auto', 'cpu', 'webgpu'].includes(options.backend)) throw new SearchInputError('expected auto, cpu, or webgpu', 'backend');
88
+ if (options.semantic !== undefined && typeof options.semantic !== 'boolean') throw new SearchInputError('expected a boolean', 'semantic');
89
+ const degraded = options.semantic !== false;
90
+ const ids = new Set<string>();
91
+ let entries: Entry[] = [];
92
+ // A plain loop also validates sparse array holes.
93
+ for (let order = 0; order < candidates.length; order++) {
94
+ const item = candidates[order], path = `candidates[${order}]`;
95
+ if (!item || typeof item !== 'object' || Array.isArray(item)) throw new SearchInputError('expected a candidate record', path);
96
+ validateText(item.id, `${path}.id`, Number.MAX_SAFE_INTEGER, true);
97
+ validateText(item.label, `${path}.label`, INPUT_LIMITS.label, true);
98
+ if (ids.has(item.id)) throw new SearchInputError('duplicate ID', `${path}.id`);
99
+ ids.add(item.id);
100
+ const fields = [field(item.label, 'label')];
101
+ if (item.aliases !== undefined) {
102
+ if (!Array.isArray(item.aliases) || item.aliases.length > INPUT_LIMITS.aliases) throw new SearchInputError('expected at most 8 aliases', `${path}.aliases`);
103
+ for (let i = 0; i < item.aliases.length; i++) {
104
+ const alias = item.aliases[i]; validateText(alias, `${path}.aliases[${i}]`, INPUT_LIMITS.alias, true); fields.push(field(alias, 'alias'));
105
+ }
106
+ }
107
+ if (item.context !== undefined) validateText(item.context, `${path}.context`, INPUT_LIMITS.context);
108
+ entries.push({ id: item.id, label: item.label, fields, order });
109
+ }
110
+ let disposed = false;
111
+ return {
112
+ async search(query, searchOptions = {}) {
113
+ if (disposed) throw new IndexDisposedError();
114
+ validateText(query, 'query', INPUT_LIMITS.query);
115
+ if (!searchOptions || typeof searchOptions !== 'object' || Array.isArray(searchOptions)) throw new SearchInputError('expected an options object', 'searchOptions');
116
+ const { limit = 10, signal } = searchOptions;
117
+ if (!Number.isSafeInteger(limit) || limit < 0) throw new SearchInputError('expected a nonnegative safe integer', 'limit');
118
+ const check = () => { if (disposed) throw new IndexDisposedError(); if (signal?.aborted) throw signal.reason ?? new DOMException('Search aborted', 'AbortError'); };
119
+ check();
120
+ const key = normalizeKey(query), tokens = tokenize(query), ranked: Ranked[] = [];
121
+ if (key && limit) for (const entry of entries) {
122
+ if (entry.fields[0]!.key === key) {
123
+ ranked.push({ result: { id: entry.id, label: entry.label, match: 'exact', reason: 'label-equality', matchedField: 'label' }, tier: 3, subclass: 0, score: 1, order: entry.order });
124
+ continue;
125
+ }
126
+ let best: Ranked | undefined;
127
+ for (const candidateField of entry.fields) {
128
+ const found = qualifies(key, tokens, candidateField);
129
+ if (!found) continue;
130
+ const item: Ranked = { result: { id: entry.id, label: entry.label, match: 'lexical', reason: found.reason, matchedField: candidateField.kind, lexicalScore: found.score }, tier: 2, subclass: found.subclass, score: found.score, order: entry.order };
131
+ if (!best || compare(item, best) < 0) best = item;
132
+ }
133
+ if (best) ranked.push(best);
134
+ }
135
+ ranked.sort(compare);
136
+ // Yield before delivery so immediate abort/dispose cannot deliver stale work.
137
+ await Promise.resolve(); check();
138
+ return { results: ranked.slice(0, limit).map(item => item.result), backend: 'lexical', degraded };
139
+ },
140
+ dispose() { disposed = true; entries = []; },
141
+ };
142
+ }
packages/core/src/lexical.ts ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ import { createIndex as createBaseIndex, type Candidate, type SearchIndex } from './index.js';
2
+ export * from './index.js';
3
+ /** Explicit lexical-only entry. No missing-model degradation is reported. */
4
+ export function createIndex(candidates: readonly Candidate[]): Promise<SearchIndex> {
5
+ return createBaseIndex(candidates, { semantic: false });
6
+ }
packages/model/feature-spec.json ADDED
@@ -0,0 +1 @@
 
 
1
+ {"version":"gpu-search-features-v1","normalizationVersion":"nfkc-ascii-v1","compositionVersion":"mean-alias-context025-v1","hash":{"name":"FNV-1a-32","offset":2166136261,"multiplier":16777619,"buckets":1024},"wordEncoding":"UTF-8 of w: followed by token","characterEncoding":"UTF-8 c: followed by tagged elements: scalar = byte 0 then uint32 little-endian; BOS = byte 1; EOS = byte 2","ngrams":[2,3,4],"order":"token, ngram length, position","maxTokens":32,"maxTokenScalars":64,"maxCharacterFeatures":512,"normalization":"NFKC; ASCII uppercase to lowercase; Unicode whitespace collapse and trim","tokens":"NFKC; insert space between ASCII [a-z0-9] and [A-Z]; ASCII lowercase; split Unicode whitespace, underscore, hyphen","pooling":"0.5 * (word family mean + character family mean); repeats retained; empty family zero","candidateComposition":"mean(valid label and alias normalized vectors) + 0.25 * valid context normalized vector, then L2 normalize","epsilon":1e-8}
packages/model/features.ts ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import { normalizeKey, tokenize } from '../core/src/index.js';
2
+ export { normalizeKey, tokenize };
3
+ const utf8 = new TextEncoder();
4
+ export function hash(bytes:number[]|Uint8Array):number {let h=2166136261;for(const b of bytes)h=Math.imul(h^b,16777619)>>>0;return h&1023}
5
+ export function features(text:string){
6
+ const all=tokenize(text),tokens=all.slice(0,32).map(t=>Array.from(t).slice(0,64).join(''));
7
+ const wordIds:number[]=[],charIds:number[]=[];let total=0;
8
+ for(const token of tokens){
9
+ wordIds.push(hash(utf8.encode('w:'+token)));
10
+ const elements=[[1],...Array.from(token).map(c=>{const n=c.codePointAt(0)!;return [0,n&255,(n>>>8)&255,(n>>>16)&255,(n>>>24)&255]}),[2]];
11
+ for(const n of [2,3,4])for(let i=0;i<=elements.length-n;i++){total++;if(charIds.length<512)charIds.push(hash([99,58,...elements.slice(i,i+n).flat()]))}
12
+ }
13
+ return {wordIds,charIds,wordCount:wordIds.length,charCount:charIds.length,truncated:all.length>32||all.slice(0,32).some(t=>Array.from(t).length>64)||total>512};
14
+ }
packages/model/runtime.ts ADDED
@@ -0,0 +1,170 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import { features } from './features.js';
2
+ import { INPUT_LIMITS, SearchInputError, normalizeKey, type Candidate } from '../core/src/index.js';
3
+
4
+ export type ModelScore = { id: string; label: string; score: number };
5
+ export type PreparedModelIndex = {
6
+ readonly size: number;
7
+ /** Raw cosine ranking using cached candidate vectors; no relevance cutoff is implied. */
8
+ score(query: string): ModelScore[];
9
+ dispose(): void;
10
+ };
11
+
12
+ export type Model = {
13
+ readonly id: string;
14
+ readonly hash: string;
15
+ encode(text: string): Float32Array;
16
+ /** Raw cosine inspection: all nonzero candidate vectors, without a quality cutoff. */
17
+ score(query: string, candidates: readonly Candidate[]): ModelScore[];
18
+ /** Snapshot and encode a bounded candidate menu once, then reuse it across queries. */
19
+ prepare(candidates: readonly Candidate[]): PreparedModelIndex;
20
+ };
21
+ export class ModelAssetError extends Error {
22
+ constructor(message: string) { super(message); this.name = 'ModelAssetError'; }
23
+ }
24
+ export class ModelIndexDisposedError extends Error {
25
+ readonly code = 'ERR_MODEL_INDEX_DISPOSED';
26
+ constructor() { super('Prepared model index has been disposed'); this.name = 'ModelIndexDisposedError'; }
27
+ }
28
+ const DIMENSION = 16, ELEMENTS = 1024 * DIMENSION;
29
+ function requireAsset(condition: unknown, message: string): asserts condition {
30
+ if (!condition) throw new ModelAssetError(message);
31
+ }
32
+ function record(value: unknown): Record<string, unknown> {
33
+ requireAsset(value !== null && typeof value === 'object' && !Array.isArray(value), 'Expected manifest object');
34
+ return value as Record<string, unknown>;
35
+ }
36
+ function normalize(vector: Float32Array): Float32Array {
37
+ let squared = 0;
38
+ for (const value of vector) squared = Math.fround(squared + Math.fround(value * value));
39
+ const norm = Math.fround(Math.sqrt(squared));
40
+ if (!Number.isFinite(norm) || norm < 1e-8) return new Float32Array(DIMENSION);
41
+ return vector.map(value => Math.fround(value / norm));
42
+ }
43
+ function valid(vector: Float32Array): boolean { return vector.some(value => value !== 0); }
44
+ function validateText(value: unknown, path: string, limit: number, nonempty = false): asserts value is string {
45
+ if (typeof value !== 'string') throw new SearchInputError('expected a string', path);
46
+ if (/[\uD800-\uDBFF](?![\uDC00-\uDFFF])|(?<![\uD800-\uDBFF])[\uDC00-\uDFFF]/u.test(value)) throw new SearchInputError('unpaired Unicode surrogate', path);
47
+ if (Array.from(value).length > limit) throw new SearchInputError(`exceeds ${limit} Unicode scalars`, path);
48
+ if (nonempty && !normalizeKey(value)) throw new SearchInputError('must not be empty', path);
49
+ }
50
+
51
+ function snapshot(candidates: readonly Candidate[]): Candidate[] {
52
+ if (!Array.isArray(candidates) || candidates.length > INPUT_LIMITS.candidates) throw new SearchInputError('expected a bounded candidate array', 'candidates');
53
+ const result: Candidate[] = [], ids = new Set<string>();
54
+ for (let i = 0; i < candidates.length; i++) {
55
+ const candidate = candidates[i], path = `candidates[${i}]`;
56
+ if (!candidate || typeof candidate !== 'object' || Array.isArray(candidate)) throw new SearchInputError('expected a candidate record', path);
57
+ const { id, label, aliases, context } = candidate;
58
+ validateText(id, `${path}.id`, Number.MAX_SAFE_INTEGER, true); validateText(label, `${path}.label`, INPUT_LIMITS.label, true);
59
+ if (ids.has(id)) throw new SearchInputError('duplicate ID', `${path}.id`);
60
+ ids.add(id);
61
+ let copiedAliases: string[] | undefined;
62
+ if (aliases !== undefined) {
63
+ if (!Array.isArray(aliases) || aliases.length > INPUT_LIMITS.aliases) throw new SearchInputError('expected at most 8 aliases', `${path}.aliases`);
64
+ copiedAliases = [];
65
+ for (let j = 0; j < aliases.length; j++) { const alias = aliases[j]; validateText(alias, `${path}.aliases[${j}]`, INPUT_LIMITS.alias, true); copiedAliases.push(alias); }
66
+ }
67
+ if (context !== undefined) validateText(context, `${path}.context`, INPUT_LIMITS.context);
68
+ result.push({ id, label, aliases: copiedAliases, context });
69
+ }
70
+ return result;
71
+ }
72
+
73
+ /** Loads the actual experimental trained pooled encoder. No model quality approval is implied. */
74
+ export async function loadModel(input: unknown, payload: ArrayBuffer): Promise<Model> {
75
+ const manifest = record(input);
76
+ const versions = {
77
+ formatVersion: 'gpu-search-experimental-v1', featureVersion: 'gpu-search-features-v1',
78
+ normalizationVersion: 'nfkc-ascii-v1', candidateCompositionVersion: 'mean-alias-context025-v1',
79
+ architecture: 'pooled', dimension: DIMENSION, featureFamily: 'both', byteOrder: 'little-endian',
80
+ };
81
+ for (const [key, expected] of Object.entries(versions)) requireAsset(manifest[key] === expected, `Unsupported ${key}`);
82
+ requireAsset(typeof manifest.modelId === 'string' && manifest.modelId.length > 0, 'Missing model ID');
83
+ requireAsset(typeof manifest.payloadSha256 === 'string' && /^[a-f0-9]{64}$/.test(manifest.payloadSha256), 'Invalid SHA-256');
84
+ requireAsset(manifest.validatedSemanticCutoff === null, 'Experimental format must not claim a validated cutoff');
85
+ requireAsset(payload instanceof ArrayBuffer, 'Expected ArrayBuffer payload');
86
+ requireAsset(manifest.payloadBytes === ELEMENTS * 2 && payload.byteLength === manifest.payloadBytes, 'Invalid payload length');
87
+ requireAsset(Array.isArray(manifest.tensors) && manifest.tensors.length === 2, 'Expected word and char tensors');
88
+ const id = manifest.modelId, hash = manifest.payloadSha256;
89
+ // Snapshot before awaiting hashing; callers cannot mutate accepted bytes or metadata during loading.
90
+ const bytes = payload.slice(0), codes = new Int8Array(bytes);
91
+ const weights: Float32Array[] = [];
92
+ for (let i = 0; i < 2; i++) {
93
+ const tensor = record(manifest.tensors[i]);
94
+ requireAsset(tensor.name === ['word', 'char'][i], 'Unexpected tensor order/name');
95
+ requireAsset(Array.isArray(tensor.shape) && tensor.shape.length === 2 && tensor.shape[0] === 1024 && tensor.shape[1] === DIMENSION, 'Unexpected tensor shape');
96
+ requireAsset(tensor.elementCount === ELEMENTS && tensor.bits === 8, 'Unexpected tensor count/quantization');
97
+ requireAsset(tensor.byteOffset === i * ELEMENTS && tensor.byteLength === ELEMENTS, 'Invalid, overlapping, or noncontiguous tensor range');
98
+ requireAsset(typeof tensor.scale === 'number' && Number.isFinite(tensor.scale) && tensor.scale > 0 && Math.fround(tensor.scale) === tensor.scale, 'Invalid f32 scale');
99
+ requireAsset(typeof tensor.scaleF32LE === 'string' && /^[0-9a-f]{8}$/.test(tensor.scaleF32LE), 'Invalid serialized scale');
100
+ const scaleBytes = new Uint8Array(tensor.scaleF32LE.match(/../g)!.map(hex => parseInt(hex, 16)));
101
+ requireAsset(new DataView(scaleBytes.buffer).getFloat32(0, true) === tensor.scale, 'Scale serialization mismatch');
102
+ const decoded = new Float32Array(ELEMENTS);
103
+ for (let j = 0; j < ELEMENTS; j++) {
104
+ const code = codes[i * ELEMENTS + j]!;
105
+ requireAsset(code !== -128, 'Reserved int8 code');
106
+ decoded[j] = Math.fround(code * tensor.scale);
107
+ requireAsset(Number.isFinite(decoded[j]), 'Dequantized weight overflow');
108
+ }
109
+ weights.push(decoded);
110
+ }
111
+ const digest = await globalThis.crypto.subtle.digest('SHA-256', bytes);
112
+ const actualHash = Array.from(new Uint8Array(digest), byte => byte.toString(16).padStart(2, '0')).join('');
113
+ requireAsset(actualHash === hash, 'Payload SHA-256 mismatch');
114
+ function encode(text: string): Float32Array {
115
+ if (typeof text !== 'string') throw new TypeError('Expected text string');
116
+ const extracted = features(text), pooled = new Float32Array(DIMENSION);
117
+ for (const [family, ids] of [extracted.wordIds, extracted.charIds].entries()) {
118
+ if (!ids.length) continue;
119
+ const mean = new Float32Array(DIMENSION), table = weights[family]!;
120
+ // Preserve repeated occurrences and collisions instead of treating IDs as a set.
121
+ for (const bucket of ids) for (let d = 0; d < DIMENSION; d++) mean[d] = Math.fround(mean[d]! + table[bucket * DIMENSION + d]!);
122
+ for (let d = 0; d < DIMENSION; d++) pooled[d] = Math.fround(pooled[d]! + Math.fround(Math.fround(mean[d]! / ids.length) * 0.5));
123
+ }
124
+ return normalize(pooled);
125
+ }
126
+ function compose(candidate: Candidate): Float32Array {
127
+ const vectors = [candidate.label, ...candidate.aliases ?? []].map(encode).filter(valid);
128
+ const composed = new Float32Array(DIMENSION);
129
+ if (vectors.length) {
130
+ for (const vector of vectors) for (let d = 0; d < DIMENSION; d++) composed[d] = Math.fround(composed[d]! + vector[d]!);
131
+ for (let d = 0; d < DIMENSION; d++) composed[d] = Math.fround(composed[d]! / vectors.length);
132
+ }
133
+ if (candidate.context) {
134
+ const context = encode(candidate.context);
135
+ for (let d = 0; d < DIMENSION; d++) composed[d] = Math.fround(composed[d]! + Math.fround(context[d]! * 0.25));
136
+ }
137
+ return normalize(composed);
138
+ }
139
+ function prepare(candidates: readonly Candidate[], checked: boolean): PreparedModelIndex {
140
+ const source = checked ? snapshot(candidates) : candidates;
141
+ const size = source.length;
142
+ let entries: { id: string; label: string; order: number }[] = [];
143
+ let vectors = new Float32Array(size * DIMENSION);
144
+ for (let order = 0; order < source.length; order++) {
145
+ const candidate = source[order]!;
146
+ const vector = compose(candidate);
147
+ if (!valid(vector)) continue;
148
+ vectors.set(vector, order * DIMENSION);
149
+ entries.push({ id: candidate.id, label: candidate.label, order });
150
+ }
151
+ let disposed = false;
152
+ return Object.freeze({ size, score(query: string) {
153
+ if (disposed) throw new ModelIndexDisposedError();
154
+ if (checked) validateText(query, 'query', INPUT_LIMITS.query);
155
+ const queryVector = encode(query);
156
+ if (!valid(queryVector)) return [];
157
+ return entries.map(entry => {
158
+ let score = 0;
159
+ for (let d = 0; d < DIMENSION; d++) score = Math.fround(score + Math.fround(queryVector[d]! * vectors[entry.order * DIMENSION + d]!));
160
+ return { ...entry, score: Math.max(-1, Math.min(1, score)) };
161
+ }).sort((a, b) => b.score - a.score || a.order - b.order).map(({ id, label, score }) => ({ id, label, score }));
162
+ }, dispose() { disposed = true; entries = []; vectors = new Float32Array(0); } });
163
+ }
164
+ return Object.freeze({ id, hash, encode, prepare(candidates: readonly Candidate[]) { return prepare(candidates, true); }, score(query: string, candidates: readonly Candidate[]) {
165
+ // Preserve the original permissive raw-vector inspection API, including empty embeddings.
166
+ if (!valid(encode(query))) return [];
167
+ const temporary = prepare(candidates, false);
168
+ try { return temporary.score(query); } finally { temporary.dispose(); }
169
+ } });
170
+ }
provenance.json ADDED
@@ -0,0 +1,136 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "source_repository": "https://github.com/maxffarrell/gpu-search",
3
+ "source_commit": "0277121c3e4127c96604b25c73a9af627663da80",
4
+ "artifact_path": "packages/model/experiments/navigation-align0p5-seed29-word085",
5
+ "model_id": "navigation-align0p5-seed29-word085",
6
+ "files": {
7
+ "LICENSE": {
8
+ "bytes": 1068,
9
+ "sha256": "ba2f860a3c2cba35f85f21ac65fea46c603900544f90ddad503e3ba9e44dce43"
10
+ },
11
+ "docs/credits.md": {
12
+ "bytes": 4326,
13
+ "sha256": "0a1a1f4fb809b301f457a9a7171cc61496b2fd414d0d9192b624810c40a7f07b"
14
+ },
15
+ "docs/data-sources.md": {
16
+ "bytes": 7453,
17
+ "sha256": "e0a5afe230d9b143d2e7ae255d61bb4373a5e421cb9035c5bdbe13bc92f09613"
18
+ },
19
+ "docs/decisions.md": {
20
+ "bytes": 4207,
21
+ "sha256": "5018474444cf16700c118588c2bb1e736792681f44d8e88b83bb502e2e23a0cb"
22
+ },
23
+ "docs/deployment.md": {
24
+ "bytes": 1131,
25
+ "sha256": "ffe531e995836161100810300e0152de4f358b5d858c3dcf57ffed3f3705052b"
26
+ },
27
+ "docs/expanded-evaluation.md": {
28
+ "bytes": 7203,
29
+ "sha256": "e96ec5d409df3c8afe71f69f0f6b10c80a51f68b72eb9a67d9bef476eec68e46"
30
+ },
31
+ "docs/experiment.md": {
32
+ "bytes": 4339,
33
+ "sha256": "32860d95ae66b3905bc5cb8865b2b9644367f9cb7113b311df318f1b4af153a2"
34
+ },
35
+ "docs/feature-experiments.md": {
36
+ "bytes": 5394,
37
+ "sha256": "fc37c5feb0963bef05b57ee82a2c94f9dc5f1db6c22e92f0f84c732270e7fe12"
38
+ },
39
+ "docs/model-card.md": {
40
+ "bytes": 7753,
41
+ "sha256": "db87bd9a511cea028c44a83c39f7ca1dfe193482219d556125a3d1911ec6e0e6"
42
+ },
43
+ "docs/model-readiness-audit.md": {
44
+ "bytes": 13746,
45
+ "sha256": "545d3e2de892a9d99ed4a9831df3dc4795ffa039efee4264642eb3da69fc7c45"
46
+ },
47
+ "docs/navigation-data.md": {
48
+ "bytes": 5518,
49
+ "sha256": "7e75afcfdb983dcdef3eadaceff9d058a46fe068b070b59837ee9baa5d828bad"
50
+ },
51
+ "docs/navigation-evaluation.md": {
52
+ "bytes": 5315,
53
+ "sha256": "2a6caef8fa554af9cec2469890b639acd13f017eeae2d2b2e0b456f95dba5892"
54
+ },
55
+ "docs/optimization-experiments.md": {
56
+ "bytes": 9237,
57
+ "sha256": "f8c1f726d0ad16d0190cbce65d6cf4ca251073f71514ab9aac6fcefb830e978a"
58
+ },
59
+ "docs/prepared-model-index.md": {
60
+ "bytes": 5460,
61
+ "sha256": "d44ec9178459710eb9972768b62958f15a4d964ef46f82cb4a7e5d7cc741cca4"
62
+ },
63
+ "docs/specification.md": {
64
+ "bytes": 43173,
65
+ "sha256": "c152daa3f1b0467316a9f4733705b084167ff9f85f22e064f717e675ad4d6aff"
66
+ },
67
+ "docs/training-experiments.md": {
68
+ "bytes": 6082,
69
+ "sha256": "dd2c9449cd97df964f4cc8fd6ca1720e343338ec256c3bf16de23ae82558087c"
70
+ },
71
+ "docs/typo-training.md": {
72
+ "bytes": 6668,
73
+ "sha256": "60414fc03803e1639d956ac76cd49cd4313032350639de1f2aec09199c33410b"
74
+ },
75
+ "docs/typo-weight-diagnosis.md": {
76
+ "bytes": 3267,
77
+ "sha256": "5bca5d7ebd5c6dd2a2c8a07ecf5dbbec9325801e367609ef260cb32d6c209a40"
78
+ },
79
+ "docs/typo-weight-experiments.md": {
80
+ "bytes": 5835,
81
+ "sha256": "1528bd9b2066474359fc6fdf213c5a578019d76052df6b0b3b48a30ce01937d1"
82
+ },
83
+ "docs/typo-word-scaling.md": {
84
+ "bytes": 5251,
85
+ "sha256": "1ea491d30987542d6f869ee38f381cd6a38545d95ffd6a1b90ba862bcf9b85cc"
86
+ },
87
+ "docs/verification.md": {
88
+ "bytes": 7540,
89
+ "sha256": "434cb7590d0e640100e774035ba63fb1b2f41c02d54783fb78f988be56bb2f4f"
90
+ },
91
+ "eval/typo-evaluation.json": {
92
+ "bytes": 487169,
93
+ "sha256": "03c9f91e46bc4d8c96f99468322ddd119f70851230e90db99cdb816a31a00690"
94
+ },
95
+ "eval/typo-release-regressions.json": {
96
+ "bytes": 1222280,
97
+ "sha256": "b2bb1cfefe7f4e003f85f4fe9d1e779140ac5dcc6001f99c4c0c6a7635f910d6"
98
+ },
99
+ "eval/typo-runtime-holdout.json": {
100
+ "bytes": 266444,
101
+ "sha256": "7b525b86b2976f18dec50370016168c820db8003f70e8a5a200a32da6a276689"
102
+ },
103
+ "eval/typo-selection.json": {
104
+ "bytes": 1255,
105
+ "sha256": "c52db2365544eb4e175f47015cf7928a8437fa6a9081890488797a83b09fe86a"
106
+ },
107
+ "manifest.json": {
108
+ "bytes": 1126,
109
+ "sha256": "019b113275824658b6ac18aa40afb7a28187ce5855620f5be393bdeda2e24388"
110
+ },
111
+ "packages/core/src/index.ts": {
112
+ "bytes": 9724,
113
+ "sha256": "e8c2a97f838a40dfd86f6604ebea67b21ea79b5e64a6c9d5be805af70c08cdc9"
114
+ },
115
+ "packages/core/src/lexical.ts": {
116
+ "bytes": 348,
117
+ "sha256": "80558807f6e96c762bdbe3b135849272e37dd1a8052ba14b4a6a37de8cde9618"
118
+ },
119
+ "packages/model/feature-spec.json": {
120
+ "bytes": 968,
121
+ "sha256": "f45d4fd99c10dc4520865d19a826453dbb3fd41362323c767ceeb0a3488e63ce"
122
+ },
123
+ "packages/model/features.ts": {
124
+ "bytes": 995,
125
+ "sha256": "48265d3cd89462afa68f88177d3a206d8352d8fdb364865a09a85e6a80f7f100"
126
+ },
127
+ "packages/model/runtime.ts": {
128
+ "bytes": 10709,
129
+ "sha256": "9c5fa93d2024d4e61244ba6bfb3df474a52bbe32b2df50057e6b7f5c57f09e2b"
130
+ },
131
+ "weights.bin": {
132
+ "bytes": 32768,
133
+ "sha256": "fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2"
134
+ }
135
+ }
136
+ }
weights.bin ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2
3
+ size 32768