File size: 4,500 Bytes
546d6ba
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
language:
- en
license: mit
tags:
- gpu-search
- experimental
- embeddings
- int8
- typescript
- navigation
---

# gpu-search

A small experimental encoder for short English software-navigation queries and candidate menus. This repository contains the **current demo checkpoint**, `navigation-align0p5-seed29-word085`, with its original **32,768-byte int8 weights**, manifest, feature specification, and TypeScript CPU runtime.

[Source code](https://github.com/maxffarrell/gpu-search) 路 [Live demo](https://gpu-search.vercel.app) 路 [Detailed model history](docs/model-card.md) 路 [Training-source attribution](docs/data-sources.md)

## Model and intended use

The encoder pools hashed word and character n-gram embeddings into a normalized 16-dimensional vector. Each table has 1,024 rows. Candidate labels and aliases are averaged, with optional context contribution. Ranking uses cosine similarity. The manifest supplies the exact int8 dequantization scales; weights alone are insufficient to reproduce the selected model.

This checkpoint reduces the previous navigation model's word-table scale by 15%. It is intended for research and explicitly labeled model suggestions in command palettes, settings, and navigation menus. The application's deterministic lexical matching takes precedence over model suggestions.

**Experimental:** no learned model has passed the project's production semantic release gate, and `validatedSemanticCutoff` is null. Scores are not confidence probabilities. Raw inference still ranks **Invoices** above **Profile** for `profle`; the demo handles that query with its separate lexical tier. No WebGPU runtime or GPU acceleration is claimed. This custom binary format is loaded by the included runtime, rather than Transformers `AutoModel`.

## Load the model

Download this repository, preserve its directory structure, and use a TypeScript-capable environment with Web Crypto. For example, save this as `example.ts` at the repository root and run `npx tsx example.ts` with Node.js 22.12+:

```ts
import { readFile } from 'node:fs/promises';
import { loadModel } from './packages/model/runtime.ts';

const manifest = JSON.parse(await readFile('./manifest.json', 'utf8'));
const bytes = await readFile('./weights.bin');
const payload = bytes.buffer.slice(bytes.byteOffset, bytes.byteOffset + bytes.byteLength);
const model = await loadModel(manifest, payload);
const index = model.prepare([
  { id: 'profile', label: 'Profile' },
  { id: 'members', label: 'Members' },
  { id: 'invoices', label: 'Invoices' },
]);
console.log(index.score('coworkers'));
index.dispose();
```

Browser applications can fetch `manifest.json` and `weights.bin`, call `loadModel(manifest, await response.arrayBuffer())`, and reuse a prepared index. The loader verifies the payload hash and manifest format. See [prepared-index documentation](docs/prepared-model-index.md) for lifecycle and input limits.

## Evidence and limitations

The selected scale adjustment improved top-one accuracy on 600 reserved **synthetic** typo cases from 64.33% to 67.33%, while retaining the 32 KiB payload. These historical results are documented in [typo-weight experiments](docs/typo-weight-experiments.md), with original evaluation receipts under `eval/`. They do not establish broad semantic generalization, human-reviewed relevance, or reliable abstention. The spelling holdout is now consumed.

Hash collisions, bag-of-features pooling, short unseen labels, ambiguous menus, and action opposites remain limitations. Preserve lexical precedence and treat learned results as suggestions. Documentation includes historical checkpoints; this repository ships only the current demo artifact identified above.

## Provenance and license

Copied from source commit `0277121c3e4127c96604b25c73a9af627663da80`, artifact path `packages/model/experiments/navigation-align0p5-seed29-word085`. `provenance.json` records SHA-256 hashes of copied files. Weight SHA-256: `fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2`.

Project code and original model artifacts are distributed under the source project's MIT license. Training and research sources retain their own terms: CLINC150 (CC BY 3.0), BANKING77 (CC BY 4.0), VS Code (MIT), and navigation metadata from GNOME, KDE, and Xfce. See [data attribution](docs/data-sources.md), [navigation provenance](docs/navigation-data.md), and [credits](docs/credits.md). Raw third-party datasets and other experimental checkpoints are not included.