gpu-search / README.md
maxffarrell's picture
Upload current experimental gpu-search model with runtime and provenance
546d6ba verified
|
Raw History Blame Contribute Delete
4.5 kB
metadata
language:
  - en
license: mit
tags:
  - gpu-search
  - experimental
  - embeddings
  - int8
  - typescript
  - navigation

gpu-search

A small experimental encoder for short English software-navigation queries and candidate menus. This repository contains the current demo checkpoint, navigation-align0p5-seed29-word085, with its original 32,768-byte int8 weights, manifest, feature specification, and TypeScript CPU runtime.

Source code 路 Live demo 路 Detailed model history 路 Training-source attribution

Model and intended use

The encoder pools hashed word and character n-gram embeddings into a normalized 16-dimensional vector. Each table has 1,024 rows. Candidate labels and aliases are averaged, with optional context contribution. Ranking uses cosine similarity. The manifest supplies the exact int8 dequantization scales; weights alone are insufficient to reproduce the selected model.

This checkpoint reduces the previous navigation model's word-table scale by 15%. It is intended for research and explicitly labeled model suggestions in command palettes, settings, and navigation menus. The application's deterministic lexical matching takes precedence over model suggestions.

Experimental: no learned model has passed the project's production semantic release gate, and validatedSemanticCutoff is null. Scores are not confidence probabilities. Raw inference still ranks Invoices above Profile for profle; the demo handles that query with its separate lexical tier. No WebGPU runtime or GPU acceleration is claimed. This custom binary format is loaded by the included runtime, rather than Transformers AutoModel.

Load the model

Download this repository, preserve its directory structure, and use a TypeScript-capable environment with Web Crypto. For example, save this as example.ts at the repository root and run npx tsx example.ts with Node.js 22.12+:

import { readFile } from 'node:fs/promises';
import { loadModel } from './packages/model/runtime.ts';

const manifest = JSON.parse(await readFile('./manifest.json', 'utf8'));
const bytes = await readFile('./weights.bin');
const payload = bytes.buffer.slice(bytes.byteOffset, bytes.byteOffset + bytes.byteLength);
const model = await loadModel(manifest, payload);
const index = model.prepare([
  { id: 'profile', label: 'Profile' },
  { id: 'members', label: 'Members' },
  { id: 'invoices', label: 'Invoices' },
]);
console.log(index.score('coworkers'));
index.dispose();

Browser applications can fetch manifest.json and weights.bin, call loadModel(manifest, await response.arrayBuffer()), and reuse a prepared index. The loader verifies the payload hash and manifest format. See prepared-index documentation for lifecycle and input limits.

Evidence and limitations

The selected scale adjustment improved top-one accuracy on 600 reserved synthetic typo cases from 64.33% to 67.33%, while retaining the 32 KiB payload. These historical results are documented in typo-weight experiments, with original evaluation receipts under eval/. They do not establish broad semantic generalization, human-reviewed relevance, or reliable abstention. The spelling holdout is now consumed.

Hash collisions, bag-of-features pooling, short unseen labels, ambiguous menus, and action opposites remain limitations. Preserve lexical precedence and treat learned results as suggestions. Documentation includes historical checkpoints; this repository ships only the current demo artifact identified above.

Provenance and license

Copied from source commit 0277121c3e4127c96604b25c73a9af627663da80, artifact path packages/model/experiments/navigation-align0p5-seed29-word085. provenance.json records SHA-256 hashes of copied files. Weight SHA-256: fcb4be56a9a8ddb82e02de9ae31f0822e2e211b4dc65e61cc24dc285620dd6a2.

Project code and original model artifacts are distributed under the source project's MIT license. Training and research sources retain their own terms: CLINC150 (CC BY 3.0), BANKING77 (CC BY 4.0), VS Code (MIT), and navigation metadata from GNOME, KDE, and Xfce. See data attribution, navigation provenance, and credits. Raw third-party datasets and other experimental checkpoints are not included.