File size: 16,964 Bytes
39371ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9c380c2
39371ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36a016c
39371ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
# Pi in the browser

A static web app running **Pi’s actual terminal UI in xterm.js**, **the real Pi Agent**, **just-bash**, and **our verified MiniCPM5-2B q4f16 conversion** in the browser. A dedicated worker owns the agent loop, WebGPU inference, and virtual shell. No inference API, API key, shell server, or Pi RPC server is involved.

```sh
cd app
npm ci
npm run build
npm run preview
```

Open the printed localhost URL and send your first message. If the complete verified model is already cached, Pi restores it and sends your message automatically without a confirmation dialog. Otherwise, a **Load model & send** dialog explains the download and closes immediately on confirmation. Progress appears in the top bar while your message stays visible as queued; it sends automatically once the model is ready. Closing the dialog or pressing Stop returns the unsent message to the editor. Loading errors reopen the dialog for retry. Automatic restoration never downloads missing model files: an incomplete or changed cache requires confirmation. You can also load ahead of time with **Load model**. The default downloads our pinned model from Hugging Face. The app needs a secure context (HTTPS or localhost), WebGPU with `shader-f16` in a worker, and OPFS storage. Allow about 2 GB of free browser storage plus device memory for inference. Actual memory availability matters on phones.

For development, use `npm run dev`. To use the already converted weights on this computer instead of downloading them, run `python scripts/serve.py` from the repository root in another terminal, then open `http://127.0.0.1:5173/?source=local`. The development and preview servers forward `/models` to that static file server. The local source option is accepted only on localhost/127.0.0.1.

## Features

- Pi’s upstream fullscreen renderer, multiline editor, user/assistant message components, Markdown, prompt history, and autocomplete. xterm.js renders their ANSI output.
- `/load`, `/help`, `/new`, and `!command` browser commands. Enter sends, Shift+Enter inserts a newline, Escape/Ctrl+C stops, Ctrl+O expands tools. Ctrl+P/N recalls prompts; Ctrl+Shift+F searches the transcript. The Send and Newline buttons also support touch input.
- Streaming conversation with genuine `@earendil-works/pi-agent-core` tool execution and automatic follow-up turns.
- `bash`, `read`, `write`, and exact-match `edit` tools connected to the same just-bash filesystem.
- File editor with Prism syntax colors for HTML/XML/SVG, CSS, JavaScript/TypeScript (including JSX/TSX), Python, Bash, JSON, Markdown, YAML, TOML, SQL, CSV, C/C++, C#, Java, Go, Rust, Dockerfiles, and diffs. Highlighting runs in a separate, lazily started worker and is debounced while typing. Native textarea editing, selection, undo, and scrolling are preserved. Unknown formats, files above 200,000 characters or 5,000 lines, excessive token output, and grammar work exceeding 1.5 seconds fall back to plain text. Text-file import, workspace export, and an interactive command entry are also included.
- The starter workspace pairs `sales.csv` (products and units sold) with `products.json` (product prices). Both are needed for total sales revenue. The cat suggests “What's our total sales revenue?” without naming files or tools; the workspace README contains no calculation instructions. Only an untouched earlier pair is upgraded together; custom data and prior deletions are preserved.
- Successful `read` tool calls open and highlight the resolved file in the workspace, including relative paths and symlinks. Following waits for newly created files, preserves unsaved edits, and keeps the conversation’s focus and page scroll position. Common file types use a local subset of 18 Seti icons from VS Code’s built-in theme, totaling about 12 KB of SVGs.
- The shell highlights Bash as you type, accepts multiline commands (Shift+Enter), recalls session history (↑/↓), and separates command/output blocks with colored exit states and errors. Clear or Ctrl+L clears the output. Its output area can be resized vertically.
- Model name, loading progress, and loaded status live in a rounded, inset top bar aligned with the workspace width. New chat sits at the bottom left of the terminal, opposite the send controls. Panels use background shading and rounded corners instead of separator borders.
- The top bar’s activity indicator labels model inference **WebGPU** and tool execution **JavaScript**. It shows loading, prompt processing, tool execution, and live generation speed. Its sparkline uses actual generated-token counts, including thinking, over a rolling second; prompt-processing time is excluded. Rates refresh four times per second and clear when generation stops. The capsule fits its current contents with a 220ms width transition, disabled when reduced motion is preferred. Nested corner radii account for their insets: the header uses a 16px radius around 8px-radius controls with 8px padding.
- The first-message dialog explains that the model, inference, and shell run entirely in the browser. Pi’s input receives focus when the workspace starts; cancelling setup returns focus and the unsent draft to Pi.
- The download dialog has a matching peeking cat with transparent paws over its top edge. The sprite preloads and decodes on page startup, stays decorative for keyboard/screen-reader users, and shrinks on short screens to keep the download note and actions visible. Subtle paw prints appear around the overlay whenever the dialog is open, including before confirmation and on retry. They do not intercept clicks, pause in hidden tabs, and stay still with reduced motion.
- After 7 seconds of inactivity, a faint paw trail occasionally crosses Pi’s transcript, with 45–60 seconds between trails. Typing, clicking, or scrolling clears it immediately; pointer movement neither resets the delay nor stops the trail. It stays above the input, never takes focus, and pauses during model activity, open dialogs, hidden or unfocused tabs, and reduced motion.
- An empty chat has a small blocky cat with a clickable starter prompt. The bubble and cat appear together once the image is decoded. The suggestion fills Pi’s editor without sending, adapts to the workspace files, and disappears after the first user message. The 104px mascot makes one gentle hop shortly after appearing, then occasionally blinks; clicking or tapping the cat replays the hop and a short meow. Both cats are decorative and never take keyboard focus. Motion pauses while hidden and is disabled for reduced-motion preferences. Its paws align with the input divider, with room reserved beside it for the status text. Its position follows the terminal’s available blank rows, keeping the input and transcript unobstructed; short viewports omit it when there is no room.
- IndexedDB persistence for chat, file bytes, executable modes, directories, and symbolic links. New chat keeps the workspace.
- Verified model downloading: six concurrent files, 32 MiB byte ranges, three attempts per chunk, cancellation, incremental SHA-256 checks, and atomic OPFS writes. Completed files survive an interrupted load; an incomplete file restarts on retry.
- Cached reloads use file size, modification time, and the stored verification receipt. Model files are hashed during download; they are not rehashed on every reload.
- Visible download / WebGPU setup / shader warm-up phases. No image requests from model-generated Markdown.
- MiniCPM5-2B's supported thinking mode, with completed reasoning retained separately in conversation history. Thoughts are collapsed in the terminal; click to expand. Tool-shaped text inside reasoning is never executed.
- Temperature 1, actual top-p 0.95 sampling, no top-k/min-p filtering, and repetition penalty 1. A tested adapter applies nucleus filtering because Transformers.js 4.2.0 does not implement its exposed `top_p` option in generation.
- An 8,192-token combined prompt/output budget, up to 2,048 generated tokens (reasoning plus answer/tool XML) per model call, and a 12-turn limit per user request. Older complete user turns can be omitted from inference context; the full saved conversation is retained. An oversized current turn produces an explicit error.
- The terminal footer shows used/total context tokens and remaining capacity for the latest model call. It includes the formatted prompt (instructions, tool definitions, and retained history) plus generated tokens, updates during streaming, survives reloads, and resets with New chat. Unsent draft text is counted when submitted; saved turns excluded from the model context are not counted. On narrow screens the counter gets its own line.

## Model quality

The [September 14 multi-turn audit](../research/multiturn-audit/REPORT.md) found that the previous no-thinking/greedy profile was not faithful to the model's documented configuration. The corrected adapter handles reasoning and sampling, but the current 4-bit conversion still fails some HTML/CDN follow-ups that the original BF16 model completed. It can produce ineffective edits or exhaust its reasoning budget. Numerical conversion checks and basic tool tests do not establish general coding reliability. Cross-call KV reuse and physical-phone validation remain outstanding.
- Responsive desktop and mobile layouts. The file workspace follows the chat on narrow screens.
- A consistent dark theme across the terminal, header, model controls, file editor, shell, and setup dialog, including native form controls and scrollbars.

## Runtime boundaries

This uses Pi's actual TUI and agent core with browser host adapters. The npm `pi-tui` package supplies the renderer, editor, input parser, layout, Markdown, search, and keybindings. Unmodified CLI presentation components are in [vendor/pi-coding-agent](vendor/pi-coding-agent/README.md), with provenance checksums and the upstream dark palette. Secondary text colors are lifted for legibility at small sizes. [The build adapter](scripts/pi-browser.mjs) replaces terminal host services; [the terminal adapter](src/pi/xterm-terminal.mjs) connects xterm input/output/resize to Pi. Browser command dispatch and just-bash tool cards are app code.

Message padding follows Pi 0.85.1 defaults: one column of output padding, one inner row above/below user and tool boxes, one separator row between turns, and zero editor column padding. The browser panel adds an 8px outer inset, with no right padding on desktop. Pi owns transcript scrolling, so xterm's unused native scrollbar gutter is disabled. The pinned `@xterm/addon-unicode-graphemes@0.4.0` addon matches Pi's grapheme widths for CJK, emoji sequences and combining accents. It uses xterm's proposed Unicode API and is covered by rendering assertions.

The full Node CLI, host filesystem, OAuth providers, extensions, package installation, and RPC server are not included. xterm.js is the terminal emulator; it does not add Node.js or a native shell. Model inference and tool execution remain in the browser worker. Terminal control sequences from model/file output are stripped before rendering, and Markdown images do not trigger downloads.

just-bash is a simulated shell, not a Linux VM. Shell scripts, pipes, redirects, `awk`, `jq`, `sed`, `grep`, `find`, and other supported text commands execute in JavaScript. Node.js, Python, npm, native executables, and SQLite are unavailable. This HTTP experiment enables curl for any HTTPS URL through a browser fetch adapter; see public/http-help.html for examples and limits. The installed browser bundle also excludes working gzip/zlib commands. Shell cwd and variables reset between commands; filesystem changes persist. Commands have a 10-second deadline, 3,000-command cap, 16 MiB filesystem budget, and 256 KiB output budget. Tool output passed to the model is capped at 6,000 characters.

Local model and filesystem tool calls can run with the browser network disabled after loading. HTTP commands require a connection and CORS permission from the destination. Reopening the app itself still requires its static HTML/JS/WASM assets; there is no offline service worker. Browser storage belongs to this site's origin and may be evicted by the browser. Export important work.

The small 4-bit model can make mistakes. These tests establish working integration on the tested device, not broad coding-agent reliability or support for every mobile browser. No physical phone has been tested.

## Verification

```sh
npm test
npm run build
# Start npm run preview, then run the terminal checks (no model needed):
node tests/terminal.mjs
# Deterministic presentation fixtures, including narrow and short viewports:
npm run test:render
# Ten consecutive actual-model turns (local model server required):
npm run test:conversation
# First message, modal cancellation/retry, and actual model auto-send:
npm run test:first-message
# Download confirmation, top-bar progress, queued messages and cancellation (controlled worker fixtures):
npm run test:load-flow
# Cached auto-start, invalidated cache, Stop and GPU failure recovery:
npm run test:cached-start
# Live token speed, generation/tool/idle phases and eight header layouts:
npm run test:runtime-meter
# Automatic read following, unsaved edits, file icons and mobile layout:
npm run test:follow-read
# File highlighting, native editing, saving, scrolling and large-file fallback:
npm run test:editor
# Companion product data and preserving existing workspaces during its addition:
npm run test:starter-products
# Opt-in real WebGPU model evaluation; requires the dedicated cached test profile:
npm run test:sales-demo
# Syntax-colored command input, actual shell execution, history and layouts:
npm run test:shell
# Actual token usage, streaming, cancellation, reload and input focus:
npm run test:context
# Start npm run preview -- --port 4173 and the root static model server first:
npm run test:browser
```

The terminal suite checks real keyboard events, multiline editing, autocomplete, history, Unicode and large clipboard pastes, shell execution, and live resize. The browser suite uses real Chromium, the actual model, and the production bundle; prompts are submitted through xterm and Pi’s editor. It disables networking while Pi reads a random value from an unseen file, calculates CSV revenue, writes and verifies a report, creates and executes a shell script, edits it, and recovers from a missing-file tool error. It also checks stopping generation, file import/edit/save, manual shell execution, workspace persistence, cached reloads, and a 390 px layout.

To exercise the public Hugging Face download in a separate browser profile:

```sh
APP_URL=http://127.0.0.1:4173/ \
BROWSER_PROFILE=../validation/agent-browser-profile/remote \
TEST_LABEL=agent-browser-hub TEST_CANCEL_LOAD=1 npm run test:browser
```

To exercise smaller GPU buffer allocations on the desktop GPU:

```sh
APP_URL='http://127.0.0.1:4173/?source=local&limit128' \
TEST_LABEL=agent-browser-limit128 npm run test:browser
```

`limit128` sets the actual WebGPU device to 128 MiB storage bindings and 256 MiB buffers. It is an allocation compatibility test; viewport resizing and GPU limits do not emulate mobile hardware performance.

Results and screenshots are in the repository's `validation/agent-browser*` files. Test logs include complete model tool calls and results. The browser profile and downloaded weights are ignored by Git.

## Static hosting

Upload **the contents of `app/dist`** to an HTTPS static host. It contains the app and local ONNX Runtime assets; weights are fetched from the pinned public Hugging Face repository and cached in OPFS. No build or inference server is needed after deployment.

This experiment is deployed at [Mike0021/MiniCPM5-2B-WebGPU-Pi-HTTP](https://huggingface.co/spaces/Mike0021/MiniCPM5-2B-WebGPU-Pi-HTTP). From this checkout, `python3 scripts/stage.py` stages the build, source, licenses, and HTTP verification results in `.space-http/`.

Browsers that deny storage in third-party embeds get an **Open workspace in a new tab** link. The top-level Space app can then use its own storage without changing browser privacy settings. Embedded-page tests restore normal Chromium storage partitioning because Playwright 1.58 disables it by default; the restricted-storage path is tested separately.

## Pinned components

| Component | Version / revision |
|---|---|
| Pi agent core and Pi AI types/event stream | `@earendil-works/*` 0.85.1 |
| Pi TUI / CLI presentation components | 0.85.1 |
| xterm.js / fit addon | 6.0.0 / 0.11.0 |
| just-bash | 3.4.2, explicit `just-bash/browser` import |
| Transformers.js | 4.2.0 |
| Model | `Mike0021/MiniCPM5-2B-ONNX` at `04a6c49fcba3a65a0351c92644c3a7e9d4343059` |
| Model provenance | official OpenBMB weights; see root README and package manifest |

Architecture decisions and primary-source research are in [the investigation](../research/browser-agent/REPORT.md).