File size: 10,640 Bytes
400054a
0c6c82c
 
 
 
 
400054a
0c6c82c
 
400054a
 
ce2d64b
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
 
 
 
 
 
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
 
 
 
 
 
 
 
 
0c6c82c
4649014
ce2d64b
4649014
 
 
 
 
ce2d64b
4649014
44745f2
4649014
 
 
 
 
 
 
44745f2
4649014
ce2d64b
4649014
 
 
 
 
 
 
 
 
44745f2
4649014
ce2d64b
4649014
ce2d64b
4649014
 
 
 
 
 
b799d1d
4649014
8f91935
4649014
b799d1d
4649014
b799d1d
4649014
8f91935
4649014
 
 
 
 
 
 
 
 
 
 
 
9916edb
0c6c82c
 
 
4649014
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44745f2
0c6c82c
 
 
4649014
0c6c82c
4649014
0c6c82c
 
 
 
 
 
4649014
0c6c82c
 
4649014
0c6c82c
 
4649014
9916edb
4649014
9916edb
4649014
9916edb
 
4649014
9916edb
 
4649014
0c6c82c
 
4649014
0c6c82c
 
4649014
0c6c82c
4649014
 
 
0c6c82c
4649014
0c6c82c
4649014
 
 
0c6c82c
 
4649014
0c6c82c
4649014
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c6c82c
 
 
44745f2
4649014
 
 
0c6c82c
 
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
 
 
 
 
0c6c82c
4649014
 
 
0c6c82c
4649014
0c6c82c
4649014
0c6c82c
4649014
 
 
0c6c82c
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
---
title: InferScale-Sim
colorFrom: indigo
colorTo: blue
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Interactive LLM serving simulator and SLO planner
---

# InferScale-Sim

A browser-based research workbench for studying LLM serving systems without provisioning a GPU.

The simulator is written in Python. Hugging Face serves a static site; Pyodide runs the same `src/inferscale` package inside a Web Worker on the visitor's CPU. There is no backend service, model download, API key, or server-side accelerator.

> **Timing scope.** The bundled L4/A10G/A100 profiles are analytical references, not measured hardware benchmarks. Queueing, scheduling, KV-cache, transfer, SLO, and policy behavior is simulated live. Hardware claims require imported measurements and held-out validation.

## Research question

InferScale asks how serving-policy choices interact with workload shape and resource constraints:

- when continuous batching or chunked prefill changes tail latency;
- when prefill/decode disaggregation helps enough to justify extra accelerators and KV transfer;
- when reusable prefixes improve goodput versus consume scarce cache capacity;
- how stateful agent sessions change KV-retention and routing decisions;
- whether online prediction improves serving outcomes, not merely prediction accuracy;
- which conclusions remain stable across matched workload seeds and timing uncertainty.

The motivation is practical. Microsoft's [Vidur](https://arxiv.org/abs/2405.05465) reported an example LLaMA2-70B configuration search that took roughly one CPU-hour through simulation versus an estimated 42,000 GPU-hours (~$218K) using deployment-based exploration.

## What is implemented

### Stateless serving

- open-loop constant, Poisson, bursty, and exact trace-replay workloads;
- static and continuous batching;
- FCFS, shortest-job-first, least-slack/SLO-aware scheduling;
- chunked prefill;
- paged KV-cache accounting and memory admission;
- controlled shared-prefix reuse;
- colocated and prefill/decode-disaggregated topologies;
- independent P/D worker counts, accelerator profiles, and serialized KV-transfer cost;
- TTFT, TPOT, E2E, queueing, throughput, goodput, SLO attainment, and KV telemetry.

### Stateful sessions

- multi-turn session dependencies separated by sampled tool gaps;
- HBM retention, TTL expiry, host-memory offload/restore, and pressure eviction;
- least-load, strict-affinity, and bounded-affinity routing;
- cross-turn cache hits, history recomputation, HBM/host GB-seconds, and session SLOs;
- online global/per-tool EWMA tool-gap prediction under a controlled distribution shift.

### Execution learning

- online first-order role-transition learning with no future-trace lookahead;
- cumulative and decayed transition models;
- confidence-gated top-1 prefetch;
- multi-step transition rollout and top-k prefetch;
- utility-aware prefetch using transfer and forecast-weighted eviction costs;
- Brier score, log loss, ECE, reliability diagrams, future-role recall@K, prefetch utilization, and speculative-transfer waste;
- confidence, forgetting-rate, forecast-horizon, and cache-budget studies.

### Research consolidation

- common-random-number A/B experiments;
- paired bootstrap intervals;
- timing-model sensitivity analysis;
- repeated-seed policy ranking;
- Pareto stability across latency, HBM residency, and unused speculative transfer;
- a bounded full-trace serving oracle over a declared policy/action-plan family;
- per-policy regret to that bounded reference;
- vLLM/SGLang-style measurement import, calibration, and held-out validation hooks;
- Markdown/JSON/CSV/PNG research-artifact export.

## Reference result

The repository includes a deterministic 12-seed analytical-reference consolidation study in [`reports/reference_consolidation.md`](reports/reference_consolidation.md).

| Policy | Median p95 TTFT | TTFT wins | Pareto stability | Median oracle regret |
|---|---:|---:|---:|---:|
| Top-1 decayed | 6,474.9 ms | 41.7% | 75.0% | 59.3 ms |
| No prefetch | 6,490.7 ms | 41.7% | **91.7%** | **34.7 ms** |
| Utility-aware multi-step | 6,507.5 ms | 8.3% | 58.3% | 133.1 ms |
| Multi-step top-k | 6,640.3 ms | 8.3% | 83.3% | 215.0 ms |

`Top-1 decayed` is the robust TTFT winner under the configured ranking protocol, while `No prefetch` is Pareto-stable on more seeds and has lower median regret to the bounded oracle. The point is not to collapse the study to one score: prediction quality, latency, speculative transfer, and HBM pressure can prefer different policies.

The bounded oracle searches 19 candidate plans per seed and has median p95 TTFT **6,408.6 ms** in this reference study. It is an information upper bound over that declared family, not a proof of globally optimal cache scheduling.

## Interface map

The public Space is organized as a research tool rather than a product dashboard:

| Section | Purpose |
|---|---|
| **Serving** | Inspect one workload and request-level behavior. |
| **Schedulers** | Compare colocated schedulers on an identical trace. |
| **Capacity** | Find the highest repeated SLO-compliant offered load. |
| **P/D + Cache** | Compare colocated/P-D serving with and without prefix reuse. |
| **Design Space** | Sweep bounded configurations and inspect performance/efficiency Pareto fronts. |
| **A/B Studies** | Run paired bootstrap and timing-sensitivity experiments. |
| **Stateful Sessions** | Study KV retention, host offload, affinity, and tool-gap prediction. |
| **Execution Model** | Study transition learning, multi-step prefetch, calibration, and cache pressure. |
| **Evidence** | Run repeated-seed consolidation and import external measurements. |
| **Methods** | Read the simulator assumptions and research lineage. |

## Architecture

```text
Hugging Face Static Space
        |
        | serves HTML / CSS / JS / Python sources
        v
Browser
  |
  +-- UI + Chart.js
  |
  +-- Web Worker
        |
        +-- Pyodide
              |
              +-- src/inferscale
                    |
                    +-- discrete-event simulation
                    +-- research protocols
                    +-- calibration / reporting
```

The browser mirror under `py/inferscale/` is generated from `src/inferscale/`; `scripts/release_check.py` fails if the copies diverge.

## Run locally

```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
pip install -e '.[dev]'

pytest -q
python scripts/release_check.py
```

For the static UI, serve the repository root with any local HTTP server:

```bash
python -m http.server 8000
```

The deployed Space needs no secrets.

## Command-line examples

Run a single simulation:

```bash
python scripts/run_simulation.py examples/balanced.json
```

Import a serving-benchmark artifact:

```bash
python scripts/import_measurements.py path/to/benchmark.json --format auto
```

Calibrate analytical timing scales and evaluate held-out cases:

```bash
python scripts/calibrate_profiles.py measurements.json --holdout 0.25
```

Generate a research report:

```bash
python scripts/generate_report.py robust-policy-study.json \
  --calibration calibration_result.json
```

## Repository layout

```text
src/inferscale/         canonical Python simulator and research code
py/inferscale/          browser mirror loaded by Pyodide
index.html              static workbench structure
styles.css              project-specific design system
app.js                  UI, charts, exports, worker orchestration
worker.mjs              Pyodide bootstrap and Python action bridge
scripts/                release, import, calibration, report utilities
tests/                  deterministic unit/integration coverage
docs/                   architecture, methodology, validation, research notes
examples/                workload and validation schemas
reports/                 reference analytical study
```

## Validation

The release suite checks both simulation behavior and deployment invariants, including:

- stateless and P/D smoke tests;
- trace replay;
- prefix reuse and design-space Pareto logic;
- stateful-session, host-tier, HBM-pressure, and affinity experiments;
- predictive tiering and execution-learning studies;
- paired bootstrap, timing sensitivity, repeated-seed consolidation, oracle regret;
- measurement import/calibration/report generation;
- source/browser Python parity;
- HF metadata and public provenance guardrails;
- DOM-reference integrity and chart-export controls.

Run:

```bash
pytest -q
python scripts/release_check.py
python -m compileall -q src scripts
node --check app.js
node --check worker.mjs
```

## Design and implementation choices

The interface deliberately avoids a product/SaaS visual language. It uses one restrained accent, flat surfaces, square geometry, system typography, and editorial hierarchy instead of gradients, floating cards, badge-heavy status UI, or decorative marketing sections. The rationale and reusable tokens are documented in [`docs/design.md`](docs/design.md).

The simulator core is dependency-light Python rather than a simulation framework. Events, requests, schedulers, caches, predictors, and studies remain inspectable from source and testable outside the browser.

## Limitations

1. Bundled device timings are analytical reference profiles.
2. The agent-session simulator intentionally isolates state/routing effects from the full dynamic-batching model used in stateless serving.
3. Transfer models are simplified serialized bandwidth + base-latency abstractions, not packet/NCCL/NIXL simulators.
4. Oracle policies are bounded information references over declared action families.
5. Calibration can correct global timing bias; it does not establish fidelity on unseen models, hardware, schedulers, or workload regimes.

These limitations are surfaced in the UI and exported reports rather than hidden.

## Research lineage

InferScale is informed by work on serving simulation, disaggregation, stateful agent workloads, and predictive cache management. See [`docs/research.md`](docs/research.md) for the annotated list and the precise distinction between implemented abstractions and cited systems.

Key starting points include:

- [Vidur](https://arxiv.org/abs/2405.05465) — predictive profiling and workload-aware serving simulation.
- [SGLang / RadixAttention](https://arxiv.org/abs/2312.07104) — structured prefix reuse and scheduling motivation.
- Recent 2026 work discussed in `docs/research.md` on P/D disaggregation, GPU-free emulation, agent-session serving, KV tiering, routing locality, and predictive prefetch.

## License

MIT.