InferScale-Sim / README.md
ArchitSharma's picture
Finalize InferScale-Sim
4649014
|
Raw History Blame Contribute Delete
10.6 kB
---
title: InferScale-Sim
colorFrom: indigo
colorTo: blue
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Interactive LLM serving simulator and SLO planner
---
# InferScale-Sim
A browser-based research workbench for studying LLM serving systems without provisioning a GPU.
The simulator is written in Python. Hugging Face serves a static site; Pyodide runs the same `src/inferscale` package inside a Web Worker on the visitor's CPU. There is no backend service, model download, API key, or server-side accelerator.
> **Timing scope.** The bundled L4/A10G/A100 profiles are analytical references, not measured hardware benchmarks. Queueing, scheduling, KV-cache, transfer, SLO, and policy behavior is simulated live. Hardware claims require imported measurements and held-out validation.
## Research question
InferScale asks how serving-policy choices interact with workload shape and resource constraints:
- when continuous batching or chunked prefill changes tail latency;
- when prefill/decode disaggregation helps enough to justify extra accelerators and KV transfer;
- when reusable prefixes improve goodput versus consume scarce cache capacity;
- how stateful agent sessions change KV-retention and routing decisions;
- whether online prediction improves serving outcomes, not merely prediction accuracy;
- which conclusions remain stable across matched workload seeds and timing uncertainty.
The motivation is practical. Microsoft's [Vidur](https://arxiv.org/abs/2405.05465) reported an example LLaMA2-70B configuration search that took roughly one CPU-hour through simulation versus an estimated 42,000 GPU-hours (~$218K) using deployment-based exploration.
## What is implemented
### Stateless serving
- open-loop constant, Poisson, bursty, and exact trace-replay workloads;
- static and continuous batching;
- FCFS, shortest-job-first, least-slack/SLO-aware scheduling;
- chunked prefill;
- paged KV-cache accounting and memory admission;
- controlled shared-prefix reuse;
- colocated and prefill/decode-disaggregated topologies;
- independent P/D worker counts, accelerator profiles, and serialized KV-transfer cost;
- TTFT, TPOT, E2E, queueing, throughput, goodput, SLO attainment, and KV telemetry.
### Stateful sessions
- multi-turn session dependencies separated by sampled tool gaps;
- HBM retention, TTL expiry, host-memory offload/restore, and pressure eviction;
- least-load, strict-affinity, and bounded-affinity routing;
- cross-turn cache hits, history recomputation, HBM/host GB-seconds, and session SLOs;
- online global/per-tool EWMA tool-gap prediction under a controlled distribution shift.
### Execution learning
- online first-order role-transition learning with no future-trace lookahead;
- cumulative and decayed transition models;
- confidence-gated top-1 prefetch;
- multi-step transition rollout and top-k prefetch;
- utility-aware prefetch using transfer and forecast-weighted eviction costs;
- Brier score, log loss, ECE, reliability diagrams, future-role recall@K, prefetch utilization, and speculative-transfer waste;
- confidence, forgetting-rate, forecast-horizon, and cache-budget studies.
### Research consolidation
- common-random-number A/B experiments;
- paired bootstrap intervals;
- timing-model sensitivity analysis;
- repeated-seed policy ranking;
- Pareto stability across latency, HBM residency, and unused speculative transfer;
- a bounded full-trace serving oracle over a declared policy/action-plan family;
- per-policy regret to that bounded reference;
- vLLM/SGLang-style measurement import, calibration, and held-out validation hooks;
- Markdown/JSON/CSV/PNG research-artifact export.
## Reference result
The repository includes a deterministic 12-seed analytical-reference consolidation study in [`reports/reference_consolidation.md`](reports/reference_consolidation.md).
| Policy | Median p95 TTFT | TTFT wins | Pareto stability | Median oracle regret |
|---|---:|---:|---:|---:|
| Top-1 decayed | 6,474.9 ms | 41.7% | 75.0% | 59.3 ms |
| No prefetch | 6,490.7 ms | 41.7% | **91.7%** | **34.7 ms** |
| Utility-aware multi-step | 6,507.5 ms | 8.3% | 58.3% | 133.1 ms |
| Multi-step top-k | 6,640.3 ms | 8.3% | 83.3% | 215.0 ms |
`Top-1 decayed` is the robust TTFT winner under the configured ranking protocol, while `No prefetch` is Pareto-stable on more seeds and has lower median regret to the bounded oracle. The point is not to collapse the study to one score: prediction quality, latency, speculative transfer, and HBM pressure can prefer different policies.
The bounded oracle searches 19 candidate plans per seed and has median p95 TTFT **6,408.6 ms** in this reference study. It is an information upper bound over that declared family, not a proof of globally optimal cache scheduling.
## Interface map
The public Space is organized as a research tool rather than a product dashboard:
| Section | Purpose |
|---|---|
| **Serving** | Inspect one workload and request-level behavior. |
| **Schedulers** | Compare colocated schedulers on an identical trace. |
| **Capacity** | Find the highest repeated SLO-compliant offered load. |
| **P/D + Cache** | Compare colocated/P-D serving with and without prefix reuse. |
| **Design Space** | Sweep bounded configurations and inspect performance/efficiency Pareto fronts. |
| **A/B Studies** | Run paired bootstrap and timing-sensitivity experiments. |
| **Stateful Sessions** | Study KV retention, host offload, affinity, and tool-gap prediction. |
| **Execution Model** | Study transition learning, multi-step prefetch, calibration, and cache pressure. |
| **Evidence** | Run repeated-seed consolidation and import external measurements. |
| **Methods** | Read the simulator assumptions and research lineage. |
## Architecture
```text
Hugging Face Static Space
|
| serves HTML / CSS / JS / Python sources
v
Browser
|
+-- UI + Chart.js
|
+-- Web Worker
|
+-- Pyodide
|
+-- src/inferscale
|
+-- discrete-event simulation
+-- research protocols
+-- calibration / reporting
```
The browser mirror under `py/inferscale/` is generated from `src/inferscale/`; `scripts/release_check.py` fails if the copies diverge.
## Run locally
```bash
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -U pip
pip install -e '.[dev]'
pytest -q
python scripts/release_check.py
```
For the static UI, serve the repository root with any local HTTP server:
```bash
python -m http.server 8000
```
The deployed Space needs no secrets.
## Command-line examples
Run a single simulation:
```bash
python scripts/run_simulation.py examples/balanced.json
```
Import a serving-benchmark artifact:
```bash
python scripts/import_measurements.py path/to/benchmark.json --format auto
```
Calibrate analytical timing scales and evaluate held-out cases:
```bash
python scripts/calibrate_profiles.py measurements.json --holdout 0.25
```
Generate a research report:
```bash
python scripts/generate_report.py robust-policy-study.json \
--calibration calibration_result.json
```
## Repository layout
```text
src/inferscale/ canonical Python simulator and research code
py/inferscale/ browser mirror loaded by Pyodide
index.html static workbench structure
styles.css project-specific design system
app.js UI, charts, exports, worker orchestration
worker.mjs Pyodide bootstrap and Python action bridge
scripts/ release, import, calibration, report utilities
tests/ deterministic unit/integration coverage
docs/ architecture, methodology, validation, research notes
examples/ workload and validation schemas
reports/ reference analytical study
```
## Validation
The release suite checks both simulation behavior and deployment invariants, including:
- stateless and P/D smoke tests;
- trace replay;
- prefix reuse and design-space Pareto logic;
- stateful-session, host-tier, HBM-pressure, and affinity experiments;
- predictive tiering and execution-learning studies;
- paired bootstrap, timing sensitivity, repeated-seed consolidation, oracle regret;
- measurement import/calibration/report generation;
- source/browser Python parity;
- HF metadata and public provenance guardrails;
- DOM-reference integrity and chart-export controls.
Run:
```bash
pytest -q
python scripts/release_check.py
python -m compileall -q src scripts
node --check app.js
node --check worker.mjs
```
## Design and implementation choices
The interface deliberately avoids a product/SaaS visual language. It uses one restrained accent, flat surfaces, square geometry, system typography, and editorial hierarchy instead of gradients, floating cards, badge-heavy status UI, or decorative marketing sections. The rationale and reusable tokens are documented in [`docs/design.md`](docs/design.md).
The simulator core is dependency-light Python rather than a simulation framework. Events, requests, schedulers, caches, predictors, and studies remain inspectable from source and testable outside the browser.
## Limitations
1. Bundled device timings are analytical reference profiles.
2. The agent-session simulator intentionally isolates state/routing effects from the full dynamic-batching model used in stateless serving.
3. Transfer models are simplified serialized bandwidth + base-latency abstractions, not packet/NCCL/NIXL simulators.
4. Oracle policies are bounded information references over declared action families.
5. Calibration can correct global timing bias; it does not establish fidelity on unseen models, hardware, schedulers, or workload regimes.
These limitations are surfaced in the UI and exported reports rather than hidden.
## Research lineage
InferScale is informed by work on serving simulation, disaggregation, stateful agent workloads, and predictive cache management. See [`docs/research.md`](docs/research.md) for the annotated list and the precise distinction between implemented abstractions and cited systems.
Key starting points include:
- [Vidur](https://arxiv.org/abs/2405.05465) — predictive profiling and workload-aware serving simulation.
- [SGLang / RadixAttention](https://arxiv.org/abs/2312.07104) — structured prefix reuse and scheduling motivation.
- Recent 2026 work discussed in `docs/research.md` on P/D disaggregation, GPU-free emulation, agent-session serving, KV tiering, routing locality, and predictive prefetch.
## License
MIT.