--- title: InferScale-Sim colorFrom: indigo colorTo: blue sdk: static app_file: index.html pinned: false license: mit short_description: Interactive LLM serving simulator and SLO planner --- # InferScale-Sim A browser-based research workbench for studying LLM serving systems without provisioning a GPU. The simulator is written in Python. Hugging Face serves a static site; Pyodide runs the same `src/inferscale` package inside a Web Worker on the visitor's CPU. There is no backend service, model download, API key, or server-side accelerator. > **Timing scope.** The bundled L4/A10G/A100 profiles are analytical references, not measured hardware benchmarks. Queueing, scheduling, KV-cache, transfer, SLO, and policy behavior is simulated live. Hardware claims require imported measurements and held-out validation. ## Research question InferScale asks how serving-policy choices interact with workload shape and resource constraints: - when continuous batching or chunked prefill changes tail latency; - when prefill/decode disaggregation helps enough to justify extra accelerators and KV transfer; - when reusable prefixes improve goodput versus consume scarce cache capacity; - how stateful agent sessions change KV-retention and routing decisions; - whether online prediction improves serving outcomes, not merely prediction accuracy; - which conclusions remain stable across matched workload seeds and timing uncertainty. The motivation is practical. Microsoft's [Vidur](https://arxiv.org/abs/2405.05465) reported an example LLaMA2-70B configuration search that took roughly one CPU-hour through simulation versus an estimated 42,000 GPU-hours (~$218K) using deployment-based exploration. ## What is implemented ### Stateless serving - open-loop constant, Poisson, bursty, and exact trace-replay workloads; - static and continuous batching; - FCFS, shortest-job-first, least-slack/SLO-aware scheduling; - chunked prefill; - paged KV-cache accounting and memory admission; - controlled shared-prefix reuse; - colocated and prefill/decode-disaggregated topologies; - independent P/D worker counts, accelerator profiles, and serialized KV-transfer cost; - TTFT, TPOT, E2E, queueing, throughput, goodput, SLO attainment, and KV telemetry. ### Stateful sessions - multi-turn session dependencies separated by sampled tool gaps; - HBM retention, TTL expiry, host-memory offload/restore, and pressure eviction; - least-load, strict-affinity, and bounded-affinity routing; - cross-turn cache hits, history recomputation, HBM/host GB-seconds, and session SLOs; - online global/per-tool EWMA tool-gap prediction under a controlled distribution shift. ### Execution learning - online first-order role-transition learning with no future-trace lookahead; - cumulative and decayed transition models; - confidence-gated top-1 prefetch; - multi-step transition rollout and top-k prefetch; - utility-aware prefetch using transfer and forecast-weighted eviction costs; - Brier score, log loss, ECE, reliability diagrams, future-role recall@K, prefetch utilization, and speculative-transfer waste; - confidence, forgetting-rate, forecast-horizon, and cache-budget studies. ### Research consolidation - common-random-number A/B experiments; - paired bootstrap intervals; - timing-model sensitivity analysis; - repeated-seed policy ranking; - Pareto stability across latency, HBM residency, and unused speculative transfer; - a bounded full-trace serving oracle over a declared policy/action-plan family; - per-policy regret to that bounded reference; - vLLM/SGLang-style measurement import, calibration, and held-out validation hooks; - Markdown/JSON/CSV/PNG research-artifact export. ## Reference result The repository includes a deterministic 12-seed analytical-reference consolidation study in [`reports/reference_consolidation.md`](reports/reference_consolidation.md). | Policy | Median p95 TTFT | TTFT wins | Pareto stability | Median oracle regret | |---|---:|---:|---:|---:| | Top-1 decayed | 6,474.9 ms | 41.7% | 75.0% | 59.3 ms | | No prefetch | 6,490.7 ms | 41.7% | **91.7%** | **34.7 ms** | | Utility-aware multi-step | 6,507.5 ms | 8.3% | 58.3% | 133.1 ms | | Multi-step top-k | 6,640.3 ms | 8.3% | 83.3% | 215.0 ms | `Top-1 decayed` is the robust TTFT winner under the configured ranking protocol, while `No prefetch` is Pareto-stable on more seeds and has lower median regret to the bounded oracle. The point is not to collapse the study to one score: prediction quality, latency, speculative transfer, and HBM pressure can prefer different policies. The bounded oracle searches 19 candidate plans per seed and has median p95 TTFT **6,408.6 ms** in this reference study. It is an information upper bound over that declared family, not a proof of globally optimal cache scheduling. ## Interface map The public Space is organized as a research tool rather than a product dashboard: | Section | Purpose | |---|---| | **Serving** | Inspect one workload and request-level behavior. | | **Schedulers** | Compare colocated schedulers on an identical trace. | | **Capacity** | Find the highest repeated SLO-compliant offered load. | | **P/D + Cache** | Compare colocated/P-D serving with and without prefix reuse. | | **Design Space** | Sweep bounded configurations and inspect performance/efficiency Pareto fronts. | | **A/B Studies** | Run paired bootstrap and timing-sensitivity experiments. | | **Stateful Sessions** | Study KV retention, host offload, affinity, and tool-gap prediction. | | **Execution Model** | Study transition learning, multi-step prefetch, calibration, and cache pressure. | | **Evidence** | Run repeated-seed consolidation and import external measurements. | | **Methods** | Read the simulator assumptions and research lineage. | ## Architecture ```text Hugging Face Static Space | | serves HTML / CSS / JS / Python sources v Browser | +-- UI + Chart.js | +-- Web Worker | +-- Pyodide | +-- src/inferscale | +-- discrete-event simulation +-- research protocols +-- calibration / reporting ``` The browser mirror under `py/inferscale/` is generated from `src/inferscale/`; `scripts/release_check.py` fails if the copies diverge. ## Run locally ```bash python3 -m venv .venv source .venv/bin/activate python -m pip install -U pip pip install -e '.[dev]' pytest -q python scripts/release_check.py ``` For the static UI, serve the repository root with any local HTTP server: ```bash python -m http.server 8000 ``` The deployed Space needs no secrets. ## Command-line examples Run a single simulation: ```bash python scripts/run_simulation.py examples/balanced.json ``` Import a serving-benchmark artifact: ```bash python scripts/import_measurements.py path/to/benchmark.json --format auto ``` Calibrate analytical timing scales and evaluate held-out cases: ```bash python scripts/calibrate_profiles.py measurements.json --holdout 0.25 ``` Generate a research report: ```bash python scripts/generate_report.py robust-policy-study.json \ --calibration calibration_result.json ``` ## Repository layout ```text src/inferscale/ canonical Python simulator and research code py/inferscale/ browser mirror loaded by Pyodide index.html static workbench structure styles.css project-specific design system app.js UI, charts, exports, worker orchestration worker.mjs Pyodide bootstrap and Python action bridge scripts/ release, import, calibration, report utilities tests/ deterministic unit/integration coverage docs/ architecture, methodology, validation, research notes examples/ workload and validation schemas reports/ reference analytical study ``` ## Validation The release suite checks both simulation behavior and deployment invariants, including: - stateless and P/D smoke tests; - trace replay; - prefix reuse and design-space Pareto logic; - stateful-session, host-tier, HBM-pressure, and affinity experiments; - predictive tiering and execution-learning studies; - paired bootstrap, timing sensitivity, repeated-seed consolidation, oracle regret; - measurement import/calibration/report generation; - source/browser Python parity; - HF metadata and public provenance guardrails; - DOM-reference integrity and chart-export controls. Run: ```bash pytest -q python scripts/release_check.py python -m compileall -q src scripts node --check app.js node --check worker.mjs ``` ## Design and implementation choices The interface deliberately avoids a product/SaaS visual language. It uses one restrained accent, flat surfaces, square geometry, system typography, and editorial hierarchy instead of gradients, floating cards, badge-heavy status UI, or decorative marketing sections. The rationale and reusable tokens are documented in [`docs/design.md`](docs/design.md). The simulator core is dependency-light Python rather than a simulation framework. Events, requests, schedulers, caches, predictors, and studies remain inspectable from source and testable outside the browser. ## Limitations 1. Bundled device timings are analytical reference profiles. 2. The agent-session simulator intentionally isolates state/routing effects from the full dynamic-batching model used in stateless serving. 3. Transfer models are simplified serialized bandwidth + base-latency abstractions, not packet/NCCL/NIXL simulators. 4. Oracle policies are bounded information references over declared action families. 5. Calibration can correct global timing bias; it does not establish fidelity on unseen models, hardware, schedulers, or workload regimes. These limitations are surfaced in the UI and exported reports rather than hidden. ## Research lineage InferScale is informed by work on serving simulation, disaggregation, stateful agent workloads, and predictive cache management. See [`docs/research.md`](docs/research.md) for the annotated list and the precise distinction between implemented abstractions and cited systems. Key starting points include: - [Vidur](https://arxiv.org/abs/2405.05465) — predictive profiling and workload-aware serving simulation. - [SGLang / RadixAttention](https://arxiv.org/abs/2312.07104) — structured prefix reuse and scheduling motivation. - Recent 2026 work discussed in `docs/research.md` on P/D disaggregation, GPU-free emulation, agent-session serving, KV tiering, routing locality, and predictive prefetch. ## License MIT.